fdo-mirrors/mesa

mirror of https://gitlab.freedesktop.org/mesa/mesa.git synced 2026-01-31 00:30:33 +01:00

Author	SHA1	Message	Date
Ben Widawsky	644c8a5151	i965/skl: Add two missing device IDs The Iris part is left unbranded because we did not have these with original SKL. v2: 0x192d is gt3, not gt4 v3: Forgot to update the temporary brand string when I did v2. Cc: "11.0 11.1" <mesa-stable@lists.freedesktop.org Signed-off-by: Ben Widawsky <benjamin.widawsky@intel.com> Acked-by: Michał Winiarski <michal.winiarski@intel.com>	2016-02-17 16:50:59 -08:00
Ilia Mirkin	f3cd62a765	mesa: allow multisampled format info to be returned on GLES 3.1 The restriction on multisampled integer texture formats only applies to GLES 3.0, so don't apply it to GLES 3.1 contexts. This fixes a slew of dEQP-GLES31.functional.state_query.internal_format.* tests, which now all pass. Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu> Reviewed-by: Timothy Arceri <timothy.arceri@collabora.com>	2016-02-17 19:30:40 -05:00
Kristian Høgsberg Kristensen	b8da261dc7	spirv: Fix SpvOpFwidth, SpvOpFwidthFine and SpvOpFwidthCoarse "Result is the same as computing the sum of the absolute values of OpDPdx and OpDPdy on P." We were doing sum of absolute values of OpDPdx of P and OpDPdx of NULL.	2016-02-17 15:28:52 -08:00
Kristian Høgsberg Kristensen	ae3e249d57	anv: Remove hacky PIPE_CONTROL in vkCmdEndRenderPass() The vkCmdPipelineBarrier() command should work as intended now and we need to pull the plug on this old hack.	2016-02-17 15:19:07 -08:00
Kristian Høgsberg Kristensen	5e92e91c61	anv: Rework vkCmdPipelineBarrier() We don't need to look at the stage flags, as we don't really support any fine-grained, stage-level synchronization. We have to do two PIPE_CONTROLs in case we're both flushing and invalidating. Additionally, if we do end up doing two PIPE_CONTROLs, the first, flusing one also has to stall and wait for the flushing to finish, so we don't re-dirty the caches with in-flight rendering after the second PIPE_CONTROL invalidates.	2016-02-17 15:18:06 -08:00
Ben Widawsky	2bf041d94f	i965: Extract push constant state to a new file Every stage has a corresponding 3DSTATE_CONSTANT_XS packet, so having the code to create and emit push constant buffers in genX_vs_state.c is a little strange. Moving it to a separate file seems more logical. v2 [Ken]: Rebase on master, explain motivation in the commit message. Signed-off-by: Ben Widawsky <ben@bwidawsk.net> Reviewed-by: Kenneth Graunke <kenneth@whitecape.org>	2016-02-17 12:34:23 -08:00
Matt Turner	0e9dc59a58	i965: Make emit_minmax return an instruction*. And use it in brw_fs_nir.cpp.	2016-02-17 12:35:27 -08:00
Matt Turner	2f2c00c727	i965: Lower min/max after optimization on Gen4/5. Gen4/5's SEL instruction cannot use conditional modifiers, so min/max are implemented as CMP + SEL. Handling that after optimization lets us CSE more. On Ironlake: total instructions in shared programs: 6426035 -> 6422753 (-0.05%) instructions in affected programs: 326604 -> 323322 (-1.00%) helped: 1411 total cycles in shared programs: 129184700 -> 129101586 (-0.06%) cycles in affected programs: 18950290 -> 18867176 (-0.44%) helped: 2419 HURT: 328 Reviewed-by: Francisco Jerez <currojerez@riseup.net>	2016-02-17 12:35:27 -08:00
Matt Turner	378d98f87e	i965/vec4: Initialize force_writemask_all in vec4_builder(). Reviewed-by: Francisco Jerez <currojerez@riseup.net>	2016-02-17 12:35:27 -08:00
Kristian Høgsberg Kristensen	3b9b908054	anv: Ignore unused dimensions in vkCreateImage We would assert on unused dimensions (eg extent.depth for VK_IMAGE_TYPE_2D) not being 1, but the specification doesn't put any constraints on those. For example, for VK_IMAGE_TYPE_1D: "If imageType is VK_IMAGE_TYPE_1D, the value of extent.width must be less than or equal to the value of VkPhysicalDeviceLimits::maxImageDimension1D, or the value of VkImageFormatProperties::maxExtent.width (as returned by vkGetPhysicalDeviceImageFormatProperties with values of format, type, tiling, usage and flags equal to those in this structure) - whichever is higher" We'll fix up the arguments to isl to keep isl strict in what it expects.	2016-02-17 12:21:51 -08:00
Kristian Høgsberg Kristensen	b63e28c0e1	anv: Set correct write domain on window system BOs We need to make sure GEM understands that we're writing to the BO, in case it needs to synchronize with other rings (blitter use in display server, for example).	2016-02-17 11:19:56 -08:00
Tom Stellard	dc7cf07af3	radeon/llvm: Add TargetLibraryInfo to the pass manager This will prevent optimization passes from introducing unsupported library calls. Tested-by: Michel Dänzer <michel.daenzer@amd.com> Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-02-17 19:06:41 +00:00
Tom Stellard	4f351a6cb1	radeon/llvm: Set the target triple on the module Tested-by: Michel Dänzer <michel.daenzer@amd.com> Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-02-17 19:06:41 +00:00
Tom Stellard	77f4e1c7ff	gallivm: Add helpers for creating and destroying TargetLibraryInfo This functionality is not exposed via the LLVM C API. Tested-by: Michel Dänzer <michel.daenzer@amd.com> Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-02-17 19:06:41 +00:00
Samuel Pitoiset	cfd1dd0500	nvc0: invalidate all buffers when switching pipe contexts Signed-off-by: Samuel Pitoiset <samuel.pitoiset@gmail.com> Reviewed-by: Ilia Mirkin <imirkin@alum.mit.edu>	2016-02-17 21:14:24 +01:00
Ilia Mirkin	49c67926c7	st/mesa: fix up result_src.type when doing i2u/u2i conversions Even though it's a no-op, it's important to keep track of the type so that we can pick the properly-signed op later on. This fixes dEQP-GLES3.functional.shaders.precision.uint.highp_div_fragment, which ended up using IDIV instead of UDIV. Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu> Reviewed-by: Kenneth Graunke <kenneth@whitecape.org> Cc: mesa-stable@lists.freedesktop.org	2016-02-17 13:30:33 -05:00
Brian Paul	5e52df2198	st/mesa: use cso_set_viewport_dims() in try_pbo_upload_common() Note that this results in a different transformation for the viewport's Z axis (depth range), but that doesn't matter for this case. Reviewed-by: Roland Scheidegger <sroland@vmware.com>	2016-02-17 11:25:02 -07:00
Jordan Justen	9a939ebb47	i965/gen7: Use predicated rendering for indirect compute On gen7 (Ivy Bridge, Haswell), we will get a GPU hang if an indirect dispatch is used, but one of the dimensions is 0. Therefore we use predicated rendering on the GPGPU_WALKER command to handle this case. Fixes piglit test: spec/arb_compute_shader/zero-dispatch-size From the ARB_compute_shader spec, under DispatchCompute: "If the work group count in any dimension is zero, no work groups are dispatched." And then for DispatchComputeIndirect: ... "is equivalent (assuming no errors are generated) to calling DispatchCompute with <num_groups_x>, <num_groups_y> and <num_groups_z>" ... Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=94100 Signed-off-by: Jordan Justen <jordan.l.justen@intel.com> Reviewed-by: Ben Widawsky <benjamin.widawsky@intel.com> Reviewed-by: Kenneth Graunke <kenneth@whitecape.org> Tested-by: Ilia Mirkin <imirkin@alum.mit.edu>	2016-02-17 09:25:47 -08:00
Rob Clark	37d540ba70	freedreno: expose time-elapsed query Signed-off-by: Rob Clark <robclark@freedesktop.org>	2016-02-17 10:41:55 -05:00
Rob Clark	ba194630cc	freedreno/a4xx: implement time-elapsed query Signed-off-by: Rob Clark <robclark@freedesktop.org>	2016-02-17 10:41:55 -05:00
Rob Clark	62fa868728	freedreno/a4xx: better occlusion/sample counting This seems to give more reliable results. More similar to what we do on a3xx, although I think it breaks the a3xx theory that the four sets of results map to each MRT (since we appear to still only have four sets on a4xx). The divide-by-two is a bit odd, but seems to be needed for some reason. Signed-off-by: Rob Clark <robclark@freedesktop.org>	2016-02-17 10:41:55 -05:00
Rob Clark	87eb406791	freedreno/query: fix refcnt'ing issue Signed-off-by: Rob Clark <robclark@freedesktop.org>	2016-02-17 10:41:55 -05:00
Rob Clark	0e91dccf9c	freedreno/query: some queries don't have ->begin_query() Signed-off-by: Rob Clark <robclark@freedesktop.org>	2016-02-17 10:41:55 -05:00
Rob Clark	9d23d7b7cb	freedreno/query: align counter snapshot locations Some hw queries need their sample memory locations to have certain alignment. At the moment that isn't an issue, since the only hw query is occlusion, so all samples have the same size. But when others are added with different sample sizes, this starts to be a problem. All current and immediately upcoming hw queries simply need their sample address aligned to their size, so let's use that for now. Signed-off-by: Rob Clark <robclark@freedesktop.org>	2016-02-17 10:41:55 -05:00
Rob Clark	8529e210ec	freedreno/query: add optional enable hook Add enable hook for hw query providers. Some will need to configure perfctr selector registers, which we want to do at the start of the submit. Signed-off-by: Rob Clark <robclark@freedesktop.org>	2016-02-17 10:41:55 -05:00
Rob Clark	45ab5b1c34	freedreno: query max gpu freq This will be needed to support converting from cycle counts to time for performance related queries (initially time-elapsed, but there are some additional performance counters that could be wired up). Signed-off-by: Rob Clark <robclark@freedesktop.org>	2016-02-17 10:41:55 -05:00
Rob Clark	dcb69185a0	freedreno: update generated headers Mostly to pull in perf ctrs. Signed-off-by: Rob Clark <robclark@freedesktop.org>	2016-02-17 10:41:55 -05:00
Rob Clark	2a7ceb5957	freedreno/ir3: fix new gcc6 errors src/gallium/drivers/freedreno/ir3/ir3_compiler_nir.c: In function ‘emit_tex’: src/gallium/drivers/freedreno/ir3/ir3_compiler_nir.c:1368:26: warning: unused variable ‘const_off’ [-Wunused-variable] struct ir3_instruction *const_off[4]; ^~~~~~~~~ unused since: commit `8750299a42` Author: Jason Ekstrand <jason.ekstrand@intel.com> Date: Tue Feb 9 14:51:28 2016 -0800 nir: Remove the const_offset from nir_tex_instr Signed-off-by: Rob Clark <robdclark@gmail.com>	2016-02-17 10:41:55 -05:00
Kristian Høgsberg Kristensen	5caa995c32	Revert "anv: Disable snooping for allocator pools again" This reverts commit `c136672c59`. We still have the intermittent missing flush for VkEvent in certain vulkancts cases: piglit.deqp-vk.api.command_buffers.execute_large_primary piglit.deqp-vk.api.command_buffers.submit_count_non_zero, Let's reenable the snooping until we figure out the root cause.	2016-02-16 23:23:49 -08:00
Kristian Høgsberg Kristensen	ecc67f1aac	anv: Make driver and icd file installable Change the name of the .so to libvulkan_intel.so and add an installable icd with the installed paths. Keep the icd file with build-tree paths, but rename to dev_icd.json to make it clear that it's for development purposes.	2016-02-16 23:23:17 -08:00
Kristian Høgsberg Kristensen	4a2d17f606	anv: Revise PhysicalDeviceFeatures and remove FINISHME	2016-02-16 15:43:12 -08:00
Karol Herbst	edf774bb7e	nv50/ir: we can't do the add to mad conversion when the mul saturates Reviewed-by: Ilia Mirkin <imirkin@alum.mit.edu>	2016-02-16 18:20:10 -05:00
Karol Herbst	068e9848ba	nv50/ir: optimize neg(and(set, 1)) to set helps shaders in saints row IV, bioshock infinite and shadow warrior total instructions in shared programs : 1914931 -> 1903900 (-0.58%) total gprs used in shared programs : 247920 -> 247785 (-0.05%) total local used in shared programs : 5673 -> 5673 (0.00%) total bytes used in shared programs : 17558272 -> 17457320 (-0.57%) local gpr inst bytes helped 0 137 719 719 hurt 0 12 0 0 v2: remove this opt for OP_SLCT and check against float for OP_SET v3: simplified the code Signed-off-by: Karol Herbst <nouveau@karolherbst.de> Reviewed-by: Ilia Mirkin <imirkin@alum.mit.edu>	2016-02-16 18:20:10 -05:00
Ilia Mirkin	ca23c8081f	nv50/ir: fix quadop emission in the presence of predication When there's a predicate, it just goes onto the sources list. If the quadop only has a single regular source, we will end up thinking that the predicate is the second source. Check explicitly for the predSrc so that we don't accidentally emit the wrong thing. This fixes a bunch of dEQP-GLES3.functional.shaders.derivate.* tests. Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu> Cc: mesa-stable@lists.freedesktop.org	2016-02-16 18:20:10 -05:00
Ilia Mirkin	1d1ddfe5f8	nv50,nvc0: enable/disable seamless cubemap texturing as requested In a situation where the seamless setting isn't available on a per-texture basis (G200+ Teslas, and all Fermis), assume that all samplers will have it identically set, and enable accordingly. This fixes arb_seamless_cubemap piglit test on Fermi and Tesla. Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu>	2016-02-16 18:20:10 -05:00
Philipp Zabel	ecd1d94d1c	anv: pCreateInfo->pApplicationInfo parameter to vkCreateInstance may be NULL Fix a NULL pointer dereference in anv_CreateInstance in case the pApplicationInfo field of the supplied VkInstanceCreateInfo structure is NULL [1]. [1] https://www.khronos.org/registry/vulkan/specs/1.0/apispec.html#VkInstanceCreateInfo Signed-off-by: Philipp Zabel <philipp.zabel@gmail.com>	2016-02-16 14:42:26 -08:00
Rob Clark	d49307435a	st/mesa: add missing ETC2 entries to format_map Noticed by Ilia when I was trying to figure out why some app was failing to use ETC2. Signed-off-by: Rob Clark <robclark@freedesktop.org> Reviewed-by: Ilia Mirkin <imirkin@alum.mit.edu>	2016-02-16 15:53:43 -05:00
Samuel Pitoiset	3d5f61a262	nvc0: enable compute support on GK110:GM200 with an envvar Without this NVF0_COMPUTE environment variable, compute support is initialized by default and this is not what we want for now because it might break 3D. It will be enabled by default once we are sure it won't break anything. Please note that compute support on GM200+ is not enabled yet because it needs to be double-checked. Signed-off-by: Samuel Pitoiset <samuel.pitoiset@gmail.com> Reviewed-by: Ilia Mirkin <imirkin@alum.mit.edu>	2016-02-16 21:39:00 +01:00
Samuel Pitoiset	6d74fa5756	nvc0: add compute support for GM107 Fortunately, compute support on GM107 is very close to GK110, except the GK110_COMPUTE.UNK02C4 which is invalid and should not be used. Signed-off-by: Samuel Pitoiset <samuel.pitoiset@gmail.com> Reviewed-by: Ilia Mirkin <imirkin@alum.mit.edu>	2016-02-16 21:39:00 +01:00
Samuel Pitoiset	bc331dd838	nvc0: fix compute state initialization on GK110+ Because our firmware doesn't support the GK110_COMPUTE.FIRMWARE[0x6] method the GPU hangs when it is used. Removing it fix the issue and allow to launch compute shaders on GK110+. Tested on GK208 and GM107. Signed-off-by: Samuel Pitoiset <samuel.pitoiset@gmail.com> Reviewed-by: Ilia Mirkin <imirkin@alum.mit.edu>	2016-02-16 21:39:00 +01:00
Timothy Arceri	a61823b584	glsl: remove duplicate interpolation_string() function We already have one in the IR code that can be used everywhere its needed in the AST code so remove the one from the AST. Reviewed-by: Samuel Iglesias Gonsálvez <siglesias@igalia.com>	2016-02-17 07:26:38 +11:00
Timothy Arceri	e70ece4eea	glsl: remove unused helper Seems to have become unused when i965 moved to NIR. Reviewed-by: Iago Toral Quiroga <itoral@igalia.com>	2016-02-17 07:25:10 +11:00
Timothy Arceri	07e6a37332	glsl: set user defined varyings to smooth by default in ES This is usually handled by the backends in order to handle the various interactions with the gl_*Color built-ins. The problem is this means linking will fail if one side on the interface adds the smooth qualifier to the varying and the other side just uses the default even though they match. This fixes various deqp tests. The spec is not clear what to for desktop GL so leave it as is for now. Reviewed-by: Iago Toral Quiroga <itoral@igalia.com> Reviewed-by: Ian Romanick <ian.d.romanick@intel.com> Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=92743	2016-02-17 07:23:49 +11:00
Samuel Pitoiset	f638512890	gm107/ir: add ATOM CAS emission This fixes the following dEQP test and the other compswap variants. dEQP-GLES31.functional.ssbo.atomic.compswap.highp_int Signed-off-by: Samuel Pitoiset <samuel.pitoiset@gmail.com> Reviewed-by: Ilia Mirkin <imirkin@alum.mit.edu>	2016-02-16 20:53:39 +01:00
Samuel Pitoiset	09446cf5f6	st/mesa: do not init limits when compute shaders are not supported When the number of uniform blocks is less than 12, ARB_uniform_buffer_object can't be enabled and the maximum GL version is not even 3.1... This fixes a regression introduced in `7c79c1e` (st/mesa: add compute shader state) if the maximum number of uniform blocks allowed for compute shaders is less than 12. This happens on Kepler but this might also affect other Gallium drivers. Signed-off-by: Samuel Pitoiset <samuel.pitoiset@gmail.com> Reported-by: Tobias Klausmann <tobias.johannes.klausmann@mni.thm.de> Tested-by: Tobias Klausmann <tobias.johannes.klausmann@mni.thm.de> Reviewed-by: Ilia Mirkin <imirkin@alum.mit.edu> Reviewed-by: Tobias Klausmann <tobias.johannes.klausmann@mni.thm.de>	2016-02-16 20:53:35 +01:00
Jordan Justen	f28d80fabf	mesa: Don't call driver when there is no compute work The ARB_compute_shader spec says: "If the work group count in any dimension is zero, no work groups are dispatched." Signed-off-by: Jordan Justen <jordan.l.justen@intel.com> Reviewed-by: Ilia Mirkin <imirkin@alum.mit.edu>	2016-02-16 09:25:20 -08:00
Jordan Justen	8514c75a26	i965: Set compute shader shared memory max to 64k See Ivy Bridge PRM, Volume 2, Part 2, 1.8.4 INTERFACE_DESCRIPTOR_DATA: DWORD 5, bits 20:16: "This field indicates how much shared local memory the thread group requires. The amount is specified in 4k blocks, but only powers of 2 are allowed: 0, 4k, 8k, 16k, 32k and 64k per half-slice." For Haswell, see Volume 2d, INTERFACE_DESCRIPTOR_DATA: DWORD 5, bits 20:16: With text identical to the Ivy Bridge PRM. For Broadwell, see Volume 2d, INTERFACE_DESCRIPTOR_DATA: DWORD 6, bits 20:16: With text identical to the Ivy Bridge PRM. Signed-off-by: Jordan Justen <jordan.l.justen@intel.com> Reviewed-by: Ben Widawsky <benjamin.widawsky@intel.com>	2016-02-16 09:25:20 -08:00
Brian Paul	f90801cd40	st/mesa: use new CSO_BITS_ALL_SHADERS Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-02-16 10:22:32 -07:00
Brian Paul	1bf8fa8277	cso: add CSO_BITS_ALL_SHADERS For saving/restoring all shader stages. Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2016-02-16 10:22:32 -07:00
Brian Paul	a0636157c4	st/mesa: simplify st->ctx, ctx->st usage in a various places	2016-02-16 10:22:32 -07:00

... 142 143 144 145 146 ...

85652 commits