fdo-mirrors/mesa

mirror of https://gitlab.freedesktop.org/mesa/mesa.git synced 2026-05-27 20:58:11 +02:00

Author	SHA1	Message	Date
Julien Isorce	89eb342def	st/va: retrieve size from the temporary img variable "image" is not ready yet since it will be set at the end of the function by: image = img; Signed-off-by: Julien Isorce <j.isorce@samsung.com> Reviewed-by: Christian K<C3><B6>nig <christian.koenig@amd.com>	2015-12-16 14:12:31 +00:00
Roland Scheidegger	8e195a6251	draw: handle edge flags in llvm path We just ignored them altogether. While this feature is rather old-fashioned supporting it is actually rather trivial. This fixes the associated piglit tests (2 gl-1.0-edgeflag, 2 gl-2.0-edgeflag and all (7) of point-vertex-id). v2: comment fixes, and make the use of the edgeflag in clipmask consistent with when it's actually there (should be impossible to hit a case where the difference would actually matter but still...) Reviewed-by: Brian Paul <brianp@vmware.com>	2015-12-16 03:55:25 +01:00
Roland Scheidegger	13c0b1c780	draw: don't set start_instance and instance id for pt emit This just adds confusion, these parameters are used when fetching vertices by translate, but certainly not when emitting hw vertices for drivers, they make no sense there (setting them has no consequences otherwise since there won't be any elements with instance_divisor set). So just set them to 0 (the draw_pipe_vbuf code for emitting vertices when the draw pipeline is run already does exactly that). Also while here do some whitespace cleanup. Reviewed-by: Brian Paul <brianp@vmware.com>	2015-12-16 03:55:14 +01:00
Samuel Pitoiset	276837cbe4	nvc0: remove old comment related to metric calculations I forgot to remove it when I refactored all performance metrics. Signed-off-by: Samuel Pitoiset <samuel.pitoiset@gmail.com>	2015-12-15 22:49:37 +01:00
Eric Anholt	3858722740	vc4: Add support for dumping executed commands to a file. The VC4_DEBUG=cl,qpu is nice and all, but I want to be able to get more detailed dumps, and to replay the same exact commands in simulation. For that I need a dump with all of the VBOs, shaders, shader recs, etc. This dump can be parsed by vc4-gpu-tools. For now this is only doable from simulator mode, because otherwise we don't have access to the RCL contents generated by the kernel.	2015-12-15 12:05:48 -08:00
Eric Anholt	07570edb98	vc4: Import updated vc4_drm.h with hang state.	2015-12-15 12:02:54 -08:00
Eric Anholt	c5b886b028	vc4: Only update vc4->msaa when the framebuffer changes. Any update here should have been the same as in vc4_set_framebuffer_state(), except for the point where vc4_blit.c temporarily sets different state for its different buffers.	2015-12-15 12:02:53 -08:00
Eric Anholt	f2cf2a63f1	vc4: Don't consider nr_samples==1 surfaces to be MSAA. This is apparently a weirdness of gallium -- nr_samples==1 is occasionally used and means the same thing as nr_samples==0. Fixes a bunch of ARB_framebuffer_srgb blit cases in piglit.	2015-12-15 12:02:53 -08:00
Eric Anholt	da92f16c50	vc4: Fix min() wrapper definition for the simulator's kernel code.	2015-12-15 12:02:53 -08:00
Eric Anholt	02bcb443ee	vc4: Warn instead of abort()ing on exec ioctl failures. It's really harsh to abort() the X Server because of a momentary failure (particularly -ENOMEM). I don't see a way to pass an -ENOMEM up the stack from here, but we can at least log to stderr before proceeding on. Cc: "11.1" <mesa-stable@lists.freedesktop.org>	2015-12-15 12:02:44 -08:00
Nicolai Hähnle	c8d9d289ff	radeonsi: fix perfcounter selection for SI_PC_MULTI_BLOCK layouts The incorrectly computed register count caused lockups. Reviewed-by: Edward O'Callaghan <eocallaghan@alterapraxis.com>	2015-12-15 11:23:40 -05:00
Nicolai Hähnle	149d049676	gallium/radeon: remove unnecessary test in r600_pc_query_add_result This test is a left-over of the initial development. It is unneeded and misleading, so let's get rid of it. Reviewed-by: Edward O'Callaghan <eocallaghan@alterapraxis.com>	2015-12-15 11:23:40 -05:00
Rob Clark	e677b3047b	freedreno/a4xx: fix fragcoord.z + fragdepth It seems like disabling earlyz on a4xx also, by defaults, disables fragcoord.z to the FS. For frag shaders that both read fragcoord(.z) and write fragdepth, we need to set some extra bits to prevent a lockup. This lets us get rid of the hack of disabling fragcoord.z (which prevented 0ad from lockups, but resulted in rendering corruption). Also fixes fbo-depth-sample-compare. Signed-off-by: Rob Clark <robclark@freedesktop.org>	2015-12-15 09:40:54 -05:00
Rob Clark	cad0920d11	freedreno: update generated headers Signed-off-by: Rob Clark <robclark@freedesktop.org>	2015-12-15 09:39:10 -05:00
Rob Clark	249b2be3bc	freedreno/ir3/cmdline: don't dump nir by default By default we only want the disasm dumped, which we get anyways. Signed-off-by: Rob Clark <robclark@freedesktop.org>	2015-12-15 09:39:10 -05:00
Christian König	10b7a7c344	st/va: remove nonesense HEVC picture id handling The picture id in this case is a VA-API surface handle, checking for a certain value can't be correct. Signed-off-by: Christian König <christian.koenig@amd.com>	2015-12-15 11:25:02 +01:00
Roland Scheidegger	8e264765a4	draw: remove clip_vertex from vertex header vertex header had both clip_pos and clip_vertex. We only really need one (clip_pos) because the draw llvm shader would overwrite the position output from the vs with the viewport transformed. However, we don't really need the second one, which was only really used for gl_ClipVertex - if the shader didn't have that the values were just duplicated to both clip_pos and clip_vertex. So, just use this from the vs output instead when we actually need it. Also change clip debug to output both the data from clip_pos and the clipVertex output (if available). Makes some things more complex, some things less complex, but seems more easy to understand what clipping actually does (and what values it uses to do its magic). Reviewed-by: Brian Paul <brianp@vmware.com Reviewed-by: Jose Fonseca <jfonseca@vmware.com>	2015-12-15 02:03:40 +01:00
Roland Scheidegger	1775400a20	draw: use clip_pos, not clip_vertex for the fake guardband xy point clipping Seems obvious now this should use the data from position and not clip_vertex (albeit might not really make a difference). Reviewed-by: Brian Paul <brianp@vmware.com Reviewed-by: Jose Fonseca <jfonseca@vmware.com>	2015-12-15 02:03:40 +01:00
Roland Scheidegger	8575ddb644	draw: rename vertex header members clip -> clip_vertex and pre_clip_pos -> clip_pos. Looks more obvious to me what these values actually represent (so use something resembling the vs output names). Reviewed-by: Brian Paul <brianp@vmware.com Reviewed-by: Jose Fonseca <jfonseca@vmware.com>	2015-12-15 02:03:40 +01:00
Roland Scheidegger	1b22815af6	draw: don't pretend have_clipdist is per-vertex This is just for code cleanup, conceptually the have_clipdist really isn't per-vertex state, so don't put it there (just dependent on the shader). Even though there wasn't really any overhead associated with this, we shouldn't store random shader information in the vertex header. Reviewed-by: Brian Paul <brianp@vmware.com Reviewed-by: Jose Fonseca <jfonseca@vmware.com>	2015-12-15 02:03:40 +01:00
Roland Scheidegger	9e3f2af3c3	draw: use position not clipVertex output for xyz view volume clipping I'm pretty sure this should use position (i.e. pre_clip_pos) and not the output from clipVertex. Albeit piglit doesn't care. It is what we use in the clip test, and it is what every other driver does (as they don't even have clipVertex output and lower the additional planes to clip distances). Reviewed-by: Brian Paul <brianp@vmware.com Reviewed-by: Jose Fonseca <jfonseca@vmware.com>	2015-12-15 02:03:40 +01:00
Samuel Pitoiset	71135e275f	nvc0: check return value of nvc0_program_validate() Spotted by Coverity. Signed-off-by: Samuel Pitoiset <samuel.pitoiset@gmail.com> Reviewed-by: Ilia Mirkin <imirkin@alum.mit.edu>	2015-12-14 19:08:42 +01:00
Samuel Pitoiset	54f58210c2	nv50: check return value of nouveau_object_new() When ret == 0, obj is not NULL. Spotted by Coverity. Signed-off-by: Samuel Pitoiset <samuel.pitoiset@gmail.com> Reviewed-by: Ilia Mirkin <imirkin@alum.mit.edu>	2015-12-14 19:08:39 +01:00
Samuel Pitoiset	3f7462b792	nv50,nvc0: make use of unreachable() when invalid texture target happens Spotted by Coverity. Signed-off-by: Samuel Pitoiset <samuel.pitoiset@gmail.com> Reviewed-by: Ilia Mirkin <imirkin@alum.mit.edu>	2015-12-14 19:08:25 +01:00
Christian König	8b52fa71ac	st/va: handle default post process regions Avoid referencing NULL pointers. Signed-off-by: Christian König <christian.koenig@amd.com> Reviewed-by: Julien Isorce <j.isorce@samsung.com> Tested-by: Julien Isorce <j.isorce@samsung.com>	2015-12-14 11:54:55 +01:00
Christian König	f6dd31c1cf	st/va: fix unused variable warning Signed-off-by: Christian König <christian.koenig@amd.com> Reviewed-by: Julien Isorce <j.isorce@samsung.com>	2015-12-14 11:54:55 +01:00
Christian König	025d97381e	st/va: clean up post process includes Signed-off-by: Christian König <christian.koenig@amd.com> Reviewed-by: Julien Isorce <j.isorce@samsung.com> Tested-by: Julien Isorce <j.isorce@samsung.com>	2015-12-14 11:54:54 +01:00
Christian König	27a276f625	st/va: cleanup filter color standard handling Signed-off-by: Christian König <christian.koenig@amd.com> Reviewed-by: Julien Isorce <j.isorce@samsung.com> Tested-by: ulien Isorce <j.isorce@samsung.com>	2015-12-14 11:54:54 +01:00
Ilia Mirkin	7752bbc44e	gk104/ir: simplify and fool-proof texbar algorithm With the current algorithm, we only look at tex uses. However there's a write-after-write hazard where we might decide to, on some path, not use a texture's output at all, but instead to write a different value to that register. However without the barrier, the texture might complete later and overwrite that value. This fixes Unreal Elemental demo on GK110/GK208, flightgear on GK10x, and likely other random-looking failures. Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu> Cc: "11.1" <mesa-stable@lists.freedesktop.org>	2015-12-12 18:10:16 -05:00
Ilia Mirkin	d35695096d	nv50/ir: combine sequences of conversions In some cases shaders want non-default rounding when converting float to integer. This can be done in one go, so merge the two ops. This comes up in the packUnorm4x8 & co functions, as well as a few random shaders. Overall shader-db impact is minimal, helping a handful of witcher2 and other misc shaders. Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu>	2015-12-12 18:10:16 -05:00
Ilia Mirkin	dbca0f3eba	nv50/ir: manually optimize multiplication expansion logic The conversion of 32-bit integer multiplies into 16-bit ones happens after the regular optimization loop. However it's fairly common to multiply by a small integer, rendering some of the expansion pointless. Firstly, propagate immediates when possible into mul ops, secondly just remove the ops when they are unnecessary. Including the change to generate imad immediates, the effect is: total instructions in shared programs : 6365463 -> 6351898 (-0.21%) total gprs used in shared programs : 728684 -> 728684 (0.00%) total local used in shared programs : 9904 -> 9904 (0.00%) total bytes used in shared programs : 44001576 -> 44036120 (0.08%) local gpr inst bytes helped 0 0 3288 4 hurt 0 0 0 842 It's easy for this to hurt bytes since we end up always generating the 8-byte form, while we can't always get rid of the immediate in question. Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu>	2015-12-12 18:10:16 -05:00
Ilia Mirkin	3af83c4bc7	nv50/ir: fix imul emission in the presence of an immediate Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu>	2015-12-12 18:10:15 -05:00
Ilia Mirkin	a0b5d5beed	nv50/ir: teach post-ra immediate folding into mad about integers There will usually be a split before the mad op, peer through that and pick out the right word of the immediate. Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu>	2015-12-12 18:10:15 -05:00
Ilia Mirkin	ab70ea1353	nv50/ir: add short imad support Support emission of the short imad, but also include it in the various logic that tries to make it possible to emit. Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu>	2015-12-12 18:10:15 -05:00
Ilia Mirkin	6aca7fecb7	nv50/ir: can't have predication and immediates Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu> Cc: "11.0 11.1" <mesa-stable@lists.freedesktop.org>	2015-12-12 18:10:15 -05:00
Ilia Mirkin	69e8b476d0	nv50/ir: fix texture grad for cubemaps We were ignoring the partial derivatives on the last dim. Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu>	2015-12-12 18:10:15 -05:00
Ilia Mirkin	a27548400e	nv50/ir: fix assumption that prog->maxGPR is in 32-bit reg units On NV50, we use 16-bit reg units (to make it all work with half-regs). A few places assumed that it was always in 32-bit units. Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu>	2015-12-12 18:10:15 -05:00
Nicolai Hähnle	d640f179d3	gallium/ddebug: regularly log the total number of draw calls This helps in the use of GALLIUM_DDEBUG_SKIP: first run a target application with skip set to a very large number and note how many draw calls happen before the bug. Then re-run, skipping the corresponding number of calls. Despite the additional run, this can still be much faster than not skipping anything. Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2015-12-12 15:23:50 -05:00
Nicolai Hähnle	b86d5ccae2	gallium/ddebug: add GALLIUM_DDEBUG_SKIP option When we know that hangs occur only very late in a reproducible run (e.g. apitrace), we can save a lot of debugging time by skipping the flush and hang detection for earlier draw calls. Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2015-12-12 15:23:34 -05:00
Roland Scheidegger	af7ba989fb	llvmpipe: fix layer/vp input into fs when not written by prior stages ARB_fragment_layer_viewport requires that if a fs reads layer or viewport index but it wasn't output by gs (or vs with other extensions), then it reads 0. This never worked for llvmpipe, and is surprisingly non-trivial to fix. The problem is the mechanism to handle non-existing outputs in draw is rather crude, it will simply redirect them to whatever is at output 0, thus later stages will just get garbage. So, rather than trying to fix this up (which looks non-trivial), fix this up in llvmpipe setup by detecting this case there and output a fixed zero directly. While here, also optimize the hw vertex layout a bit - previously if the gs outputted layer (or vp) and the fs read those inputs, we'd add them twice to the vertex layout, which is unnecessary. And do some minor cleanup, slots don't require that many bits, there was some bogus (but harmless) float/int mixup for psize slot too, make the slots all unsigned (we always put pos at pos zero thus everything else has to be positive if it exists), and make sure they are properly initialized (layer and vp index slot were not which looked fishy as they might not have got set back to zero when changing from a gs which outputs them to one which does not). This fixes the failures in piglit's arb_fragment_layer_viewport group (3 each for layer and vp). Reviewed-by: Jose Fonseca <jfonseca@vmware.com>	2015-12-12 01:59:15 +01:00
Brian Paul	27d5be0b8f	svga: avoid emitting redundant SetSamplers() commands This greatly reduces the number of SetSamplers() commands for some applications. Reviewed-by: José Fonseca <jfonseca@vmware.com> Reviewed-by: Charmaine Lee <charmainel@vmware.com>	2015-12-11 16:54:58 -07:00
Brian Paul	1291e910d5	svga: avoid emitting redundant SetIndexBuffer commands Reviewed-by: Charmaine Lee <charmainel@vmware.com> Reviewed-by: José Fonseca <jfonseca@vmware.com>	2015-12-11 16:54:44 -07:00
Brian Paul	c877f1aeef	util/blitter: minor formatting fixes	2015-12-11 16:53:20 -07:00
Eric Anholt	076551116e	vc4: Add quick algebraic optimization for clamping of unpacked values. GL likes to saturate your incoming color, but if that color's coming from unpacking from unorms, there's no point. Ideally we'd have a range propagation pass that cleans these up in NIR, but that doesn't seem to be going to land soon. It seems like we could do a one-off optimization in nir_opt_algebraic, except that doesn't want to operate on expressions involving unpack_unorm_4x8, since it's sized. total instructions in shared programs: 87879 -> 87761 (-0.13%) instructions in affected programs: 6044 -> 5926 (-1.95%) total estimated cycles in shared programs: 349457 -> 349252 (-0.06%) estimated cycles in affected programs: 6172 -> 5967 (-3.32%) No SSPD on openarena (which had the biggest gains, in its VS/CSes), n=15.	2015-12-11 12:36:16 -08:00
Eric Anholt	e3efc4b023	vc4: When doing algebraic optimization into a MOV, use the right MOV. If there were src unpacks, changing to the integer MOV instead of float (for example) would change the unpack operation.	2015-12-11 12:21:22 -08:00
Eric Anholt	2591beef89	vc4: Fix handling of src packs on in qir_follow_movs(). The caller isn't going to expect it from a return, so it would probably get misinterpreted. If the caller had an unpack in its reg, that's fine, but don't lose track of it.	2015-12-11 12:21:22 -08:00
Eric Anholt	b70a2f4d81	vc4: Add missing progress note in opt_algebraic.	2015-12-11 12:21:22 -08:00
Eric Anholt	5989ef2b0f	vc4: Add debugging of the estimated time to run the shader to shader-db.	2015-12-11 12:21:22 -08:00
Eric Anholt	53b2523c6e	vc4: Fix handling of sample_mask output. I apparently broke this in a late refactor, in such a way that I decided its tests were some of those interminable ones that I should just blacklist from my testing. As a result, the refactors related to it were totally wrong.	2015-12-11 12:21:22 -08:00
Edward O'Callaghan	53609de762	softpipe: enable GL_ARB_viewport_array support, update GL3.txt doc Signed-off-by: Edward O'Callaghan <eocallaghan@alterapraxis.com> Reviewed-by: Roland Scheidegger <sroland@vmware.com>	2015-12-11 20:09:21 +01:00

1 2 3 4 5 ...

25573 commits