fdo-mirrors/mesa

mirror of https://gitlab.freedesktop.org/mesa/mesa.git synced 2026-04-23 09:30:36 +02:00

Author	SHA1	Message	Date
Francisco Jerez	dfe957c02b	i965: Move up fs_inst::flag_subreg to backend_instruction. Reviewed-by: Matt Turner <mattst88@gmail.com>	2015-02-10 16:05:51 +02:00
Francisco Jerez	639696aa05	i965: Move up fs_inst::regs_written to backend_instruction. It will also be useful in the VEC4 back-end. Reviewed-by: Matt Turner <mattst88@gmail.com>	2015-02-10 16:05:51 +02:00
Francisco Jerez	4ed52e8bc4	i965/vec4: Remove dependency of vec4_instruction on the visitor class. The only reason why you need a vec4_visitor to construct a vec4_instruction is to initialize vec4_instruction::ir and ::annotation. Instead set them from vec4_visitor::emit() just like fs_visitor does. Reviewed-by: Matt Turner <mattst88@gmail.com>	2015-02-10 16:05:50 +02:00
Francisco Jerez	a3ee6c7d19	i965/fs: Remove dependency of fs_inst on the visitor class. The fs_visitor argument of fs_inst::regs_read() wasn't used at all. Reviewed-by: Matt Turner <mattst88@gmail.com>	2015-02-10 16:05:50 +02:00
Francisco Jerez	bfbb0e84e1	i965: Move IR object definitions to separate header files. One should be able to manipulate i965 IR without pulling the whole FS/VEC4 visitor classes -- Optimization passes and other transformations would ideally be visitor-agnostic. Among other issues this avoids a circular dependency between the header file where such visitor-agnostic code will be defined and the main FS/VEC4 header where both IR (layer below) and visitor (layer above) happen to be defined. Reviewed-by: Matt Turner <mattst88@gmail.com>	2015-02-10 16:05:50 +02:00
Francisco Jerez	447879eb88	i965: Factor out virtual GRF allocation to a separate object. Right now virtual GRF book-keeping and allocation is performed in each visitor class separately (among other hundred different things), leading to duplicated logic in each visitor and preventing layering as it forces any code that manipulates i965 IR and needs to allocate virtual registers to depend on the specific visitor that happens to be used to translate from GLSL IR. v2: Use realloc()/free() to allocate VGRF book-keeping arrays (Connor). Reviewed-by: Matt Turner <mattst88@gmail.com>	2015-02-10 16:05:47 +02:00
Francisco Jerez	e6146e6f14	glsl: Forbid calling the constructor of any opaque type. The spec doesn't define any opaque type constructors. Reviewed-by: Ian Romanick <ian.d.romanick@intel.com>	2015-02-10 15:49:43 +02:00
Francisco Jerez	c4111dfa0a	glsl: Return correct number of coordinate components for cubemap array images. Cubemap array images are unlike cubemap array samplers in that they don't need an additional coordinate to index individual cubemaps in the array, instead they behave like a 2D array of 6n layers, with n the number of cubemaps in the array. Take this exception into account. Reviewed-by: Ian Romanick <ian.d.romanick@intel.com>	2015-02-10 15:49:43 +02:00
Francisco Jerez	fcc2fd53df	mesa: Bump MAX_IMAGE_UNIFORMS to 32. So the i965 driver can expose 32 image uniforms per shader stage. Reviewed-by: Kenneth Graunke <kenneth@whitecape.org>	2015-02-10 15:37:56 +02:00
Francisco Jerez	818585b9f9	mesa: Rename the CEILING() macro to DIV_ROUND_UP(). Some people have complained that code using the CEILING() macro is difficult to understand because it's not immediately obvious what it is supposed to do until you go and look up its definition. Use a more descriptive name that matches the similar utility macro in the Linux kernel. Reviewed-by: Matt Turner <mattst88@gmail.com>	2015-02-10 15:37:47 +02:00
Tiziano Bacocco	1e02f2badf	nv50,nvc0: Mark PIPE_QUERY_TIMESTAMP_DISJOINT as ready immediately Without this when an application issues that query, it would try to wait the result from the gpu, and since no query has been actually issued, it will wait forever. Signed-off-by: Tiziano Bacocco <tizbac2@gmail.com> Reviewed-by: Ilia Mirkin <imirkin@alum.mit.edu>	2015-02-10 08:02:17 -05:00
Roy Spliet	09ee907266	nv50/ir: Fold IMM into MAD Add a specific optimisation pass for NV50 to check whether SRC0 or SRC1 is a MOV dst, IMM. If so: fold the IMM in and try to drop the MOV. Must be done post-RA because it requires that SDST == SSRC2. V2: improve readability and add comments to clarify decisions V3: Remove redundant code... compiler already attempts to put the IMM in SSRC1 Signed-off-by: Roy Spliet <rspliet@eclipso.eu> Reviewed-by: Ilia Mirkin <imirkin@alum.mit.edu>	2015-02-10 08:02:13 -05:00
Roy Spliet	3dc39d0bca	nv50/ir: Add emit support for MAD IMM format But don't enable generation of it in the opProperties, because we can't guarantee the SDST==SRC2 constraint until after register assignment. We'll add a post-RA folding pass to utilise this. Signed-off-by: Roy Spliet <rspliet@eclipso.eu> Reviewed-by: Ilia Mirkin <imirkin@alum.mit.edu>	2015-02-10 08:02:02 -05:00
Roy Spliet	fb63df2215	nv50/ir: Add support for MAD 4-byte opcode Add emission rules for negative and saturate flags for MAD 4-byte opcodes, and get rid of some of the constraints. Obviously tested with a wide variety of shaders. V2: Document MAD as supported short form V3: Split up IMM from short-form modifiers Signed-off-by: Roy Spliet <rspliet@eclipso.eu> Reviewed-by: Ilia Mirkin <imirkin@alum.mit.edu>	2015-02-10 08:01:46 -05:00
Ilia Mirkin	354206f407	nv50/ir: change the way float face is returned The old way made it impossible for the optimizer to reason about what was going on. The new way is the same number of instructions (the neg gets folded into the cvt) but enables the optimizer to be cleverer if comparing to a constant (most common case). [The optimizer is presently not sufficiently clever to work this out, but it could relatively easily be made to be. The old way would have required significant complexity to work out.] Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu>	2015-02-10 08:01:46 -05:00
Kenneth Graunke	480ee1f0b4	nir: Mark nir_print_instr's instr pointer as const. Printing instructions doesn't modify them, so we can mark the parameter const. Signed-off-by: Kenneth Graunke <kenneth@whitecape.org> Reviewed-by: Jason Ekstrand <jason.ekstrand@intel.com>	2015-02-10 03:37:55 -08:00
Kenneth Graunke	08a06b6b89	i965: Fix integer border color on Haswell. +82 Piglits - 100% of border color tests now pass on Haswell. Signed-off-by: Kenneth Graunke <kenneth@whitecape.org> Reviewed-by: Jason Ekstrand <jason.ekstrand@intel.com> Reviewed-by: Chris Forbes <chrisf@ijw.co.nz> Cc: mesa-stable@lists.freedesktop.org	2015-02-09 13:18:58 -08:00
Kenneth Graunke	e1e73443c5	i965: Use a gl_color_union for sampler border color. This should have no effect, but will make it easier to implement other bug fixes. v2: Eliminate "unsigned one" local; just use the value where necessary. Signed-off-by: Kenneth Graunke <kenneth@whitecape.org> Reviewed-by: Ian Romanick <ian.d.romanick@intel.com> Cc: mesa-stable@lists.freedesktop.org	2015-02-09 13:18:58 -08:00
Kenneth Graunke	8cb18760cc	i965: Override swizzles for integer luminance formats. The hardware's integer luminance formats are completely unusable; currently we fall back to RGBA. This means we need to override the texture swizzle to obtain the XXX1 values expected for luminance formats. Fixes spec/EXT_texture_integer/texwrap formats bordercolor [swizzled] on Broadwell - 100% of border color tests now pass on Broadwell. Signed-off-by: Kenneth Graunke <kenneth@whitecape.org> Reviewed-by: Ian Romanick <ian.d.romanick@intel.com> Reviewed-by: Chris Forbes <chrisf@ijw.co.nz> Cc: mesa-stable@lists.freedesktop.org	2015-02-09 13:18:54 -08:00
Carl Worth	b16de0b713	util/u_atomic: Add new macro p_atomic_add This provides for atomic addition, which will be used by an upcoming shader-cache patch. A simple test is added to "make check" as well. Note: The various O/S functions differ on whether they return the original value or the value after the addition, so I did not provide an add_return() macro which would be sensitive to that difference. Reviewed-by: Matt Turner <mattst88@gmail.com> Reviewed-by: Aaron Watry <awatry@gmail.com> Reviewed-by: José Fonseca <jfonseca@vmware.com>	2015-02-09 10:47:44 -08:00
Jason Ekstrand	345e8cc849	util/hash_table: Try to hit a double-insertion bug in the collision test Reviewed-by: Eric Anholt <eric@anholt.net>	2015-02-07 17:01:05 -08:00
Jason Ekstrand	623c3a858d	util/set: Do a full search when adding new items Previously, the set_insert function would bail early if it found a deleted slot that it could re-use. However, this is a problem if the key being inserted is already in the set but further down the list. If this happens, the element ends up getting inserted in the set twice. This commit makes it so that we walk over all of the possible entries for the given key and then, if we don't find the key, place it in the available free entry we found. Reviewed-by: Eric Anholt <eric@anholt.net>	2015-02-07 17:01:05 -08:00
Jason Ekstrand	c9287e797b	util/hash_table: Do a full search when adding new items Previously, the hash_table_insert function would bail early if it found a deleted slot that it could re-use. However, this is a problem if the key being inserted is already in the hash table but further down the list. If this happens, the element ends up getting inserted in the hash table twice. This commit makes it so that we walk over all of the possible entries for the given key and then, if we don't find the key, place it in the available free entry we found. Reviewed-by: Eric Anholt <eric@anholt.net>	2015-02-07 17:01:05 -08:00
James Legg	1581e12aba	mesa: Make renderbuffer FBO attachments not layered For framebuffer completeness checks, consider renderbuffers as not layered. Previously, they would have counted as layered if a layered textured had previously been bound to the same attachment point. This could cause framebuffer completeness checks to incorrectly fail with GL_FRAMEBUFFER_INCOMPLETE_LAYER_TARGETS, even if no layered attachments were present. Reviewed-by: Chris Forbes <chrisf@ijw.co.nz> Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=89026	2015-02-08 13:54:15 +13:00
Brian Paul	d1e21325cf	gallium/hud: also try R8_UNORM format for font texture Convert the code to try formats from an array rather than a bunch of if/else cases. Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2015-02-07 11:03:37 -07:00
Brian Paul	6447e9dbfa	gallium/hud: flush stdout in print_help(), for Windows Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2015-02-07 11:03:37 -07:00
Ben Widawsky	7ea1e37497	i965: Add more stringent blitter assertions Blits to or from a y-tiled surface must always be a multiple of the tile size. From page 16 of the HSW PRM (https://01.org/linuxgraphics/sites/default/files/documentation/intel-gfx-prm-osrc-hsw-memory-views.pdf#16) "The pitch of a tiled enclosing region must be an integral number of tile widths" Signed-off-by: Ben Widawsky <ben@bwidawsk.net> Reviewed-by: Kenneth Graunke <kenneth@whitecape.org> Reviewed-by: Jason Ekstrand <jason.ekstrand@intel.com>	2015-02-07 08:08:59 -08:00
Ben Widawsky	efde74c89d	i965: Consolidate some of the intel_blit logic An upcoming patch is going to introduce some code here, and having this code organized as the patch does makes it a bit easier to read later. There should be no functional change here. Signed-off-by: Ben Widawsky <ben@bwidawsk.net> Reviewed-by: Kenneth Graunke <kenneth@whitecape.org> Reviewed-by: Jason Ekstrand <jason.ekstrand@intel.com>	2015-02-07 08:07:56 -08:00
Park, Jeongmin	0467a52dc3	st/dri: Make depth buffer optional for postprocessing Since only pp_jimenezmlaa uses depth buffer, we can make it optional. Signed-off-by: Marek Olšák <marek.olsak@amd.com>	2015-02-07 12:12:00 +01:00
Park, Jeongmin	2e6ba6afdb	postprocess: Check for depth buffer in pp_jimenezmlaa Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=88962 Signed-off-by: Marek Olšák <marek.olsak@amd.com>	2015-02-07 12:12:00 +01:00
Ben Widawsky	8030e269e9	i965/vec4: Correct MUL destination hazard As it turns out, we were over-thinking the cause of the hang on Cherryview. It's simply errata for Cherryview. commit `88fea85f09` Author: Ben Widawsky <benjamin.widawsky@intel.com> Date: Fri Nov 21 10:47:41 2014 -0800 i965/vec4/gen8: Handle the MUL dest hazard exception This is an explanation to why we never saw the hang on BDW. NOTE: The problem the original patch was trying to fix does still exist. It will have to be fixed at some point. v2: Modify commit message, s/CHV/BDW Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=84212 Signed-off-by: Ben Widawsky <ben@bwidawsk.net> Reviewed-by: Matt Turner <mattst88@gmail.com>	2015-02-06 17:54:17 -08:00
Eric Anholt	bff4cbdafa	nir: Fix broken fsat recognizer. We've probably never seen this ridiculous pattern in the wild, so it didn't matter. Reviewed-by: Connor Abbott <cwabbott0@gmail.com> Reviewed-by: Jason Ekstrand <jason.ekstrand@intel.com>	2015-02-06 15:57:55 -08:00
Eric Anholt	6706537dd4	nir: Slightly simplify algebraic code generation by reusing a struct. Reviewed-by: Jason Ekstrand <jason.ekstrand@intel.com>	2015-02-06 15:57:55 -08:00
Eric Anholt	9e35af08af	tgsi/ureg: Add missing some missing opcodes opcode_tmp.h I wanted all of these for NIR-to-TGSI. Reviewed-by: Roland Scheidegger <sroland@vmware.com>	2015-02-06 15:50:07 -08:00
Eric Anholt	f3dbf3689a	tgsi/ureg: Move ureg_dst_register() to the header. I wanted to use it for nir-to-tgsi. The equivalent ureg_src_register() is also located here. Reviewed-by: Roland Scheidegger <sroland@vmware.com>	2015-02-06 15:50:07 -08:00
Marek Olšák	40fa7d44ab	gallium/u_tests: test a NULL buffer sampler view Reviewed-by: Glenn Kennard <glenn.kennard@gmail.com>	2015-02-06 22:27:07 +01:00
Marek Olšák	56e709bffb	gallium/u_tests: test a NULL constant buffer This expects (0,0,0,0), though it can be changed to something else or allow more than one set of values to be considered correct. This is currently the radeonsi behavior. Reviewed-by: Glenn Kennard <glenn.kennard@gmail.com>	2015-02-06 22:27:07 +01:00
Marek Olšák	9e8a6d8486	gallium/u_tests: test a NULL texture sampler view v2: allow one of the two values	2015-02-06 22:27:06 +01:00
Marek Olšák	63e51baedc	gallium/u_tests: restructure the only test, refactor out reusable code Reviewed-by: Glenn Kennard <glenn.kennard@gmail.com>	2015-02-06 22:27:06 +01:00
Marek Olšák	dcf996c31e	gallium: run gallium tests if GALLIUM_TESTS=1 is set Reviewed-by: Glenn Kennard <glenn.kennard@gmail.com>	2015-02-06 22:27:06 +01:00
Marek Olšák	0271ac72d1	gallium/postprocessing: fix crash at context destruction Reviewed-by: Michel Dänzer <michel.daenzer@amd.com>	2015-02-06 20:03:06 +01:00
Xavier Bouchoux	2fd21c4098	r600g/sb: fix a bug in constants folding optimisation pass ADD R6.y.1, R5.w.1, ~1\|3f800000 ADD R6.y.2, \|R6.y.1\|, -0.0001\|b8d1b717 was wrongly being converted to ADD R6.y.1, R5.w.1, ~1\|3f800000 ADD R6.y.2, R5.w.1, -1.0001\|bf800347 because abs() modifier was ignored. Signed-off-by: Xavier Bouchoux <xavierb@gmail.com> Reviewed-by: Glenn Kennard <glenn.kennard@gmail.com>	2015-02-06 20:03:06 +01:00
Xavier Bouchoux	acef65503e	r600g: fix abs() support on ALU 3 source operands instructions Since alu does not support abs() modifier on source operands, spill and apply the modifiers to a temp register when needed. Signed-off-by: Xavier Bouchoux <xavierb@gmail.com> Reviewed-by: Glenn Kennard <glenn.kennard@gmail.com>	2015-02-06 20:03:06 +01:00
David Heidelberg	bae23a1756	r300g: small code cleanup (v2) v2: incorporated changes from Marek Olšák Reviewed-by: Marek Olšák <marek.olsak@amd.com> Signed-off-by: David Heidelberg <david@ixit.cz>	2015-02-06 18:27:30 +01:00
Iago Toral Quiroga	71a36e0a2c	glsl: GLSL ES identifiers cannot exceed 1024 characters v2 (Ian Romanick) - Move the check to the lexer before rallocing a copy of the large string. Fixes the following 2 dEQP tests: dEQP-GLES3.functional.shaders.keywords.invalid_identifiers.max_length_vertex dEQP-GLES3.functional.shaders.keywords.invalid_identifiers.max_length_fragment Reviewed-by: Ian Romanick <ian.d.romanick@intel.com>	2015-02-06 12:21:42 +01:00
Kenneth Graunke	d4a461caaf	i965: Fix INTEL_DEBUG=shader_time for SIMD8 VS (and GS). We were incorrectly attributing VS time to FS8 on Gen8+, which now use fs_visitor for vertex shaders. We don't hit this for geometry shaders yet, but we may as well add support now - the fix is obvious, and we'll just forget later. Signed-off-by: Kenneth Graunke <kenneth@whitecape.org> Reviewed-by: Francisco Jerez <currojerez@riseup.net> Reviewed-by: Jordan Justen <jordan.l.justen@intel.com>	2015-02-05 20:01:03 -08:00
Kenneth Graunke	32f1d4e286	i965/fs: Use inst->eot rather than opcodes in register allocation. Previously, we special cased FB writes and URB writes in the register allocation code. What we really wanted was to handle any message with EOT set. This saves us from extending the list with new opcodes in the future. Signed-off-by: Kenneth Graunke <kenneth@whitecape.org> Reviewed-by: Jordan Justen <jordan.l.justen@intel.com> Reviewed-by: Ben Widawsky <ben@bwidawsk.net>	2015-02-05 20:01:02 -08:00
Kenneth Graunke	10d8a1a88e	i965/fs: Delete is_last_send(); just check inst->eot. This helper function basically just checks inst->eot, but also asserts that only opcodes we expect to terminate threads have EOT set. As far as I'm aware, we've never had such a bug. Removing it means that we don't have to extend the list for new opcodes. Cherryview and Skylake introduce an optimization where sampler messages can have EOT set; scalar GS/HS/DS will likely introduce new opcodes as well. Signed-off-by: Kenneth Graunke <kenneth@whitecape.org> Reviewed-by: Jordan Justen <jordan.l.justen@intel.com> Reviewed-by: Ben Widawsky <ben@bwidawsk.net>	2015-02-05 20:00:42 -08:00
Michel Dänzer	a338dc0186	st/mesa: Don't use PIPE_USAGE_STREAM for GL_PIXEL_UNPACK_BUFFER_ARB The latter currently implies CPU read access, so only PIPE_USAGE_STAGING can be expected to be fast. Mesa demos src/tests/streaming_rect on Kaveri (radeonsi): Unpatched: 42 frames in 1.023 seconds = 41.056 FPS Patched: 615 frames in 1.000 seconds = 615.000 FPS Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=88658 Cc: "10.3 10.4" <mesa-stable@lists.freedestkop.org> Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2015-02-06 10:55:53 +09:00
Tiziano Bacocco	17abefa12b	st/nine: Implement dummy vbo behaviour when vs is missing inputs Use a dummy vertex buffer object when vs inputs have no corresponding entries in the vertex declaration. This dummy buffer will give to the shader float4(0,0,0,0). This fixes several artifacts on some games. Signed-off-by: Axel Davy <axel.davy@ens.fr> Signed-off-by: Tiziano Bacocco <tizbac2@gmail.com>	2015-02-06 00:07:20 +01:00

1 2 3 4 5 ...

61503 commits