Mesa 26.2.0 Release Notes / 2026-08-05¶
Mesa 26.2.0 is a new development release. People who are concerned with stability and reliability should stick with a previous release or wait for Mesa 26.2.1.
Mesa 26.2.0 implements the OpenGL 4.6 API, but the version reported by glGetString(GL_VERSION) or glGetIntegerv(GL_MAJOR_VERSION) / glGetIntegerv(GL_MINOR_VERSION) depends on the particular driver being used. Some drivers don’t support all the features required in OpenGL 4.6. OpenGL 4.6 is only available if requested at context creation. Compatibility contexts may report a lower version depending on each driver.
Mesa 26.2.0 implements the OpenCL 3.1 API, but the version reported by the CL_DEVICE_VERSION, CL_DEVICE_NUMERIC_VERSION and CL_DEVICE_OPENCL_C_ALL_VERSIONS clGetDeviceInfo queries depends on the particular driver being used.
Mesa 26.2.0 implements the Vulkan 1.4 API, but the version reported by the apiVersion property of the VkPhysicalDeviceProperties struct depends on the particular driver being used.
SHA checksums¶
SHA256: efd4bb08cdb7c365a812cd4e6c9202ab55b2f22cdcd13c7d6c4f9647b799a4ef mesa-26.2.0.tar.xz
SHA512: c7c810c6958fa18eb75eea9968d84d0edd29b579d351a7f22d8c2b2b13b6e04e5b0df31dae3aa943672c792d124325bab03ccbd475a16d27267418bddbec4936 mesa-26.2.0.tar.xz
New features¶
cl_khr_subgroup_rotate on radeonsi
cl_khr_subgroup_rotate on iris
VK_EXT_shader_uniform_buffer_unsized_array on panvk
VK_KHR_shader_constant_data on RADV
VK_EXT_dynamic_rendering_unused_attachments on panvk
protectedMemory support on RADV/GFX10+ and VEGA10
VK_KHR_performance_query on RADV/GFX11
VK_EXT_conservative_rasterization on panvk
shaderImageGatherExtended on pvr
static C++ stdlib required on rusticl to workaround applications using their own C++ stdlib
VK_EXT_pipeline_protected_access on RADV
VK_EXT_extended_dynamic_state3 on panvk
GL_ARB_texture_query_lod on panfrost/v9+
VK_KHR_maintenance11 on RADV
OpenCL 3.1 support for rusticl on asahi, iris, radeonsi, llvmpipe and zink
VK_KHR_workgroup_memory_explicit_layout on pvr
VK_KHR_maintenance5 on pvr
VK_KHR_calibrated_timestamps on hasvk
VK_KHR_present_id on pvr
VK_KHR_present_wait on pvr
VK_KHR_present_id2 on hasvk
VK_KHR_present_wait2 on hasvk
VK_KHR_shader_fma on RADV
VK_EXT_shader_split_barrier on RADV/GFX12
VK_KHR_shader_fma on nvk
Support for G1-Ultra, G1-Premium and G1-Pro GPUs on Panfrost and PanVK
VK_EXT_shader_atomic_float on nvk
VK_{KHR,EXT}_index_type_uint8 on pvr
VK_EXT_debug_marker in vulkan runtime
VK_EXT_mesh_shader on NVK
VK_KHR_device_fault on RADV
VK_EXT_present_timing now also on wsi/x11
VK_EXT_present_timing on hasvk
VK_GOOGLE_display_timing for KHR_display (and opt-in for x11, wayland) on anv, hasvk, hk, nvk, panvk, pvr, radv, tu, v3dv
VK_KHR_shader_abort on RADV
VK_KHR_shader_fma on panvk
VK_NV_shader_atomic_float16_vector on NVK
VK_KHR_compute_shader_derivatives on panvk
VK_EXT_shader_subgroup_ballot on pvr
VK_EXT_shader_subgroup_vote on pvr
VK_KHR_shader_subgroup_rotate on pvr
VK_KHR_shader_subgroup_uniform_control_flow on pvr
VK_EXT_subgroup_size_control on pvr
VK_EXT_rasterization_order_attachment_access on panvk
VK_ARM_rasterization_order_attachment_access on panvk
GL_OES_texture_float on etnaviv/HALF_FLOAT
GL_ARB_texture_float on etnaviv/HALF_FLOAT
VK_NVX_binary_import on NVK
VK_EXT_image_sliced_view_of_3d on panvk
VK_EXT_shader_tile_image on panvk
VK_KHR_internally_synchronized_queues on panvk
VK_KHR_workgroup_memory_explicit_layout on panvk
VK_EXT_device_memory_report on pvr
VK_EXT_descriptor_heap enabled by default on anv, RADV
VK_EXT_device_address_binding_report on panvk
VK_EXT_multisampled_render_to_single_sampled on RADV/Android
VK_KHR_shader_fma on anv
VK_KHR_extended_flags on RADV
VK_EXT_display_surface_counter on pvr
VK_EXT_display_control on pvr
VK_EXT_direct_mode_display on pvr
VK_KHR_surface_maintenance1 on pvr
VK_EXT_surface_maintenance1 on pvr
VK_KHR_swapchain_maintenance1 on pvr
VK_EXT_swapchain_maintenance1 on pvr
VK_EXT_swapchain_colorspace on pvr
VK_EXT_acquire_drm_display on pvr
VK_KHR_unified_image_layouts on pvr
VK_KHR_shader_quad_control on v3dv
VK_KHR_shader_subgroup_rotate on v3dv
VK_KHR_shader_maximal_reconvergence on v3dv
VK_EXT_shader_image_atomic_int64 on panvk
VK_EXT_host_image_copy on RADV/GFX10.3+
Bug fixes¶
“bad target in _mesa_select_tex_object()” when using glXBindTexImageEXT
7900XT, hardware acceleration crashes chromium-based apps unless LIBVA_DRIVER_NAME=null
A702: assertions at A6XX_TEX_MEMOBJ_4_BASE_LO
ANV: dEQP ASTC tests crash w/ FPE_INTDIV on Xe3
AV1 videos dropping frames with AMD card
After updating GStreamer, all videos in Showtime are green/purple
Ambient occlusion is broken in The Chronicles of Riddick - Assault on Dark Athena
Blockland crashes when shader quality turned up
Broken rendering in Isonzo (Unity) since vulkan-radeon 26.1.0
Celestia hit fallback on r300g from git?
Confidential issue #15670
Copy paste bug in `gallium/drivers/radeonsi/si_texture.c`
Copy paste bug in `vulkan/runtime/vk_debug_utils.c`
Crash linking cached GLSL shader with optimized array with explicit uniform location
D3D12: last vertex’s non-zero-offset attribute fetched as 0 with a tightly-packed interleaved vertex buffer (AMD D3D12 only, regression since 24.1)
DaVinci Resolve 20.2 opencl-mesa crash
Dishonored 2 flickering menus on BMG and DG2
Doom Eternal - Page Fault - somewhat reproducable. (RX9070)
Gunfire Reborn crashes with ring gfx_0.0.0 timeout since Mesa 26.0.0
Horizon Forbidden West misrendred lighting effects on BMG
Intel binaries using GLX crash on ARM Macs using llvmpipe
Is maxFragmentCombinedOutputResources=16 in Honeykrisp reflects an actual HW limit?
Issues with shadows on DOOM: The Dark Ages - Revelations DLC
Mesa fails to build due to rust bindings error
Minecraft with Complementary Shaders and Voxy LOD rendering issue on Radeon 680M on latest mesa main branch
NIR: loop unrolling should not create phis for constants defined in the loop
NVK The Surge 2 misrendering near text
NVK: Rendering corruption in Shadow of the Tomb Raider
OpenGL app deadlocks in loader_dri3_swap_buffers_msc → xcb_wait_for_special_event on radeonsi / Mesa 26.1 / X11 (Telegram Desktop)
Qt application TrenchBroom hangs in glXSwapBuffers in recent version of Mesa
RADV NIR optimization bug
RADV/RX 9070 XT: WUCHANG: Fallen Feathers gpu ring reset
RADV: RDNA2 Raytracing regression after “aco/lower_branches: Add try_rotate_latch_block() optimization”
RADV_DEBUG=bo_history crashes any vkd3d/dxvk game at startup with heap corruption (double free)
RADV_PERFTEST=transfer_queue causes assertion failure in prime scenario
Regression. VTK polydata. gl_PrimitiveID requires explicit geometry shader with versions 25.3.6 and later.
Rusticl causes a crash after compiling an OpenCL program
SIGABRT: Invalid free in zink_destroy_resource_surface_cache()
Spike in mmap count for vk_cmd_queue regressed dEQP-VK.api.object_management.max_concurrent#command_buffer_*
System freeze when seeking in h264 files with gst-play-1.0 using the VA plugin
The End is Nigh (Wine): No lighting in The Hollows
Uplink text rendering still bugged out
[AC/NIR] XPlane 12 failure to start
[ANV] Intel arc b580 | Halo Infinite misrenders and GPU hangs
[ANV][ARC] Space engineers 2 Artifacts
[ANV][LNL] - Split Fiction (2001120) - Vertex explosion on white cloth during tutorial.
[ANV][PTL] - Elden Ring (1245620) - Blue artifacts and flickering lights with raytracing enabled
[ANV][PTL] - Horizon: Forbidden West regression in shader quirk
[ANV][PTL] - Persona 3 Reload (2161700) - Blue color on some objects + reflections do not reflect
[Arc B580][Baldurs Gate 3] Hang on opening inventory
[BMG] regression: Artifacts in GTK applications in 26.0.x
[GR Breakpoint] Inverted frustum culling for grass meshes on Linux
[PTL] dEQP-VK.synchronization2.op.single_queue.event.write_fill_buffer_read_ubo_tess_control.buffer_16384_maintenance9 fails
[RADV/ACO]: (Bisected) Regression causes flashing reflections in CyperPunk 2077
[RADV] REGRESSION: Lighting Bleed-through in Need for Speed games on mesa 1:26.1.1-2 ArchLinux
[RADV] REGRESSION: Weird graphics glitches and Lighting Bleed-through in Saints Row 2
[RADV] Regression: Infinite loop in NIR compiler during nir_opt_dead_write_vars loading Gemma 4 via llama.cpp on gfx1100
[RADV] Regression: Sackboy - A big Adventure, crash if RT is enabled
[RADV] Video color artifacts in mpv
[RADV][RDNA3][regression] GPU context lost in Sushi Ben (Steam 2419240) when using in-game camera
[Security] gallium VA-API AV1 tile_info(): unbounded loop writes past fixed tile_col_start_sb/width_in_sbs arrays on heap
[VAAPI] VRAM leak on Polaris, 26.1 regression
[VAAPI][Feature Request] Support of VAProcPipelineCaps.blend_flags in radeonsi (vf_overlay_vaapi)
[VAOn12] Thread safety issues
[VC4/V3D] GL_EXT_shadow_samplers exposed but GL_TEXTURE_COMPARE_MODE/FUNC returns GL_INVALID_ENUM in GLES2
[Windows][arm64] Building windows ARM64 with MSVC fails on src\util\u_math.c
[Xe][ARC770][Baldurs Gate 3] Hang during shader compilation
[amdgpu] Little Inferno rendering issues starting with Mesa 21.1
[anv] Intel ARC B390 | Hitman 2 | DX11 | Blinking corruptions seen
[anv] Intel ARC B390 | Starfield | DX12 | Crash after starting the game
[anv] [bmg] Halo Infinite will not start rendering on an Arc B580
[anv] dpas/Vulkan coopmat uses ~double registers due to scalar lowering
[anv] negate of INT_MAX calculated as INT_MIN instead of INT_MIN + 1
[anv] wrong value for local variable after nested switch merge
[meson][etnaviv] Build target etnaviv_isa_rs has no sources
[r300] [big-endian] bad colors under memory pressure
[radeonsi/VCN] RX 9070 XT (Navi 48, gfx1201): VCN unified ring timeout during VAAPI HEVC encode from Steam Game Recording
[radeonsi] SIGSEGV in gallivm during software fallback for GL_FEEDBACK with lighting
[radeonsi] eglinfo ‘libEGL warning: failed to get driver name for fd -1’
[radv] Regression causes GPU page faults in Crimson Desert
[radv] Regression causes black patches on the ground in DOOM The Dark Ages Revelations DLC
android14 gpu virgl: Shader memory leak
anv: Assert in alias vkCreateImage
anv: Missing null check in vkCmdEndTransformFeedback
anv: NULL deref in anv_AllocateMemory when importing DMA-BUF without VkMemoryDedicatedAllocateInfo on Xe2 (regression)
anv: World of Warcraft dx11 CMAA 2 option buggy on a b580
anv: anv_load_fp64_shader consumes 12MB of heap at device creation
anv: atan shader compiler bug leads to negative infinity
anv: bottom-of-pipe timestamp latched before vkCmdDispatch completes (Arc B390 / Panther Lake, Mesa 26.1.4), breaking wgpu compute-pass timings
anv: cooperative matrix loads from shared memory return the wrong tile on Arc B390 (Xe3/PTL), matmul results are wrong
asahi: broken image? rendering in firefox with active hw-acceleration since mesa-26.0.5
brw/decoder: fragment shader decoding issue
brw: Blender performance regression for Barbershop benchmark scene
build: intel_eu_stall_viewer fails on 32-bit platform
ci: arm64 lava kernel missing zstd support
d3d12: stretched + black-bar artifact after fast offscreen FBO resize on Adreno (regression in MR 41322)
ethosu: Build error on 32-bit due to %lu on uint64_t
gallium/vl: scale_vaapi limited range RGB to limited range YUV broken with red tint
glcpp: incorrect macro expansion in token pasting
intel/blorp: VK_ANDROID_external_format regression on buildtype=debug
intel/blorp: shader_pipeline initialization results in incorrect blorp_key values
intel: Investigate HIZ Plane Optimization disable bit for gfx12.5
ir3: THREAD64→THREAD128 heuristic change in 25.2.0 causes vertical line artifacts in No Man’s Sky on Adreno 740 (a7xx)
ir3: ir3_cf brokenness
iris: Blender wireframe rendering broken
iris: unhandled -EAGAIN in iris_batch_flush (26.1 regression)
kk: Implement timestamps
kk: stencil test unexpectedly not working
lavapipe: dynamic state for advanced blend is broken
llvm23 breaks build of clc_helpers.cpp
llvm23 commit d50631f breaks build of ac_llvm_helper.cpp
macOS build stuck in infinite loop since 26.1
mediafoundation: error C2039: ‘step’: is not a member of ‘_inputQPSettings’
mesa 26.1.3 does not compile with ARM64 MSVC
mesa: clean up st_context caps flags
mesa: glthread crash with MESA_VERBOSE=api
nir: possible exactness bug in reassociate
nir_opt_copy_prop_vars validation failure with Rusticl OpenCL kernel shaders
nvk: various float_controls cts test fails on Turing only
nvk: zcull causes SAVE_RESTORE_ADDR_OOB crashes in Horizon Forbidden West
panvk: Failures in new (1.4.5.0) texturequerylod*trilinear CTS tests
r300 bisected : Broken rendering with applications using MSAA visuals
r300: dEQP-GLES2.functional.shaders.struct.uniform.(not_)equal_fragment regression
r600, sfn: Lowering to assembly failed on R700 while running ShooterGame demo
r600: SFN assersion failed in Transport fever 2
radeonsi+ACO jobs should use ACO_DEBUG=validatera
radeonsi/VCN: AV1/HEVC encode reports a coded size larger than the coded buffer, client segfaults reading the mapped buffer
radeonsi: sqtt missing data for viewperf
radv: Use more efficient cache uuid for hardware identifiers
radv: acceleration structure update (refit) of AABB geometry loses intersection candidates (NAVI32, Mesa 26.1.4)
radv: descriptor heaps do not support non-uniform indexing on buffer pointers
radv: incorrect memory accounting for imported bufffers
radv: use preload sgprs for 16/8bit push constant loads
regression;bisected;radeonsi/video: corrupted VA-API H.264 video encoding (bisected to 4487162a)
rusticl + v3d assertion failure
rusticl/radeonsi OpenCL behavior regression after 5a298f3560629943e9140c74547183da1635352e
rusticl: incorrect float16 constant folding
segfault v3d raspberry Pi5 h265 hw decoding / regression mesa 26.1.3/26.1.4
some dEQP-VK.image.host_image_copy failures on xe2/Xe3 with 32bit
std430 layout not supported on older GLSL profiles with GL_ARB_shader_storage_buffer_object extension
the new mesa 26.1.0-1 is causing mutliple graphical issues on intel i5-2400 and higher (integrated graphics)
tu: assertions in ir3_ra.c: Assertion `physreg != (physreg_t)~0’ failed
turnip: Blender: viewport contents invisible with UBWC
v3dv: Compute shader crashes for unknown reason
venus: rare flake in dEQP-VK.wsi.android.swapchain.render.basic
venus: typo in vn_queue_submit_2_to_1
vulkan/runtime: GetPipelineBinaryDataKHR incorrectly assumes `pPipelineBinaryDataSize` must be 0 initialized
vulkan/wsi/win32: bgcolor of Vulkan app window becomes lighter under venus
vulkan: 5.10 and 5.15 LTS kernels require implicit sync support
wsi: Suspicious assertion in wsi_create_buffer_blit_context
zink/ci: switch all traces jobs to surfaceless+gbm
Changes¶
Adam Jackson (7):
zink: consolidate resource_create error paths
zink: extract format list setup from create_image
zink: extract pNext chain construction from create_image
zink: extract memory binding from create_image
zink: replace image negotiation with candidate-based approach
zink: stop find_good_mod from mutating ici in place
nvk: use unsigned comparison in UBO bounds checks
Adam Stylinski (1):
nv30: fix an issue when this push buffer is NULL
Aditya Swarup (2):
anv/pps: Use counter block to stay consistent with Perfetto
intel/test: Add support for Perfetto counter groups
Adrián Larumbe (10):
pan/kmod: Fix minor version number check for USER_MMIO_OFFSET ioctl
pan/kmod: fix double syncop count sum when populating vm_bind syncs
drm-uapi: Sync the panthor header
pan/kmod: Use kernel-reported page sizes for new VM when available
pan/kmod: Pass signal and wait syncs separately
pan/kmod: Handle sync object signals in Panthor’s vm_bind
pan/kmod: Introduce sparse binding
pan/kmod: Introduce vm_op buffering and sparse mapping emulation
panvk: Use pankmod instead of panthor drm interfaces in bind queues
panvk: Talk directly to pankmod when binding sparse resources
Agate, Jesse (2):
amd/vpelib: separating frontend programming
amd/vpelib: Fix blending hang issue
Ahmed Hesham (10):
pan/bi: Restore b3210 as a valid swizzle
pan/nir: Fix 8 and 16 bool reduction lowering
pan/bi: Fix function temp lowering with 64-bit pointers
pan/bi: Fix MKVEC.v2i8 src2 swizzle lowering
pan: report async CSF group faults via context reset status
clc: fix fp16 fallback mask for remquo
nir: fix vectorising phis with mixed chased sources
rusticl: enable panfrost by default
ci: Add OpenCL-CTS to GL test infrastructure
pan/ci: Add OpenCL-CTS quick job for Mali-G610
Aitor Camacho (37):
kk: Add poly dependency to KK
kk: Reuse as much poly utilities as possible for unrolling
kk: Increase maxFragmentCombinedOutputResources to KK_MAX_DESCRIPTORS
kk: Add ds state to fragment key since it’s part of the pipeline we compile
kk: Fix subgroup failures on M1/2 due to bcsel
kk: Fix global_store writemask
kk: Rewrite force position output pass to use lowered io
kk: Add residency set to queues
kk: Add grid struct for dispatches for convenience
kk: Correctly report failures when compiling precompiled shaders
kk: Rework draw dispatch
kk: Rework shader compilation to handle more than 2 stages
kk: Implement tessellation
kk: Use index element size instead of Metal enum to avoid asserts
kk: Use subgroups for tessellation prefix count since they are now fixed
kk: Move poly data out of root buffer
kk: Clean up per draw upload for tessellation stage
kk: Move to Metal4 command encoding
kk: Add GPU hang detection
kk: Disable workarounds 1-6 in macOS 27
kk: Reduce root buffer pointer by replacing it with the GPU address
kk: Expose texture max dimensions based on GPU family
poly/lower_tcs: Preserve TCS barriers against lowered outputs
kk: refold combined image/sampler packing after vars_to_ssa
kk: defer cmd buffer submission and lighten compute barriers
kk: Record command buffers live and replay only on resubmit
kk: Fix flrp signed-zero preservation with float_controls2
nir: compute acos in 32-bit then downgrade to 16-bit
kk: Fix metal import assert
kk: Handle per alu math controls in MSL
kk: Use isnan(x) for x != x and !isnan(x) for x == x
kk: Compile all shaders with fast math
kk: Expose shaderSignedZeroInfNanPreserveFloat16/32
kk: Implement VK_KHR_shader_float_controls2
kk: Metal’s precise functions are only fp32
kk: expose Vulkan 1.4
kk: Fix icd json api version
Alejandro Piñeiro (9):
pan/midgard: reorder nir_shader_compiler_options alphabetically
panfrost: add explicit casts when assigning ~0 and 0 to enum-typed fields
cs_builder: fix trailing comma in cs_builder_init
panfrost: track active endpoint scoreboard slot
panfrost: define scoreboard slots
panfrost: add Perfetto render stage tracing support (v10+)
panfrost: add u_trace indirect capture for compute dispatches (v10+)
panfrost: add Batch, Barrier and CacheFlush Perfetto render stage (v10+)
docs/perfetto: panfrost now supports render stages
Aleksi Sapon (2):
llvmpipe: remove unused SSE rasterization code
llvmpipe: fix overflow in rasterizer
Alessandro Astone (4):
gallivm: Fix armhf build against LLVM 22
radv: Support explicit DRM format modifier from android gralloc
anv: Support VK_ANDROID_native_buffer older than version 11
hasvk: Fix android build with android-strict=false
Alexander Slobodeniuk (1):
radeonsi: fix conformance window emission in the SPS
Ali, Nawwar (1):
amd/vpelib: update shaper config size
Allen Ballway (2):
vulkan/android: Set COLOR_ATTACHMENT_BIT for external format resolve
vulkan/android: Map AHARDWAREBUFFER_FORMAT_Y8 to VK_FORMAT_R8_UNORM
Alyssa Rosenzweig (288):
nir/opt_generate_bfi: avoid trivial instructions
jay: strengthen assert
jay: drop dead code
jay: generalize last kill code
jay: reduce zeroing
jay: reduce calloc to malloc when memsetting after
jay: fix SEL implied pipe
jay: fix the source pinning code
jay/register_allocate: use standard builder name
jay/opt_dead_code: handle predication
jay/lower_pre_ra: skip predication
jay/assign_flags: handle predicated CMP
jay/register_allocate: tie predicated-defaults
jay: allow predication of pure-flag instrs
jay: fold logic ops
jay/test-optimizer: fuse before/after cases
jay: test logic op fusing
jay: fix simd32 deswizzle
jay: improve spiller debug
jay: fix spiller coupling code
jay: don’t print internal without the flag
jay/register_allocate: don’t depend on indexing
jay: call DCE an extra time
jay: refuse to propagate ADDRESS copies
jay: relax mov type check
jay: fix SEL types
jay/print: deal with bare r0 copies
jay/to_binary: handle packing accumulators
jay: validate non-SSA accumulators
jay/register_allocate: start using accumulators
jay/lower_post_ra: remove SWAP macro
jay/lower_post_ra: drop old 2<–>8 lowering
jay/ra: don’t reserve registers when not spilling
jay/ra: use accumulator for memory copies
jay/ra: use accumulator for memory swaps
jay/ra: use accumulator for stride=4 swaps
jay/ra: drop memory copy reordering
jay/ra: only use stride=4 temps
gallium: Drop users of post-processing filters
gallium: Drop post-processing filters
nir/opt_algebraic: add redundant u2u32/unpack_64_2x32_split_x patterns
brw/nir_lower_cs_intrinsics: do some math at 16-bit
nir/opt_reassociate: fix exactness bug
jay/assign_flags: refactor for next commit
jay/assign_flags: don’t burn a null flag
jay/assign_flags: don’t burn a flag for ballots
jay: shrink stack allocation
jay: jayize swsb print
jay: consolidate file prefixes
jay: fix 16-bit predicated compares
jay: drop jay_exec_mask
jay: inline jay_control()
jay/opt_propagate: fold uflag copies
jay/opt_propagate: disable f64 opts for now
jay: introduce a physical control flow graph
jay: drop UGPR->UMEM spilling path
jay: convert to LCSSA
jay: do not copyprop ballots globally
jay: propagate inverse-ballots only locally
jay: check for inverse-ballots in jay_uses_flag
jay: smarten predication pass
jay: adjust flag replication
jay: predicate NoMask instructions in uniform IF’s
jay: drop a bunch of stale TODO and XXX
jay/to_binary: rename grf -> phys_reg
jay/to_binary: fix packing of simd-split accumulators
jay: model MAC
jay: do moves on the float pipe where possible
jay: move simd32 deswizzling to float pipe
jay: assign accumulators post-RA
jay/lower_scoreboard: elide more dependencies
jay/lower_scoreboard: refactor wait pipe code
jay/lower_scoreboard: fix tracking for A@* and *@7
jay/lower_scoreboard: refactor SYNC.nop insertion
jay/lower_scoreboard: use .src annotations
jay/lower_scoreboard: be the sole emitter of SYNC
jay/lower_scoreboard: use SYNC.allrd/allwr
jay: swap predication/acc pass order
jay: add JAY_DEBUG=noacc option
jay: fix bfn with 0xffff constant
jay: elide atomic dests
jay: optimize pack_32_2x16_split(#0, x)
jay: make indirect push data blow up more obviously
jay: fix comment
jay: have proper UNDEF
jay: clarify development model
jay/lower_scoreboard: fix trivial scheduling
jay/lower_scoreboard: refactor
jay/lower_scoreboard: run RegDist globally
jay/lower_scoreboard: factor regdist logic out
jay/lower_scoreboard: control flow is int pipe
jay/lower_scoreboard: compact inst_exec_pipe
jay/lower_scoreboard: rename gpr_range -> key
jay/lower_scoreboard: use CFG for RegDist scoreboarding
jay/lower_scoreboard: use sbid syncs to elide regdist deps
jay/register_allocate: set num_regs[MEM] properly
jay/register_allocate: tweak roundrobin heuristic
jay/opt_propagate: fix NOT propagation
jay/opt_propagate: propagate undefs
pan/mdg: make clang warning quiet
jay: relax fragment payload layout
jay: insert simd32 deswizzle in a dedicated pass
jay/liveness: remove pointless bitset init
jay/liveness: speed up physical CFG merging
jay/liveness: drop redundant source filtering
jay/lower_scoreboard: handle accumulator hazard
jay/lower_scoreboard: add asserts on key bounds
jay/opt_propagate: avoid branching on poison
jay: annotate pure sends
jay: factor jay_op_(starts,ends)_block queries
jay: schedule for pressure
jay: fix omask on single sample
jay: hack for sample position
jay: use new fs payload variable more
brw/eu_validate: relax EOT requirements on Xe2
CODEOWNERS: add Jay
brw,jay: add use_src_xy prog data field
jay/liveness: use jay_foreach_preload
jay: legalize shuffle(ugpr) for now
jay/spill: fix reload array size issues
jay/register_allocate: inline silly helper
jay/register_allocate: remove out of date comment
jay/lower_spill: use 1 less temporary
jay: fix FS reading too many sysvals
jay: add zip_ugpr16 instruction
jay: pack jay_stride
jay: stop asking for stride=4 ugpr’s
jay/validate_ra: use jay_def_stride
jay: drop unneeded #include
jay: rewrite partition handling
jay: merge partition blocks
jay: remove send split hack
jay: allow simd32 gl_SamplePosition
jay: avoid overflow affinities with large UGPR vecs
jay: allow SIMD1 imageStore()
jay: hide MAD->MAC behind !JAY_DEBUG=strict
jay: gate early EOT code behind =strict
jay: renumber reg files predictably
jay: simplify uniformity checks
jay: allow npot operands in RA
brw: nir_lower_constant_convert_alu_types only once
intel/gen: remove dead #include
intel/gen: drop noisy build spam
jay/validate_ra: validate against partition
jay/register_allocate: do not treat reserved regs as free
jay/register_allocate: remove remnant of old partition code
jay/register_allocate: split out jay_stride.c
jay/register_allocate: drop #include
jay/partition: validate we don’t generate g127<2>
jay/partition: pick better partitions
jay: limit stencil export to simd16
jay/to_binary: relax packed float restriction
jay: generalize jay_extract_range_post_ra
jay: rework lane ID calculations
jay: replace GPR_FROM_UGPRs with a simple CVT
jay: replace BYTE/WORD_PACK with a simple MOV
jay/register_allocate: don’t hang if a block is missing
jay/partition: reduce 16-bit partitioning more
jay: drop dead if
jay: clang-format
jay/lower_scoreboard: allow multiple jumps
jay: lower JAY_OPCODE_LOOP_ONCE earlier
jay/to_binary: big clean up post-gen
jay: avoid bogus copyprop with cmods
jay: add unit test for bogus copyprop case
jay/validate: add validation for bogus uflag cases
jay: fix 8/16-bit inline_data loads
jay/assign_flags: require ballots to be in the balloted src
jay: fix mismatched files with predication
jay: workaround the while bug
jay: follow source order for mad/bfe
jay: fix last-use accounting with ARF sources
jay: allow null in jay_collect_vectors
jay: uniformize bti indirects
jay: optimize out more early eot related copies
jay/register_allocate: make phi webs conservative
bin: add drm-shim script
jay/lower_pre_ra: allow immediate on bfe
anv: enable jay ray query
nir/opt_sink: sink more Intel block instructions
nir/lower_terminate_to_demote: tweak terminate_if lowering
jay/spill: spill at definitions
jay/spill: do initial find-and-replace for ugpr spilling
jay/spill: unstub rematerialization
jay/spill: refactor
jay/spill: drop sketchy heuristic
jay/spill: simplify limit()
jay: fix bogus unit tests
jay/validate: check for mixed ugpr/gpr problems
jay/register_allocate: simplify split copy logic
jay/register_allocate: don’t search a 2nd UGPR temp
jay: allow more 3-src imms
jay: introduce accumulators into the partition
jay: pool constants per block under pressure
jay: remove #include
jay: cache message headers locally
jay: forbid 8-bit immediate prop
jay: autopep8
jay: manually format jay_type_for_glsl_base_type
jay: clang-format
jay: track skip_helpers
jay: rewrite demote/terminate/helper/halt handling
nir/opt_dead_cf: delete redundant returns/halts
intel, nir: Add {load,store}_global_intel intrinsics
jay: distinguish physical & logical loop headers
jay/validate: validate backedges
jay/lower_scoreboard: simplify trivial swsb
jay/lower_scoreboard: fix barriers in trivial SWSB
jay/test: drop SSA repair tests
jay/to_binary: use dedicated addr reg for shuffle
jay/lower_pre_ra: fix oob read
jay/register_allocate: fix file prefixes
jay/opt_dead_code: drop stale todo
jay/opt_dead_code: handle phi properly
jay/spill: add an assert()
jay/spill: repair as we go
jay/spill: implement ugpr spilling
jay/spill: don’t try to remat mov_imm64
jay/spill: do lazy reloading instead
jay: remove unused SSA repair pass
jay: drop #include
jay/assign_flags: fix ballot handling
jay: fix barycentrics
jay: improve the stride partition heuristic
jay/lower_spill: rename to make easier to follow regs
jay/lower_pre_ra: fix f64 negate
jay: clang-format
Revert “rusticl: fix leak in `util_queue`”
jay: fix EOT with indirect message descriptors
intel: make more multisampling/coarse state static
jay/to_binary: avoid overflowing address reg
jay: add jay_bare_regs helper
jay: use jay_bare_regs
jay: drop jay_extract_range_post_ra
jay: rename post-sched lowering
jay: zero a0 before divergent shuffles
jay: consolidate spilling call into a single file
jay: fix JAY_DEBUG=spill
jay/lower_spill: set cursor explicitly
jay: fix ugpr reloads in divergent control flow
jay: fix printing internal shaders
people: add Lucas Fryzek
util: add __bitset_zero, __bitset_copy helpers
treewide: use __bitset_{zero,copy}
jay: remove #define
jay: unify inline_data and push_data loads
jay: remove duplicated bf16 check
jay: remove pointless (and wrong) bf16 check
jay: drop #include
jay: treat init_helpers as a start instruction
jay: print the shader after spilling before RA
jay/print: make predication syntax more explicit
jay/schedule: move schedule into ctx
jay/schedule: move function into context
jay/dag: inline dag code
jay/dag: defer parent_count calculation
jay/dag: split into a mutable and immutable part
jay/dag: add jay_dag_print helper for debug
jay/dag: add jay_dag_iterator_reset helper
jay/schedule: add per-block data structure
jay/schedule: edit comment about pressure
jay/schedule: split block analysis from scheduling
jay/schedule: flip the signs on pressure calcs
jay/schedule: account for demand per-file
jay/schedule: fix accounting for dead flags
jay/schedule: add missing whitespace
jay/schedule: early-out invalid schedules as we go
jay: add mlen to SEND
jay: add real cycle model
nir: lower boolean shuffle without subgroup size
nir/opt_algebraic: optimize ~x != x
gen/print: include pipes for jay_print
intel/gen: allow dead loops
iris: use u_default_get_sample_position
crocus: use u_default_get_sample_position
intel: remove intel_get_sample_positions getter
anv: use vk_standard_sample_locations
hasvk: use vk_standard_sample_locations
jay/spill: fix variable shadowing
jay/lower_post_ra: fix simd16 flag zeroing
jay/opt_propagate: fix inverse_ballot(bfn)
jay/lower_helpers: fix unconditional discard
jay: lower boolean shuffles
jay: make uniformity explicit
jay: support as_uniform
jay: drop #include
jay: rewrite flags
intel: fuse off Jay in Mesa 26.2
Andrzej Datczuk (2):
radv: enable advertising of VK_KHR_pipeline_library under llvm
radv/rra,rmv: fix device id written into trace files
Anna Maniscalco (1):
ir3: Skip preftech and and warmups for non bindless earlier
Arjob Mukherjee (2):
pvr: increase value of maxPerStageDescriptorStorageBuffers
pvr: increase maxPerStageDescriptorStorageBuffers to 16
Arkady Shlykov (1):
Add offset getter for image intrinsics
Arzaq Naufail Khan (1):
spirv: fix resource leak in spirv shader replacement
Ashley Smith (2):
panfrost: Fix tiler_desc assignment
panfrost: Avoid race between bo import and unref
Assadian, Navid (1):
amd/vpelib: FP16 non linear handling
Autumn Ashton (5):
nak: Expose max_warps_per_sm
nvk: Allow nvk_cmd_upload_qmd to take a custom root descriptor
nvk: Add nvk_cmd_dispatch_with_root
nouveau/cubin: Add cubin and fatbin parsers
nvk: Implement VK_NVX_binary_import
Benjamin Cheng (19):
ac/vcn: Rename VCN5 swizzle mode to GFX12
radv/video_enc: Use correct swizzle mode for VCN5 with GFX11
radv/wsi: Re-use transfer queue if it exists
ac/parse_ib: Add parsing for variable slice mode
radeonsi/video: Cleanup dpb buffer
radeonsi/mm: Disable variable slices when bad input is found
util/ycbcr: Fix adjust_to_range
util/ycbcr: Add a narrow range RGB coeff helper
gallium/vl: Fix RGB narrow range conversions
draw: Add lower_opcodes NIR pass
mesa/st: run the lower_opcodes pass for draw shaders
radv/video: Set accurate minQp/QIndex
radv/video: Report MULTIPLE_SLICE_SEGMENTS_PER_TILE_BIT
ac/video: Add {min,max}_qp to video enc caps
radv/video: Use {min,max}_qp caps from ac
mesa/st: use col0 attrib from provoking vertex for feedback
ac/surface: Remove GFX10 limitation for FORCE_SWIZZLE_MODE
ac/vcn_enc: Disable var slice for preencode with VCN4
vl/proc: Guard compositor creation with HAVE_GFX_COMPUTE
Benjamin Gaignard (1):
pan/format: Advertise support for AFBC(32x8,sparse)
Benjamin Otte (1):
Revert “lavapipe: Don’t advertise support for multiplane drm formats”
Benoît du Garreau (1):
docs: Add many missing features
Bo Hu (3):
codegen: update scripts/cereal/decoder.py
vk-snapshot: do not generate code to save vkQueueFlushCommandsGOOGLE
gfxstream: vk-snapshot: update handling of bufferview in vkUpdateDescriptorSets
Boris Brezillon (9):
pan/kmod: Don’t pass drmVersionPtr objects around
kraid: Fix cross-build
pan/kmod: Add a pan_kmod_timestamp_cycles_to_ns() helper
pan/props: Make pan_query_core_count() safe with wide shader_present bitmaps
pan/props: Split core_count/core_id_range into two helpers
pan/props: Add pan_query_perf_counter_per_block()
pan/perf: Start relying on canonical mali_perf definitions
pan/perf: Replace pan_perf_init() by pan_perf_{create,destroy}()
pan/perf: Transition to auto-generated derived counters
Boyuan Zhang (1):
radeonsi/mm: use non-tmz buf when session tmz size is 0
Brandon Jones (1):
nir/opt_algebraic: fix fabs optimization
Brendan King (1):
pvr: move some asserts in pvr_srv_alloc_display_pmr
Caio Oliveira (115):
brw: Don’t set saturate for SYNC instruction
brw: Use brw prefix to LSC helpers tied to brw
brw: Remove various unused fields
brw: Fix max_dispatch_width collection for CS with variable size
brw: Stop tracking inline parameter usage in prog_key/prog_data
intel/dev: Expose list of known platform names
brw: Fix some indentation in brw_generator.cpp
brw: Move brw_prog_data_init to a different file
brw: Remove references to SIMD4x2
intel/executor: Map the DPAS check to has_systolic
brw/tests: Remove redundant parser test
brw/tests: Stop using regions/type for null in assembler tests
brw/tests: Stop using regions/type for non-null SEND sources in tests
anv: Remove saturating cmat configurations when INTEL_LOWER_DPAS=1
anv: When using INTEL_LOWER_DPAS disable BFloat16 cmat configurations
intel: Move cmat configurations to anv_physical_device
intel/perf: Use intel_perf_context as ralloc parent of sample buffers
nir/instr_set: Fix multi-slot intrinsic index equality
intel/perf: Add helpers to get names of enums
intel/perf: Show type, data type and units in intel_perf_query_layout
nir/instr_set: Consider normalization when calculating hash
brw/scoreboard: Add disabled tests for RegDist baking on Xe2+
brw: Save original regs_written() value in register coalesce
brw: Call size_read() once in regs_read()
brw: Avoid unnecessary calls to size_read() in flags_read()
brw: Pass VGRF numbers to liveness helpers
brw: Don’t directly use regs_read/regs_written/size_read as bound for non-trivial loops
brw: Bound register coalesce rewrites by live range
nir: Add print for other cmat_description slots
compiler: Support more than 255 cols/rows in cmat descriptions
intel/compiler: Move bison command to shared meson.build
intel/executor: Add an overflow check for alloc function
intel/executor: Add performance counter support
brw: Use a single brw_compile entrypoint
brw: Move key and prog_data to base compile params
anv: Simplify code that calls brw/jay
iris: Simplify code that calls brw/jay
spirv: Stop warning about ignored invalid ArrayStride decorations
anv, brw: Use previous shader VUE map for FS input layout when available
anv: Fill VERTEX_ELEMENT_STATE before further emissions
anv: Use empty_vs_input for default VERTEX_ELEMENT_STATE
jay: Use TGL_PIPE_NONE when RegDist is zero
intel/gen: Add gen encoding module
intel/gen: Add various to_string/from_string functions
intel/gen: Add validation
intel/gen: Add function to finish structured control flow
intel/gen: Add gen_print()
intel/gen: Add gen_parse()
Revert “intel/dev: Remove unused intel_get_device_info_for_build() function”
intel/gen: Add integrated `gentool` CLI for asm/disasm
brw: Add temporary workarounds for compatibility with old parser
brw/tests: Port assembler tests to gen module
intel/compiler: Stop replicating narrow immediate values in brw_asm “compat” mode
intel/compiler: Stop forcing null source <0;1,0> region in brw_asm “compat” mode
intel/compiler: Stop forcing null destination HSTRIDE=1 on pre-xe SEND in brw_asm “compat” mode
brw, jay: Use lsc_* symbols from gen
brw, jay: Use gen_swsb and related enums
brw, jay: Use gen_sfid instead of brw_sfid
brw, jay: Use various enums from gen instead of brw
jay: Use gen_condition enum and helper
jay: Use gen module
intel/decoder: Convert to use gen module
intel/executor: Update to use gen module
brw: Make a copy of brw_generator for using gen
brw: Port the copy of generator to use gen_encoding
brw: Use the new gen based generator
brw: Add brw_to_binary() as the single codegen entrypoint
brw: Make brw_generator an implementation detail of brw_to_binary.cpp
anv: Use gen_print instead of brw_disasm
intel/tools: Use gen_print instead of brw_disasm
iris: Use gen_print instead of brw_disasm
intel/compiler: Add and use gen_update_reloc_imm()
brw: Remove old encoder, generator and related tools
intel/gen: Change validation test code to use the parser
jay: Unify macro for NIR passes
jay: Add INTEL_DEBUG=mda support
util: Add runtime parser for boolean lookup tables
intel/gen: Support symbolic print/parse of BFN function
jay: Add helpers for managing unordered instructions
jay: Handle dpas_intel intrinsic
util: Fix float8 denorm rounding to min-normal
intel: build compiler before blorp
intel/gen: Generate opcodes and their metadata
intel/gen: Drop unused format parameter from gen_inst_has_dst
intel/gen: Don’t encode zero exec_size
intel/gen: Add a gentool ‘check-roundtrip’ subcommand
intel/gen: Replace gen_parse_print_test.cpp with text tests
intel/gen: Import test cases from the brw assembler tests
brw: Remove the brw assembler tests
nir: Handle nir_var_mem_push_const in divergence analysis
intel: Change dpas_intel source order to follow DPAS
jay: Handle convert_cmat_intel intrinsic
jay: Add SIMD restriction for Math with HF
brw: Drop dead spill_writes in brw_opt_fill_and_spill
brw: Add unit tests for brw_opt_predicate_logic
brw: Refactor DPAS lowering for HF case
brw: Fix INTEL_LOWER_DPAS=1 for Xe2+
util: Add util_is_half_subnormal()
brw, elk: Fix invalid case using float-negation in combine constants
intel/gen: Use lookup tables for Gfx12+ short type encoding/decoding
anv: Initialize shader debug archive key size
brw: Fix comma placement when printing memory logical sources
brw: Use shared LSC opcode names in IR printing
brw: Remove dead mixed-size MOV special case from def copy propagation
brw: Fix initializing matrix with uniform but non-constant values
intel/gen: Fix case where IF is a loop header
brw: Report scratch memory size in shader stats
brw: Track logical scratch offsets for spill cleanup
brw: Reuse scratch slots between non-interfering spilled VGRFs
nir: Account for cmat memory accesses in copy_prop_vars
anv: Include build and device identity in shader binary UUID
brw: Match fill/spill optimization scratch accesses by logical offset
intel/gen: Fix Gfx11 3-src accumulator file encoding
nir, spirv: Match debug printf argument alignment with u_printf
nir, spirv: Pad 3-component debug printf arguments to 4 components
Caius-Moldovan-img (5):
pco: Replace nir_shader_lower_instructions with nir_shader_*_pass
pvr: set SMP component count for TQ frag load shaders
pco: Fix metadata invalidation
nir: Fix trailing comment generation for variable naming
pco: Remove hardcoded metadata location
Calder Young (40):
anv: Fix address bit masking for indirect SBTs
anv: Fix support for indirect SBTs on Xe3+
anv: Store batch buffers in a null-initialized VMA heap
anv: Add padding to the shader heap to manage EU prefetch
isl: Add usage flag to force SurfaceArray to false
isl: Add additional alignment/padding requirements to prevent overfetch
isl: Optimize the sampler cache to overlap as few 64B cachelines as possible
isl: Add function to calculate the amount of overfetch for an unpadded surface
isl: Add and use isl_tiling_get_intratile_range_el/sa
blorp: Work around sampler overfetch for buffer copies
anv: Make sure robust UBO access does not fault
brw: Avoid rounding every convergent block load up to a full register
brw: Avoid vectorizing loads in NIR if it could extend into a different page
anv: Disable scratch page by default on Xe KMD
intel_hang_replay: Don’t force scratch page on Xe KMD unless explicitly requested
isl: Make sure isl_device::requires_padding is always initialized
anv: Fix some usage flags not propagated to ISL for explicit layouts
brw: Allow instruction reordering around memory writes
brw: Add support for ACCESS_CAN_REORDER memory ordering
spirv: Fix debugPrintfEXT not working with multiple arguments
brw: Add workaround pass for shaders using derivatives in control flow
anv: Add workaround for vertex explosions in Split Fiction
jay: Use gen_arf enums instead of jay_arf
jay: Use gen_names.h to print CMODs and ARFs
jay: Do not propagate ARF src unless its src0
jay: Add support for saturating f2i16 and f2i8 NIR opcodes
jay: Disable avoid_ternary_with_two_constants when using jay
brw: Move ray payload bitfield generation to NIR
brw: Move topology id helper intrinsics to NIR
jay: Disable SIMD32 if ray queries are used
jay: Implement ray tracing topology id intrinsics
jay: Implement ray tracing trace intrinsics
nir: Do not mask helper lanes of writes if ACCESS_INCLUDE_HELPERS is set
anv: Track more error codes from certain IOCTLs
intel: Add common utils for page fault reporting
anv: Add function to get the list of page faults
anv: Print page faults whenever a queue gets banned
anv: Enable support for VK_EXT_device_fault/VK_KHR_device_fault
intel: Add function to get the number of SBIDs from device info
jay: Implement the global SBID scoreboarding pass
Caleb Callaway (6):
docs: fix Intel tracepoints.py path
perfetto: v56.1 update
pps: generate forward-looking counters
perfetto: suppress array-bounds warning
perfetto: suppress stringop-overflow warning
anv: hide Intel vendor ID for Cyberpunk
Casey Bowman (2):
anv: Add option to disable HiZ via drirc
anv: Set anv_disable_hiz for Sons of the Forest
Charmaine Lee (1):
nir_to_tgsi: fix shared memory index
Christian Gmeiner (149):
mesa/st: Extend st_context_invalidate_state with meta-op flags
mesa/st: Convert st_cb_drawtex to use st_context_invalidate_state
mesa/st: Convert st_cb_bitmap to use st_context_invalidate_state
mesa/st: Convert st_cb_clear to use st_context_invalidate_state
mesa/st: Convert st_cb_drawpixels to use st_context_invalidate_state
mesa/st: Convert st_cb_readpixels to use st_context_invalidate_state
mesa/st: Convert st_cb_texture to use st_context_invalidate_state
panvk: Advertise VK_EXT_shader_uniform_buffer_unsized_array
panvk: Advertise VK_EXT_dynamic_rendering_unused_attachments
panvk: Implement vkCmdFillBuffer with panlib kernels
panvk: Wire up VK_EXT_conservative_rasterization on v11+
egl: Switch to mesa_log(..)
etnaviv: Bypass BGRA-internal optimization for shared resources
etnaviv: Convert PE-internal BGRA to RGBA when flushing shared resources
etnaviv: Select texture format dynamically for shared RB_SWAP resources
etnaviv: Add per-RT frag_rb_swap shader key and NIR lowering
etnaviv: Use shader R/B swap for LINEAR_PE shared resources
panvk: Apply sample mask in single-sample mode
panvk: Advertise VK_EXT_extended_dynamic_state3
compiler/rust: Move VecPair from NAK to shared compiler crate
st/mesa: Zero MaxTextureImageUnits for unsupported stages
etnaviv: blt: Add sRGB support to blt_imginfo
etnaviv: Map R8G8B8A8_SRGB to BLT_FORMAT_A8R8G8B8
etnaviv: blt: Add BLT format conversion support
lavapipe: Skip advanced blend lowering when blending is disabled
lavapipe: Lower advanced blend at draw time when its state is dynamic
lavapipe: Enable extendedDynamicState3ColorBlendAdvanced
mesa: Allow GL_TEXTURE_IMMUTABLE_LEVELS query on GLES3
mesa/main: Add trace dispatch plumbing for MESA_VERBOSE=api
mesa/main: Auto-generate MESA_VERBOSE=api trace dispatch
etnaviv: Update headers from rnndb
etnaviv: Use the full 64-bit clear value for 64bpp render targets
etnaviv: Use integer texture formats for R32/RG32 integer textures
etnaviv: blt: Don’t sRGB-roundtrip same-encoding copies
etnaviv: Drop unused num_loops shader stat
vulkan/wsi: Constify wsi_instance_supports_google_display_timing(..)
panvk: Advertise VK_GOOGLE_display_timing
panvk: Advertise VK_KHR_shader_fma
panvk: Shuffle local ids for quad derivatives
panvk: Advertise VK_KHR_compute_shader_derivatives
compiler/rust: move ACORN PRNG to shared location
panvk: Move maxFramebuffer limits to defines
panvk: Derive viewport limits from the framebuffer dimension
vulkan/runtime: Track rasterization_order_access in pipeline state
vulkan/runtime: Add rasterization_order_access to dynamic graphics state
panvk: Disable FPK and force late ZS for rasterization order access
panvk: Advertise VK_EXT_rasterization_order_attachment_access
etnaviv: Flush texture caches after clears
etnaviv: Add a helper for the 128-bit second-plane offset
etnaviv: Wrap pipe_framebuffer_state in etna_framebuffer_state
etnaviv: Lay out 128-bit color FBO as paired G32R32F render targets
etnaviv: Emit paired 128-bit sampler descriptors
etnaviv: NIR pass to lower 128-bit color RT and texture access
etnaviv: blt: Use block-layout offset for 128-bit second-plane blit
etnaviv: Save the framebuffer without 128-bit companion slots
etnaviv: rs: Support 128-bit color clears
etnaviv: Limit nir_lower_fragcolor(..) to advertised render targets
etnaviv: Advertise 128-bit color formats as renderable and samplable
etnaviv: Disable TS per render target on mixed TS modes
etnaviv: Update headers from rnndb
etnaviv: Set per-RT sRGB bit on non-zero render target slots
etnaviv: Support split sampler for 128-bit formats on the state path
etnaviv: Gate 128-bit render targets on HALF_FLOAT
pan/texture: Reuse the layer range as a Z-slice range on 3D views
panvk: Slice 3D storage image views on Valhall
panvk: Slice 3D storage image views on Bifrost
panvk: Advertise VK_EXT_image_sliced_view_of_3d
nir: Add load_tile_image intrinsic
spirv: Implement SPV_EXT_shader_tile_image
panvk: Lower tile image reads
panvk: Advertise VK_EXT_shader_tile_image
mesa/main: Keep RealPublished in sync when glthread toggles with api trace
panvk: Iterate the common queue list instead of a per-family array
panvk: Advertise VK_KHR_internally_synchronized_queues
pan/va: Only widen constants on 32-bit instructions
panvk: Advertise VK_KHR_workgroup_memory_explicit_layout
etnaviv: blt: Zero-initialize conv_swizzle
mr-label-maker: Add rule for rust files
nir: Fix lower_fround_even to round half to even
panvk: Add panvk_address_binding_report() helper
panvk: Report address binding for device memory
panvk: Report address binding for buffers
panvk: Report address binding for images
panvk: Report address binding for internal allocations
panvk: Report address binding for descriptor sets
panvk: Advertise VK_EXT_device_address_binding_report
u_transfer_helper: Add U_TRANSFER_HELPER_Z32F_S8_IN_Z24S8
u_transfer_helper: Convert Z32_FLOAT in the Z32F_S8_IN_Z24S8 path
etnaviv: Route resources through u_transfer_helper
gallium: Add native_fp32_depth cap to gate ARB_depth_buffer_float
etnaviv: Emulate Z32_FLOAT (DEPTH_COMPONENT32F) as D24S8
etnaviv: Lower depth32f shadow compare in the shader
etnaviv: Emulate Z32_FLOAT_S8X24_UINT (DEPTH32F_STENCIL8) as D24S8
etnaviv: blt: Address emulated depth32f as D24S8
etnaviv: Size depth32f transfer staging by the internal format
etnaviv: Support stencil blit of depth32f_stencil8
etnaviv: rs: Resolve depth surfaces
etnaviv: blt: Resolve depth_component24
etnaviv: Address emulated depth32f as D24S8 in resource copies
etnaviv: Decide transfer tileability on the physical format
etnaviv: Support MSAA resolve of emulated depth32f
etnaviv/isa: Add bit_insert instruction
nir: Add bitfield_insert_etna opcode
etnaviv: Support native bitfield_insert
etnaviv: hwdb: Add UNIFIED_SAMPLERS feature
etnaviv: Detect unified sampler support
etnaviv: Extract sampler descriptor emit helpers
etnaviv: Rework descriptor emit into a single loop
etnaviv: Implement unified sampler allocation
etnaviv: Allow depth only or stencil only MSAA resolves
etnaviv: Allow MSAA resolve of stencil only buffers
etnaviv: Fix sampler view leak on unsupported texture target
etnaviv: Extract texture descriptor fill helper
etnaviv: Keep texture descriptor template CPU-side
etnaviv: Compose texture descriptors at emit time
etnaviv: Rename etna_sampler_view_update_descriptor()
etnaviv: Add perf debug for texture descriptor recompose
util: Add u_shader_variant_cache
util/u_shader_variant_cache: Add hash + equal callback pair
etnaviv: Adopt u_shader_variant_cache
draw/llvm: Adopt u_shader_variant_cache
llvmpipe: Move FS variant LLVM types to temporary JIT struct
util/u_shader_variant_cache: Add cap + refcount + per-list eviction
llvmpipe: Migrate FS variants to u_shader_variant_cache
llvmpipe: Move setup variant LLVM function pointer to a local
llvmpipe: Migrate setup variants to u_shader_variant_cache
llvmpipe: Move CS variant LLVM types to temporary JIT struct
llvmpipe: Migrate CS/task/mesh variants to u_shader_variant_cache
llvmpipe: Remove dead USE_GLOBAL_LLVM_CONTEXT path
llvmpipe: Move shader caches and LLVMContext to screen
draw: Decouple GS/TCS/TES shader CSOs from draw_context
draw: Decouple VS shader CSO from draw_context
draw: Move current_variant slot from CSO to draw_context
draw/llvm: Bound per-stage variant caches with cap + pin slots
draw: Move per-draw state off the GS/TCS/TES shader CSOs
draw: Decouple mesh shader CSO from draw_context
llvmpipe: Enable PIPE_CAP_SHAREABLE_SHADERS
panvk: Restore push descriptor dirty bit after meta operations
panfrost: Add pan_nir_lower_image_64bit NIR pass
panvk: Alias R64 storage images to R32G32_UINT for color clear
panvk: Bounds-check SSBO accesses for robust storage buffer access
pan/format: Add R64_UINT/R64_SINT format entries for v9+
nir/lower_robust_access: drop out-of-bounds buffer stores
panvk: Wire up VK_EXT_shader_image_atomic_int64 on v9+
etnaviv: Only emit index buffer state when it changed
etnaviv: Skip state emission when nothing is dirty
mesa/st: Invalidate FS sampler views after PBO texture transfers
nir/opt_vectorize: Remove the phis that were combined
etnaviv: Drop stale stream output relocs in etna_update_hwxfb(..)
Christian Meissl (1):
nir/lower_tex: skip external texture YUV lowering for query instructions
Christoph Neuhauser (1):
anv: Add compute only divergent atomics fusion optimization for Blender Blender uses atomic operations as part of its virtual shadow mapping implementation. Virtual shadow mapping page tagging in compute shaders benefits from divergent atomics fusion, while fragment shaders doing the atomic raster step in general have worse performance with this optimization turned on. Thus, an option is added to only apply divergent atomics fusion to compute shaders in ANV, and this option is enabled for Blender.
Christoph Pillmayer (19):
pan/bi: Fix source swizzle in bi_repair_ssa
pan/bi: Fix format in bi_repair_ssa
pan/kmod: Fix uninitialized timestamp info
pan: Set nir_shader::source_blake3 for internal shaders
pan/clc: Set source_blake3 for each precompiled variant
pan: Add BIFROST_MESA_DUMP_DIR option to dump shader binaries
pan/kmod: Add l2 features to pan_kmod_dev_props
pan/props: Add BUS_WIDTH query to pan_props
pan/perf: Generate derived counter definition from Arm’s XMLs
pan/kmod: Add perf counter api
pan/kmod: Add a perf implementation to the panfrost backend
pan/pps: Delegate more tasks to PanfrostPerf
pan/perf: Make counter sum across block instances optional
pan/pps: Output counters per block
pan/perf: Add timing related getters/setters and use them
pan/perf: Use new kmod api
pan: Fix BIFROST_MESA_DUMP_DIR
pan/kraid: Fix 32bit hw_runner builds
pan/perf: Fix 32bit build of panquick
Collabora’s Gfx CI Team (21):
Uprev VVL to 8474616c3095756c52c1b810b21bd1366b3fc909
Uprev VVL to 4acd00c7a0665c9b1d01604e5fe1454837f87134
Uprev VVL to 6a6182c0edb35cba7bab0abc61eaff82d11022fb
Uprev VVL to d55be6264a17cd28f436805973b12f12a5d22f2f
Uprev ANGLE to 7772c5602d59140204494967ba8ebdf801180054
Uprev Piglit to 6fd29fe44f8857b876a67bee962919635f22ecc8
Uprev VVL to 36187ee9f2074609d3ec56fa1a315c191366b688
Uprev ANGLE to a793c75398c746f3f8a08fd2e74dfc4dff07a0c9
Uprev VVL to 315d28985ebd1ff9a2e4380e34ed2d8ebe487531
Uprev ANGLE to 196d1b79eadbd8fdbf0590092266bb2d87264988
Uprev VVL to d2b091858d12802cc1c3722c81a9f3d865d833d4
Uprev VVL to 2ab77a01659e3e46d6ca8a25425b19b3425adb11
Uprev ANGLE to 8e09325ebad45c7e11630a79754361e965e5fab0
Uprev VVL to e17d63f8fcd967b2ff91efcb8607d2c9ab962e23
Uprev ANGLE to 836636df1b06034b39a4a2a68b811ecf6f9b674f
Uprev VVL to 5e72d40395d07930c935d0624cd7db5f1a144c5e
Uprev ANGLE to a4eea1fbedace7a03bae52cd1bb9d6ebdfffb0f7
Uprev VVL to 6906fd94f2f422beb682d43d8b1af872aaff97b0
Uprev Piglit to 9a0eab5e1f7f009b4f72c25d23faf937b38354a6
Uprev VVL to e875181c4f0fab1bd0c4926c21a21de0bd967bb3
Uprev ANGLE to 7a960f346c57daae4b1615ead302b4c58af12258
Connor Abbott (35):
ir3: Don’t reset immediate count to 0 after lowering
ir3: Use correct immediate size for constlen calculation
tu: Optimize sync2 event handling in the non-asymmetric case
tu: Don’t zero-initialize query pool
turnip, ir3: Use shader for vertex input count
tu: Support VK_KHR_maintenance9
tu: Fix LRZ+FDM offset+secondaries
tu: Disable LRZ when resuming if the GPU doesn’t support tracking
tu: Zero out unused parts of descriptors
compiler/shader_info: Introduce occupancy_bounded_workgroup_fairness
spirv: Remove redundant OpExtInst handling
spirv: Use correct opcode in non-semantic OpExtInst handling
spirv: Implement concurrent workgroup hint from vkd3d-proton
freedreno: Document SP_CS_CNTL_0::COMPUTERRMODEEN
ir3, tu, freedreno: Plumb through round-robin mode
tu: Add TU_DEBUG=computeroundrobin
freedreno: Add round_robin_errata to device info
ir3: Implement round-robin workaround
nir/lower_amul: fix infinity recursion with different phi source order
freedreno/qrisc: Refactor section disassembly
freedreno/qrisc: Emulate multiple firmwares more accurately
freedreno/qrisc: ISA changes for gen8
freedreno/qrisc: Support DDE on gen8
freedreno/qrisc: Don’t write CP_LPAC_SQE_CNTL on a7xx+
freedreno/qrisc: Add support for gen8 firmwares
freedreno/qrisc: Add missing <decode/> to #sqe-base
freedreno/qrisc: Add extra preempt entrypoint
freedreno: Add some gen8 control registers
freedreno/qrisc: Allow limited relocation expressions
freedreno/qrisc: Allow labels on two-src ALU instructions
freedreno/qrisc: Bump label/instruction count
freedreno/qrisc: Add support for “absolute” label references
freedreno/qrisc: Test new label features
tu: Fix resetting command streams with writeable BOs
tu: Fix condition for skipping emitting aprons
Daivik Bhatia (6):
broadcom/compiler: Add explicit NOP instruction at page boundaries
pan/nir: fix GNU compilation error with clang
broadcom/compiler: add support for null descriptors
v3dv: Implement and enable nullDescriptor support
nir/opt_copy_prop_vars: kill stale entries when source deref is written
rocket: simplify input/output tensor creation
Daniel Lang (3):
radv/meta: fix samples datatype in radv_meta_nir
nir: change type_size return type to unsigned in nir_lower_{amul,io}
docs: Add GL_ARB_map_buffer_range and GL_ARB_vertex_array_object to etnaviv
Daniel Schürmann (39):
nir: add nir_loop::do_while to indicate do-while loops
vtn: set nir_loop::do_while during spirv_to_nir()
glsl_to_nir: set nir_loop::do_while
nir/opt_loop: Don’t peel initial break from do-while loops
nir/opt_loop: stop recursion at loop header phi in can_constant_fold()
nir/opt_loop: always try to peel initial break from loops with unrolling hint
nir/opt_algebraic: use imul24_relaxed for lowered dot4x8_add
nir/opt_algebraic: add some imul24_relaxed pattern
nir/opt_constant_folding: create const_value_for_alu() helper
nir/opt_constant_folding: constant-fold op(bcsel(), #c) -> bcsel(.., #c1, #c2)
nir/opt_algebraic: optimize downcast followed by upcast to extract
nir/opt_algebraic: extend some extract_u8 pattern to extract_i8
nir/builder: constant-fold nir_mov_alu() if requested
nir/lower_bit_size: use nir_builder::constant_fold_alu
nir/lower_bit_size: use nir_def_replace() instead of nir_def_rewrite_uses()
nir/lower_bit_size: skip conversion for more opcodes
anti-lag: rework wait time calculation
aco/assembler: Fix s_inst_prefetch insertion after loop latch rotation
aco/assembler: pass std::vector to insert_code
nir: remove fixed-sized nir_op_bitz / nir_op_bitnz
nir: disallow converting to sized booleans from nir_type_convert()
nouveau: don’t handle 8- and 16-bit comparisons
panfrost: remove fixed-sized binop reductions
panfrost: replace fixed-sized with unsized comparison opcodes
panfrost: replace fixed-sized bcsel with unsized bcsel_pan
nir,panfrost: remove 8-bit and 16-bit booleans
aco/ra: Fix get_reg_impl() for operand registers
aco: encode unused VOP3 operands as inline constant 0 on RDNA
aco/isel: move add64_32() to aco_isel_helpers.cpp
aco/isel: use add64_32() for nir_iadd(nir_u2u64(), ..)
nir/range_analysis: handle read_first_invocation and friends in nir_def_num_lsb_zero()
nir/range_analysis: handle phis in nir_def_num_lsb_zero()
amd/lower_global_access: refactor using state struct
amd/lower_global_access: only consider constant offsets that are aligned
amd/lower_global_access: Only consider 32-bit offsets which are aligned
amd/lower_global_access: lower SMEM offsets according to hw capabilities
aco: remove alignment handling for global SMEM loads
aco: emit global SMEM loads directly
aco: Remove SMEM offset optimization for non-buffer loads
Daniel Stone (9):
pan/afbc: Code motion for split modifier queries
pan/mod: Protect against no usage flags for 64k
pan/mod: Reorder linear modifier checks
pan/afbc: Properly validate format/parameter combinations
ci/panfrost: Switch T860 jobs to another RK3399 device type
ci/panfrost: Add two T860 OpenCL fails
symbols-check: Ignore more pthread symbols
doc/ci: Add custom-kernel testing workflow
draw: Avoid warnings for maybe-unused variable
Danylo Piliaiev (54):
tu: Fix draw call offset for LRZ warnings in secondaries
tu/perfetto: Move away from single timeline for all apps
tu: Fix CP_CCHE_INVALIDATE not being applied at the right point
freedreno: Fix CP_CCHE_INVALIDATE not being applied at the right point
tu/u_trace: Use correct u_trace destination in tu_clone_trace_range
tu/u_trace: Prevent cloning stale RB_DONE_TS results
tu/u_trace: Correct the order of tracepoints clonning for binning
tu/u_trace: Fix explicit toggle_name not being used
tu/perfetto: Add a performance warning track to perfetto
tu/perfetto: Add performance warning tracepoints
tu: Fix tu_bo_make_zombie without queues
tu: Fix double free of timestamp_copy_data->trace
tu: Don’t leak pre_chain.rp_trace, and correct u_trace_move
tu: Fix BV/BR race in tu_clone_trace_range when waiting on barrier
tu/a8xx: Fix reading border_color from sampler memory
tu: Don’t disable UBWC for D24S8+USAGE_SAMPLED+customBorderColorWithoutFormat
tu: Disable concurrent binning by default due to perf regressions
tu: Don’t enable FDM when there is FDM attachment is UNUSED
tu: Always lazy_init_vsc for tiler rendering
u_trace: Lazy init ut->linear_alloc
tu: Fix TU_CMD_DIRTY_DRAW_STATE value collision
tu: Start/End occlusion query should force depth state recalculation
tu: Change of disable_fs state should force depth state recalculation
tu/a7xx: Don’t force enable IJ_LINEAR_PIXEL for FragFace/FragCoord
freedreno/a7xx: Don’t force enable IJ_LINEAR_PIXEL for FragFace/FragCoord
tu: Disable FS in some cases even when FS explicitly writes D/S
ir3: Add resbase_ir3 intrinsic
tu: Add allow_oob_indirect_ubo_loads to device cache uuid
tu: Specify max texel buffer and storage buffer limits via GPU props
tu/a8xx: Set real storage/texel buffer size limits
tu: Add option to raise the maximum texel buffer size
tu: Enable texel buffer / SSBO emulation for known problematic games
tu: Match SW depth clear value packing with HW
tu: Match SW color clear value packing with HW
tu: Don’t process A2R10G10B10 clear values via new pack function
tu/lrz: Pick correct depth attachments in msrtss case
tu: Force GMEM mode when renderpass has MSRTSS attachments
tu: Refactor separate D32S8 and ignore D/S aspect mask for RP attachments
tu: Remove depth/stencil-specific blit src/dst helpers
tu: Remove event_blit_dst_view
tu: Enable tu_dont_care_as_load for all Kex Engine games
tu: Use application_name_match instead of exe match for workarounds
tu/a6xx: Work around D32S8 EARLY_Z_LATE_Z hang
tu: Enable tu_allow_oob_indirect_ubo_loads for Clausewitz engine
tu: Custom resolve should always use AVOID_CCU layout
tu: Fix subsampled metadata and blit emission for separate stencil
tu: Fix gfx_write_access checking for TRANSFORM_FEEDBACK_COUNTER_READ_BIT
tu: Fix blit_cache_cleaned never being set to true
tu: Fix LRZ handling for VK_EXT_custom_resolve
tu: Dirty LRZ after changing attachment locations disable LRZ writes
nir: Include scalarized component offsets in UBO ranges
tu: Fix tu_event not being reset on creation
tu: Merge disable_write_for_rp from secondary to primary
tu: Fix memory leak of FDM patch-points
Dave Airlie (15):
nouveau: drop sector promotion.
gallivm: handle llvm 22 coroutine end change
gallivm: handle llvm 22 scatter/gather intrinsic changes.
lavapipe: treat NULL pColorAttachmentLocations as no handles
nak: fix image size for multisample arrays
nak: add more sizes to assert in bindless_image_sparse_load
nvk: enable subgroupQuadOperationsInAllStages
ci: vmware farm is offline, stop using it
st: fix get tex subimage fallback for 1D ARRAY
st: drop ununsed arguments to copy_to_staging_dest.
u_blitter: only set texcoord.w to sample for multisample sources
nak: block pipe_format from nak bindings.
ir3: use the correct builder for adding preamble to main.
nir: add impl pointer to block to avoid recursive linked list
nir: remove a lot of nir_cf_node_get_function calls.
David Airlie (3):
nir/coopmat: refactor the split vars to clean it up
nir/coopmat: move the row/col into a box and add some helpers.
nir/coopmat: rename the box split variables.
David Rosca (70):
d3d12: Use HEVC RefPicSet order from frontend
ac/parse_ib: Fix printing enc recon VAs on VCN5
radv: Fix uint32 overflow in slice offset calculation
radv/video: Fix initializing rc structs with default rate control
radeonsi: Always use 2D tiling for video dpb
frontends/va: Fix finding LTRs from POCs in HEVC decode
frontends/va: Fix out of bounds write in AV1 decode tile info
frontends/va: Fix setting output color properties from color standard
frontends/va: Fix dereference before NULL check in postproc
frontends/va: Add missing NULL check for additional output surface
vl: Use NV12 as deint format instead of preferred format
vl: Don’t check npot textures support when creating buffers
pipe/video: Remove unused PIPE_VIDEO_CAP_PREFERRED_FORMAT
pipe/video: Remove unused PIPE_VIDEO_CAP_NPOT_TEXTURES
pipe/video: Remove unused PIPE_VIDEO_CAP_MAX_LEVEL
pipe/video: Remove unused PIPE_VIDEO_CAP_STACKED_FRAMES
radeonsi/uvd_enc: Skip extra padding bytes in output bitstream
radeonsi: Move si_vpe.* to mm subfolder
ac/info: Add video codec caps
radeonsi/video: Use new video codec caps
radv/video: Use new video codec caps
ac/info: Remove old video codec caps
ac/info: Print number of VPE instances
radeonsi: Add RADEON_FLUSH_FORCE and use it to force flush
ac/vcn_dec: Add ac_vcn_dec_init_regs to get register offsets
ac/vcn: Add ac_vcn_sq_header/tail and use it for decode
ac/cmdbuf: Add ac_emit_video_write_memory
radv: Use ac_emit_video_write_memory
ac/vcn_dec: Move register defines to ac_vcn_dec.c
ac/cmdbuf: Add ac_emit_video_write_timestamp
radv: Add support for timestamps on video queue
radeonsi/mm: Add support for 2-ref H264 encode
radeonsi/mm: Remove comment about kernel AV1 instance scheduling bug
ac/parse_ib: Add VCN decode queue parsing
ac/parse_ib: Add VCN timestamp command
radeonsi/mm: Add si_vid_create_buffer and use it
radeonsi/mm: Set PIPE_RESOURCE_FLAG_UNMAPPABLE for buffers
va: Set contiguous_planes for DMA-BUF imported surfaces
ac/vcn_dec: Add 10 to 8 bit dithering support
radeonsi/mm: Select DPB format independently from decode surface format
vl: Skip transfer function and primaries conversion when not needed
va: Always reset compositor chroma location
va: Use RGB format with matching bit depth for YUV->YUV matrices
radeonsi/mm: Return error when decoding H264 P/B frame with no refs
radeonsi/mm: Only setup ref surfaces with tier3
radeonsi/mm: Set correct usage in si_dec_fill_surface
radeonsi/mm: Fix setting VPE rotation when horizontal flip is enabled
va: Implement vaPutImage for derived images
pipe/video: Add out_pipe_fence to pipe_picture_desc
vl: Support blending with gfx compositor
vl: Add pipe_video_codec proc using vl_compositor
vl: Add vl_proc pipe_video_codec using vl_compositor as fallback
va: Use vl_proc for processing context
va: Add vlVaDestroySurface
va: Add vlVaPostProc and use it instead of compositor and vid engine blit
va: Use vlVaPostProc in vlVaPutSurface and for subpictures
va: Stop using vl_compositor
pipe: Remove pipe_video_codec::expect_chunked_decode
r600/uvd: Set correct h264 chroma format
pipe: Remove pipe_video_codec::chroma_format
vulkan/video: Fix coding AV1 decoder/encoder_buffer_delay
vulkan/video: Don’t code AV1 decoder model info when not present
vulkan/video: Fix coding AV1 operating points
d3d12/video: Don’t reset batches in fence_wait
vulkan/video: Fix coding H265 ref pic list modification lists
vulkan/video: Fix coding H265 SPS pcm block sizes, inter ref pic set and lt refs
va: Fix leak when vlVaUploadImage fails
va: Ensure templat is valid for temporary surfaces
radeonsi: Stop forcing GTT with no gfx/compute
ac/video: Fix number of AV1 single refs
Derek Lesho (2):
zink: Guard bo map/unmap on map_count.
zink: Fix zink_bo_unmap synchronization for client pointer support.
Dhruv Mark Collins (7):
tu/autotune: Fail gracefully when CP counters are unavailable
fd/pps: Allocate performance counters from high-to-low
tu/autotune: Allocate performance counters from low-to-high
tu/query_pool: Avoid CP counter conflict with autotune
freedreno: Update A6XX_PC_MODE_CNTL definition and values
tu/util: Fix tile division algorithm
tu: Propagate allocation failures for tu_cs_* functions
Dmitry Baryshkov (3):
rusticl: enable freedreno by default
tu: limit KHR_internally_synchronized_queues to Vulkan 1.1+
tu: limit VALVE_fragment_density_map_layered to Vulkan 1.1 devices
Dmitry Osipenko (3):
intel/virtio: Preserve errno properly when handling ioctl
drm-uapi: Update virtio-gpu with new hinting field
intel/virtio: Support DRM_VIRTGPU_BLOB_FLAG_HINT_DEFER_MAPPING
Dorinda Bassey (1):
util/rust: Add atomic memory synchronization support
Duncan Brawley (7):
pco: Fix pco_last_igrp returning the first element instead of the last
pco: Refactor internal shader pass skipping
pco: Add propagating coherent/volatile access qualifiers
pco: Add DMA ld/st caching support for ssbo/ubo operations
pco: Add DMA sampling caching support
pco: Fix smp instruction encoding map order
pco: Add DMA ld/st caching support for all ld/st instructions
Dylan Baker (3):
intel/brw: Add assert for error case
meson: ensure that libdrm auto-features match requirements
intel/gen: decode type of src1 in basic 2 source after setting IMM
Emma Anholt (87):
spirv: Demote the SPIRV 1.6 OpTypeSampledImage on Buffer failure to a warning.
ci: Bump apitrace version to 14.0.
ci: Don’t set wine vars in deqp-runner.sh/vkd3d-runner.sh.
ci: Build a working wine installation in build-wine.sh.
ci/test-vk: Install win64 apitrace 14.0 along with setting up wine.
ci/test-vk: Install DXVK 2.7.1 to our wine installation.
ci/lava: Fix the name of the fluster overlay.
ci/lava: Add a note about an otherwise-mysterious error you can encounter.
ci/gfxreconstruct: Disable OpenXR support.
ci: Build gpu-trace-perf and include a script to use it.
ci: Bump the image tags for the previous build script changes.
bin/update-traces_checksum.py: Pull out per-job work to a helper function.
ci/update_traces_checksum: Parse gpu-trace-perf’s format for hash changes.
ci/update_traces_checksum: Default to updating for the current HEAD.
ci/update_traces_checksum: Make it work on restricted traces jobs, too.
ci/llvmpipe: Use anholt’s new GPU trace snapshot comparison tool.
ci/lavapipe: Use anholt’s new GPU trace snapshot comparison tool.
ci/turnip: Drop two 660 vk jobs and tune down the vk coverage fraction.
ci/turnip: add an a660 VK restricted traces job.
ci: Delete references to various broken traces.
ci/amd: Switch radv-raven-traces-restricted over to gpu-trace-replay.sh
ci/intel: Switch over to the new tool for restricted traces.
ci/piglit-traces: Remove ANGLE trace support.
ci/llvmpipe: Disable some traces too close to the timeout.
tu: Set HALF_PRECISION on blits to R11G11B10.
ir3: Fix shared IMAD24 lowering.
tu: Add capture/replay for sparse buffers and descriptor buffer.
screenshot-layer: Fix leftover VK queues in the map at DeviceDestroy.
screenshot-layer: Fix a bunch of unused variable warnings.
screenshot-layer: Fix rename() to final png before the file is flushed.
screenshot-layer: Clean up the lifetime management of the copyDone fence.
screenshot-layer: Fix race on writing the .pngs vs device destroy.
screenshot-layer: Wait on the fence before fallible operations.
tu/ci: Drop some old xfails that don’t trigger any more.
tu: Report missing layout support for host_image_copy with unifiedLayouts.
zink/ci/tu: Fix up skips/xfails for GLCTS testcases that got divided up.
tu: Disable storage image support for depth/stencil.
ci/panfrost: Drop a set of flakes whose fix had landed.
panfrost/ci: Skip dEQP-VK.wsi.wayland.swapchain.render.10swapchains on g52.
lvp/ci: Drop an old skip long since fixed in the CTS.
tu/ci: Drop a750 VKCTS to 50% coverage.
ci: Update VK CTS to 1.4.5.3 with fixes.
ir3: Add an env var to prefer single wavesize.
ir3: Fix shader bisect crashing out when too many shaders get bisected.
ir3: Give some feedback as we shader bisect.
ir3/shader_bisect: Allow a ‘r’ response to retry a run mid-bisect.
ir3: Drop the “SIMD0” debug print that was apparently added for frameretrace.
ir3: Deduplicate shader disassembly generation.
ir3: If we’re dumping IR3_SHADER_BISECT=[hash] disasm, include the NIR.
tu: Disable 128-wide subgroups on No Man’s Sky.
screenshot-layer: Log when we can’t open the output directory.
weston: Run at a more reasonable 1920x1080 resolution, not 1024x640.
ci: Include Windows renderdoc in with wine and update gpu-trace-perf.
ci: Build the Vulkan screenshot layer as part of VK test builds.
freedreno/ci: Add restricted traces testing of D3D11 traces on a660.
freedreno/ci: Add more explanation of a trace failure that’s not our fault.
tu: Always set the kernel’s name for BOs.
util/drirc_gen: Add a little documentation of what this does.
radv/drirc_gen: Clean up the dependency handling.
util/drirc_gen: Reduce manual importing of functions.
util/drirc_gen: Move the common VK WSI options to a core helper function.
util/drirc_gen: Make the header usable from C++.
tu: Move to using drirc_gen.
drm-shim: Include the hex of the driver ioctl for unimplemented ioctls.
drm-shim/freedreno: Provide a dummy set of UBWC config params.
drm-shim/freedreno: report a 48-bit address space.
drm-shim/freedreno: Report VM_BIND support.
vulkan: Enable GOOGLE_display_timing on KHR_display across multiple drivers.
drm-shim/freedreno: Fix VM_BIND support.
docs: Link in particular to the difficulty: * issue tags in Help Wanted.
zink: Use the new common code for nearest consistency in blits.
zink: Also enable the nearest consistency workaround on turnip.
zink: Also enable the nearest consistency workaround on anv.
.mailmap: Switch to anholt’s current work address.
intel/device_info_override_test: Make sure we actually find our device.
drm-shim: Share common code for PCI and platform device setup.
drm-shim: Generalize overriding of links.
drm-shim: Give the device/subsystem links real link values.
drm-shim: Lock access to shim_device.fd_map.
drm-shim: Remove drm_shim_driver_prefers_first_render_node.
drm-shim: Remove unnecessary runtime setup of drm_device_path_prefix.
drm-shim: Remove unnecessary runtime setup of various device strings.
drm-shim: Fix racy initialization.
freedreno/ci: Clear the xfail for texture-immutable-levels.
etnaviv/ci: Fix flakes lists that are breaking gc2000 CI.
intel/ci: Fix xfails for nightlies.
freedreno: Don’t force image component A=1 substitution on R/RG textures.
Emre Cecanpunar (1):
jay: allocate shader under memctx
Eric Engestrom (131):
VERSION: bump to 26.2
docs: reset new_features.txt
docs: update calendar for 26.1.0-rc1
docs: update calendar for 26.0.5
docs: add release notes for 26.0.5
docs: add sha sum for 26.0.5
docs: add stub of vk_struct_type_cast.h for vk_util.h
ci/bare-metal: drop duplicate timestamps now that gitlab-runner has per-line timestamps
docs: update calendar for 26.1.0-rc2
docs: update calendar for 26.1.0-rc3
docs: update calendar for 26.0.6
docs: add release notes for 26.0.6
docs: add sha sum for 26.0.6
docs: update calendar for 26.1.0
docs: add release notes for 26.1.0
docs: add sha sum for 26.1.0
docs: add calendar for the 26.1 cycle, and 26.2 branchpoint and release candidates
docs: fix unescaped `*`
docs/submittingpatches: fix section nesting
docs/ci: explain what Marge saying “Manual Step encountered” means
zink+nvk/ci: update expected fails
docs: update calendar for 26.0.7
docs: add release notes for 26.0.7
docs: add sha sum for 26.0.7
docs/ci: ignore docs.redhat.com & registry.khronos.org links
etnaviv: initialize value before calling etna_gpu_get_param(), in case it fails
meson/libmesa: ensure shader_replacement.h is generated before using it
meson/amd: only build libaco when requested
meson/asahi: only build libagx2_disasm when requested
meson/freedreno: only build libfreedreno_common when requested
meson/intel: only build libblorp_elk when requested
ci/build: restore riscv64 build as it works again
Revert “ci/build: restore riscv64 build as it works again”
docs: update calendar for 26.1.1
docs: add release notes for 26.1.1
docs: add sha sum for 26.1.1
docs: update calendar for 26.0.8
docs: add release notes for 26.0.8
docs: add sha sum for 26.0.8
util/meson: simplify list of per-driver drirc files
drirc: move 00-$drv-defaults.conf to each driver’s folder
Revert “drirc: move 00-$drv-defaults.conf to each driver’s folder”
docs: update calendar for 26.1.2
docs: add release notes for 26.1.2
docs: add sha sum for 26.1.2
rusticl: skip bindgen for pipe_shader_state_from_tgsi
meson: exclude known buggy versions of bindgen
ci: bump rust version from 1.90 to 1.96
ci: bump bindgen version from 0.71.1 to 0.72.1
ci: bump fedora from 42 to 44
meson: drop non-existent platforms=xcb check
Revert “egl: fix _EGL_NATIVE_PLATFORM fallback for unrecognized native displays”
docs: update calendar for 26.1.3
docs: add release notes for 26.1.3
docs: add sha sum for 26.1.3
gen_release_notes_test: don’t evaluate backslash
gen_release_notes: add support for “work_items” links
docs: fix release notes for 26.1.0
docs: fix release notes for 26.1.1
docs: fix release notes for 26.1.2
docs: fix release notes for 26.1.3
ci: fix perfetto download in `make-git-archive` nightly job
ci: fix perfetto download in build-perfetto.sh
ci: fix the fix for perfetto download in `make-git-archive` nightly job
etnaviv/ci: document two fixed tests
nvk/ci: document fixed tests, new failures, and recent flakes
zink+nvk/ci: fix duplicate fails
docs: drop x.org -> x.org/wiki/ redirect and expected url
docs: s/issues/work_items/
docs/ci: mark yet another domain as blocking linkcheck
docs/ci: disable auto-retry on nightly linkcheck
docs/ci: use full/explicit option names in linkcheck job
docs/ci: only print the linkcheck issues, not the thousands of non-issues
mr-label-maker: add ~drirc label on all drirc files
util: add support for multiple colon-separated DRIRC_CONFIGDIR entries
drirc: move 00-$drv-defaults.conf to each driver’s folder
meson: merge two consecutive `if with_egl`
meson: add native platform to the summary
meson: ensure native platform is one of the undetectable ones
meson: drop misleading `-D egl-native-platform` values
zink/ci: drop leftover anv-cml deqp suite
docs: update calendar for 26.1.4
docs: add release notes for 26.1.4
docs: add sha sum for 26.1.4
docs: fix x.org url
etnaviv/ci: update nightly job expectations
zink+nvk/ci: update nightly job expectations
lvp/ci: update nightly job expectations
llvmpipe/ci: update nightly job expectations
nvk/ci: update nightly job expectations
drm-shim: name the driver name `driver_name` consistently
drm-shim: set `driver_name` in `drm_shim_*_device_setup()`
docs: gitignore the contents of the `_generated` folder
ci: disable auto-retry on rustfmt job
docs/helpwanted: url-encode `[]` to avoid a pointless redirection
rusticl: document api@clgetmemobjectinfo as fixed for all drivers
lavapipe/ci: document fixed dEQP-VK.mesh_shader.ext.misc.emit_in_control_flow_bad_emit_last
freedreno/ci: document fixed KHR-GL46.copy_image.smoke_test
freedreno/ci: document a recent flake
zink+nvk/ci: document a couple of recent flakes
ci/piglit: fix nightly expectations after piglit uprev
img/ci: add `farm:imagination` tag to all jobs
loader: move variable to correct scope
broadcom/ci: mark fixed tests as such
ci/video: move two single-thread tests to global list
ci/video: install the current version of gstreamer
ci/video: uprev fluster
ci/video: download fluster test suites by codec name
ci/video: update comment with the new blocker for AV1 support
radv/ci: enable VP9 testing in fluster
anv/ci: enable VP9 testing in fluster
nvk/ci: document fixed dEQP-VK test
nvk/ci: document fixed vkd3d tests
nvk/ci: document two vkd3d regressions
freedreno/ci: document fixed tests
radeonsi: fix truncated cache key
mailmap: update my email address
zink+nvk/ci: document two fixed tests
VERSION: bump for 26.2.0-rc1
.pick_status.json: Update to d49a00bdf15fd48b31af93aaf5feed3eebcbda12
VERSION: bump for 26.2.0-rc2
.pick_status.json: Update to 8b00adbe72f2705985146b057f6fde9256d0dcb0
.pick_status.json: Mark 47efd739121d51e2f9049cec715a70c13767a67c as denominated
.pick_status.json: Mark 7999060992e9cee91d1962faf65dc4e5c6fe4f69 as denominated
.pick_status.json: Mark 2515024a5919ed14fe05471e3f1f89c54a454610 as denominated
.pick_status.json: Mark 4fd93a0039a07ec2027f2a6d1d252c73ed033e0a as denominated
pick-ui: turn commit.date into a (cached) property
pick-ui: show MR number for additional context
VERSION: bump for 26.2.0-rc3
.pick_status.json: Update to 85c082ddbed727940535911e6bf87f7d274525bf
[26.2 only] docs/new_features: mention that VK_EXT_host_image_copy was exposed on RADV/GFX10.3+
Eric Guo (3):
compiler: Add missing MESA_SHADER_KERNEL case for SPIR-V dump
pan/compiler: Clamp fp16 ldexp exponent range
pan/bi: Lower 64-bit hadd on v9/v10
Eric R. Smith (4):
glsl, spirv: Improve accuracy of asin() and acos()
panfrost: add some sanity checks
panfrost: make sure INDEX_OFFSET is cleared
panfrost: add helper function for checking for active queries
Erico Nunes (4):
ci: lima farm maintenance
Revert “ci: lima farm maintenance”
CODEOWNERS: add lima maintainers
ci: lima farm maintenance
Erik Faye-Lund (88):
panvk: drop out-of-date TODO
panfrost: use perf-trilinear when doing anisotropic sampling
panvk: use perf-trilinear when doing anisotropic sampling
pan/lib: fix up afbc and linear layout
pan/lib: emit high bits of buffer-size
pan/lib: validate data_size_B in drivers
panvk: do not artificially limit image dimensions
panvk: increase maxResourceSize on v11 and later
panvk: increase maxBufferSize on v11 and later
nouveau: do not report unsupported feature
radeonsi: remove old, unsupported cap
d3d12: remove benign but unsupported cap
iris,crocus: remove benign but unsupported cap
llvmpipe: drop support for tgsi_tex_txf_lz cap
ntt: stop emitting TXF_LZ
gallium/u_blitter: stop emitting TEX_LZ
gallium: remove defunct pipe-cap
ttn: do not handle T{EX,XF}_LZ
gallium: completely remove T{EX,XF}_LZ opcode
panvk: do not enable extension without required feature
panvk: do not enable extension without required feature
haiku: remove unfinished post-processing support
gallium: delete leftovers of post-processing infrastructure
pan/ci: add a flake from nightly
util/format: make Y8_UNORM an alias of Y8_400_UNORM
util/format: make subsampling explicit
util/format: mark subsampled RGB formats as actually subsampled
util/format: verify subsampling in name
pan/va: do not allow force_delta_enable on v9
pan/bi: correct computation of lod.x
panfrost: enable ARB_texture_query_lod on v9+
mesa/main: remove stale prototypes
mesa/main: remove incorrect debug-output
mesa/main: do not gate performance warning
mesa/main: remove low-value debug-output
mesa/main: remove unused verbose-flags
mesa/main: remove VERBOSE_API
mesa/main: remove mesa_print_display_list function
mesa/main: remove low-value verbose-switch
Revert “mesa: check for ARB_ES3_compatibility in format checks”
mesa/main: remove unused array
pan/ci: update flakes based on nightly ci
pan/ci: remove benign typoed flake
meson: update libdrm wrap
pan/ci: add missing gitlab rules
pan/ci: remove outdated gitlab rule
pan/ci: add missing gitlab rule
pan/ci: fix gitlab rules after move
pan/genxml: correct size of field
pan/genxml: add missing modifier
pan/genxml: correct size of field
pan/genxml: correct size of field
pan/genxml: correct size of field
pan/genxml: add missing enum value
pan/genxml: sort CS structs by enum-value
pan/genxml: use consistent name for scissor
pan/genxml: use an enum for progress increment
pan/genxml: consistently use bool for error reject
pan/genxml: consistently use hex for masks
pan/genxml: consistently use uint for signal slot
pan/genxml: consistently use uint for chunk indexes
pan/genxml: remove needless defaults
pan/genxml: keep enum ordering from v10
pan/genxml: correct casing of names/types
pan/genxml: consistently use hex for uint immediates
pan/genxml: consistently set default
pan/genxml: clean up whitespace
pan/genxml: make field consistent
pan/genxml: remove some pointless comments
pan/genxml: use consistent attribute order
pan/ci: add a couple of flakes
pan/ci: use slow-skips to only skip slow tests for merge-requests
pan/ci: move cts-bug-fails to skips
pan/ci: stuff some breadcrumbs in the fails-list
pan/ci: reenable passing tests
pan/ci: add a few new g925 flakes
pan/ci: add back missing skip-list heading
pan/ci: drop needless skips
pan/ci: move skip to flakes
pan/ci: skip slow test
pan/ci: move common flake to common flake-file
pan/ci: recognize flaking test
pan/ci: mark missing xfails
pan/ci: move longprim flake into common flake-file
pan/ci: add new flake
pan/ci: just mark all random-max draw-tests as flakes
ci/vulkan: remove long outdated skips
panvk: simplify non_polygon calculation
Etaash Mathamsetty (4):
vulkan/wsi/wayland: Fix error handling for tearing control.
vulkan/wsi/wayland: Move drm syncobj to swapchain.
vulkan/wsi/wayland: Move color management surface to swapchain.
vulkan/wsi/wayland: Do a roundtrip after retiring the old swapchain.
Faith Ekstrand (399):
panvk/csf: Emit INDEX_BUFFER[_SIZE] even for non-indexed draws
pan/bi: Improve swizzle propagation
zink: Assert if we try to use a dedicated allocation with offset > 0
panfrost: Add and use a new pan_nir_res_handle() helper
pan,nir: Add cube face intrinsics
nir/builder: Allow backend1/2 in nir_build_tex()
nir: Add a new nir_texop_gradient_pan
panvk: Implement bitfield_select
pan/nir: Add a pass for lowering texture ops in NIR on Valhall+
pan/nir: Use the NIR lowering on Valhall+
nir: Add a new nir_op_f2u32_rtne
pan/bi: Implement nir_op_f2[iu]32_rtne
pan,nir: Add Bifrost texturing intrinsics
pan/nir: Add bifrost support to pan_nir_lower_tex()
pan/nir: Lower texturing ops in NIR on Bifrost
pan/nir: Load texel buffer conversion descriptors in NIR
pan/bi: Allow setting the table on lea_attr_pan
pan/nir: Use HW NIR intrinsics for texel buffer addresses
pan/bi: Delete the old texel buffer intrinsics
pan/nir: Lower texel buffers in nir_lower_tex()
pan/nir: Lower texture queries in nir_lower_tex() on Valhall+
panfrost: Also remap image handles for image_size/samples
pan/nir: Lower image queries in NIR on Valhall+
panvk: Let the compiler handle texture queries on v9+
pan/nir/tex: Support full index+offset
panvk: Add MAX_VS_ATTRIBS to image indices in panvk_nir_lower_descriptors
panfrost: Take texture/sampler_index into account in lower_res_indices
panfrost: Prefix valhall bits of lower_res_indices
panfrost: Handle pre-Valhall images and texel buffers in lower_res_indices
pan/bi: Drop lower_index_to_offset from preprocess
util/half: Use explicit RTNE rounding for the C++ float16_t
util/half: Stop whacking CPU flags to test float_to_half_slow()
util/half: Rename the tests
util/half: Re-organize the tests a bit
util/half: Add float_to_half rounding tests
util/half: Add double_to_half tests
util/half: Add a simpler double_to_float16()
util/half: Add double_to_float16_ru/rd helpers
nak: Move Srcs/DstsAsSlice implementations
nak: Implement Srcs/DstsAsType directly for Op
nak: Implement Srcs/DstsAsSlice directly on ops
nak,compiler: Move AttrList into NAK
nak: Don’t use the proc macro to implement auto-boxing of ops
nak,compiler: Move FromVariants to common code
pan/bi: Use LOD_MODE_EXPLICIT for the 2nd half of textureGrad() on Bifrost
docs: Move and rename “Development Notes”
docs: Add docs with Vulkan/SPIR-V extensions basics
docs: Add docs for drafting new MESA extensions
panvk/csf: fix VERTEX_SPD dirty tracking when topology changes
panvk/csf: Inline the SPD addr helpers
nouveau/push: Rename push_method to push_mthd
nouveau: Don’t build NAK tests on Android
compiler/rust: Add a float16 wrapper
etnaviv: Remove f32_to_f16_fallback() in favor of float16::F16
meson: Bump the minimum rust version to 1.85.0
compiler/rust: Add LowerBoundedU32[Array] types
nak: Use LowerBoundedU32 for SSAValue
nak: Allow SSA value 0 again
nak: Simplify SSARef construction with try_push()
meson: Suffix compiler/rust bindings with _compiler_rs_extern
compiler/rust/bindings: Add util_dyarray
compiler/rust: Add a nir_shader::get_entrypoint() helper
compiler/rust: Add a nir_shader::to_string()
compiler/rust/nir: Add structured block iterators
compiler/rust/nir: Add helpers for getting ALU input/output types
compiler/rust/bitset: Add a BitIndex helper struct
compiler/rust/bitset: Don’t reserve space in remove()
compiler/rust/bitset: Add find_next_[un]set() helpers
compiler/rust/bitset: Generalize BitSetIterator
compiler/rust/bitset: Implement Into/FromBitIndex for more types
compiler/rust/bitset: Add a new ConstBitSet type
compiler/rust: Add an EnumAsU8 trait
nak: Use EnumAsU8 for RegFile
panfrost: Initial rust build system support
panfrost: Add the basis for the new Kraid compiler
kraid: Add a GPU model abstraction
kraid: Add DataType and NumericType enums
kraid: Add a swizzle struct
kraid: Add SSAValue and SSARef structs
kraid: Add Src/Dst data types
kraid: Add an Opcode trait and Op enum
kraid: Add Instr, BasicBlock, and Shader structs
kraid: Add a builder
kraid: Start parsing NIR shaders
kraid: Parse the NIR CFG
kraid: Handle load_const instructions
kraid: Handle nir_op_mov/vec/[un]pack
kraid: Add some float alu ops
kraid: Implement nir_op_iadd
kraid: Handle a few NIR intrinsics
Kraid: re-indent shaders for prettier printing
kraid: Add a validator to check IR invariants
kraid: Add a super simple register allocator
kraid: Plumb through Model::encode_shader()
kraid: Rework swizzles
kraid: Print ASM swizzles when we have them
kraid: Copy the bitview module from nouveau
kraid: Add a FlowCtrl struct
kraid: Replace OpEnd with OpNop.end
kraid: Move proc/lib.rs to proc/macros.rs
subprojects: Pull in the Rust xml crate
kraid: Add ISA XML for v9-15
kraid: Add the start of encoder code-gen
kraid/isa: Add a simple XML parser
kraid/isa: Generate enums with [Try]Encode/Decode
kraid/isa: Add an encoder for expressiosn
kraid/isa: Add an encoder for instructions
kraid/isa: Add support for field modifiers
kraid: Add the start of a v9 encoder
kraid/isa: Specially handle small_constant_t
kraid: Add a SmallConstant struct and a Model::small_constants() hook
kraid: Add a lower_small_constants() pass
kraid: Add a very dumb message slot assignment pass
kraid: Implement shifts and logic ops
kraid: Implement integer comparisons
krai/isa: Expose a new InstructionInfo struct per-instruction
kraid: Break v9 instruction encoding out into traits
kraid: Use instruction info to implement op_is_message()
kraid/isa: Add a special case in to_snake/camel_case() for data types
kraid/isa: Emit TryFrom<DataType> for all data-type-like enums
kraid: Clean up the data type mess in the encoder
kraid: Implement OpCSel and nir_op_[ui]min/max
kraid: Claim we use 64 registers
kraid: Support signless IAdd
kraid: Implement nir_op_u2u/i2i
kraid: Add a SrcRef::Zero
kraid: Add a 16-bit ALU lowering pass
kraid: Implement nir_op_extract_*
kraid: Be more lax about immediates
kraid: Map H01 and B0123 to None in the encoder
kraid: Implement nir_op_f2f*
kraid: Make Instruction::get_info() more ergonamic
kraid: Add a Model::op_src_supports_imm32() query
kraid/isa: Handle field restrictions
kraid: Box ops inside Op
compiler/rust/smallvec: Implement Clone, Default, and new()
compiler/rust/smallvec: Add a push_mut() method
compiler/rust/smallvec: Implement Deref[Mut]<Target = [T]>
compiler/rust/smallvec: Implement Extend<T> for SmallVec<T>
compiler/rust/smallvec: Implement From<Vec<T>>
compiler/rust/smallvec: Implement FromIterator and From<[T; N]>
compiler/rust/smallvec: Implement IntoIterator
compiler/rust/smallvec: Implement From<SmallVec<T>> for Vec<T>
nak: Simplify BasicBlock::map_instrs()
nak/builder: Use some of the SmallVec improvements
nak: Simplify our SmallVec usage
compiler/rust/smallvec: Hide the enum
compiler/rust/smallvec: Optimize extend()
nir: Allow atomic intrinsics to have multiple components
spirv,nir: Add support for AtomicFloat16VectorNV
nak/nir: Lower f16vec4 atomics to 2xf16v2
nak: Rename AtomType::F16x2 to F16v2
nak/from_nir: Handle f16v2 atomics
nvk: Advertise VK_NV_shader_atomic_float16_vector
kraid: Make SrcRef::Imm32 explicitly non-zero
kraid: Make SrcRef PartialEq
kraid: Add map_instrs() methods to Shader and BasicBlock
kraid/builder: Store the model in builders
kraid: Split DataType into two enums
kraid/v9: Fix encoding of high register numbers
kraid/v9: Rework the shift_lop encode macro
kraid/v9: Add the rest of the shift/lop ops
kraid: Add None logic and shift ops
kraid/v9: Allow immediates in logic ops
compiler/rust/bitset: Implement Eq and PartialEq for ConstBitSet
compiler/rust/enum_as_u8: Add an EnumAsU8::MAX_DISCRIMINANT
compiler/rust/enum_as_u8: Add an ConstU8EnumSet struct
compiler/rust/as_slice: Document AsSlice
compiler/rust/as_slice: Add a new AsArray trait
kraid: Add a VirtualOpcode trait
kraid: Add a Model::op_is_supported() query
kraid: Add a lanes to Dst
kraid/isa: Make Enum::meta a weak reference
kraid/isa: Rework enum literals
kraid/isa: Make Swizzle EnumAsU8
kraid/isa: Treat exact= as a field restriction
kraid/isa: Expose allowed swizzles through InstructionInfo
kraid/isa: Expose allowed lanes through InstructionInfo
kraid: Add a Model::op_src_supports_swizzle() helper
kraid: Add a Model::op_dst_supports_lanes() helper
kraid: Add the hardware MkVec ops
kraid/v9: Fix OpShiftLop::src_supports_imm32()
kraid/v9: Fold swizzles and modifiers on imm1w sources
kraid: Add a virtual OpCopy and the relevant lowering pass
kraid/nir: Emit OpCopy instead of OpMov
kraid: Allow 8-bit SSA values
kraid: RA per-byte
kraid/ops: Claim even more variants
kraid/nir: Emit 8-bit ops
kraid: Widen ALU ops before RA
kraid: Expose the guts of Swizzle
kraid/builder: Add copy_iN_to() helpers
kraid: Add an OpSwz and a lower_mkvec_swz() pas
kraid/nir: Use OpSwz for nir_op_u2uN and nir_op_i2iN
kraid/nir: Fix 2x16 extract_[iu]8
kraid/nir: Use OpSwz op_extract_*
kraid/nir: Implement nir_op_unpack_32_*
kraid: Add a new legalize_src_swizzles() pass
kraid: Add word() helpers to Src/Dst types
kraid/nir: Implement nir_op_unpack_64_*
kraid: Better document swizzles
kraid: Only dump shaders if KRAID_DEBUG=print is set
kraid: Fix RA for dead destinations
kraid: Add a Model::op_src_is_staging_reg() helper
kraid: Add a Model::op_dst_is_staging_reg() helper
kraid: Allocate whole registers for staging destinations
kraid: Re-materialize constants
panfrost: Set the rustfmt edition to 2024
kraid/swizzle: Add a Swizzle::is_none() helper
kraid/swizzle: Add an is_none() special case in fold_u32()
kraid/swizzle: Take a src_bytes param in Swizzle::bytes_read()
kraid/validate: Fix 64-bit destination validation
kraid/hw_tests: Allow the test to specify swizzles and lanes
kraid/swizzle: Return Option<Swizzle> from AsmSwizzleWiden::to_swizzle()
kraid: OpShiftLop is unsigned
kraid: Add an SSAValue::bytes() helper
kraid: Use a tuple struct for SSAValue
kraid: Add OpRegIn and OpRegOut
panvk/jm: De-duplicate most of cmd_draw[_indirect]
panvk/jm: Re-group setting desc tables and SSBOs
panvk/jm: Take a desc_info in meta_get_copy_desc_job
panvk/jm: Take a desc_info in prepare_desc/dyn_ssbo()
panvk/csf: Take a desc_info in fill_dyn_bufs() and prepare_res_table()
panvk: Move desc_info to panvk_shader
panvk: Call panvk_lower_nir() before lowering multiview
panvk: Add a central panvk_cmd_draw() helper
panvk/csf: Make various panvk_draw_info pointers const
panvk: Plumb index buffers through panvk_draw_info
panvk/jm: Plumb IA state through draw_info
panvk/csf: Plumb IA state through draw_info
panvk: Improve base instance tracking for indirect draws
panvk/csf: Add some sanity assertions in prepare_push_uniforms
panvk/csf: Break FS descriptor setup into a new helper
panvk/csf: Break VS descriptor setup into a new helper
panvk/csf: Prepare descriptors first
panvk: Patch VS attribute descriptors as a separate step
panvk: Improve panvk_shader_foreach_variant()
panvk: Add a helper for uploading to cmd mem
panvk/csf: Add a helper for dispatching compute shaders with 3D state
compiler/rust: Re-add From<Box<T>> to FromVariants
compiler/rust: Only allow FromVariants on enums
kraid: Make PAN_USE_KRAID per-stage
kraid: Use unsafe with no_mangle
kraid/v9: Simplify DstLanes logic for staging registers
kraid/ir: Rework some RegRange helpers
kraid: Automatically swizzle in From<SrcRef> for Src
kraid: Add lowering for COPY.i64
kraid: Add OpFMul and plumb it through
kraid/nir: Implement nir_op_inot
kraid/data_types: Add message types
kraid/data_types: Add unit tests
kraid: Add OpLea/LdTex and plumb them through
kraid: Add OpLd/StCvt and plumb them through
compiler/rust: Implement Eq/Hash/PartialEq for LowerBoundedU32Array
compiler/rust/bitset: Add an iteration test
compiler/rust/bitset: Further generalize find_next_set()
compiler/rust/bitset: Add a next_set() method
compiler/rust/bitset: Further generalize find_aligned_unset_range()
compiler/rust/bitset: Add a find_aligned_set_range() method
compiler/rust/bitset: Don’t write past the end in insert_range()
compiler/rust/bitset: Generalize ConstBitSet::insert_range()
compiler/rust/bitset: Add some range methods to BitSet<usize>
compiler/rust/bitset: Add an iter_bit_indices() method
compiler/rust: Add more methods/traits to U8EnumSet
kraid/nir: Implement load_local_invocation_id
kraid: Add OpMux and plumb it through
kraid/isa: Handle 16-bit replicated destinations
kraid: Add OpFrcp/Frsq and plumb them through
kraid/ir: Add a Opcode::set_variant() method
kraid: Widen more ops
kraid/hw_tests: Use a single basic block
kraid: Store blocks in a CFG
compiler/rust/bitset: Enable From/IntoBitSet for u32
kraid: Copy the SimpleLiveness and LiveSet from NAK
kraid: Add a parallel copy builder
kraid: Use Swizzle::is_none() more
kraid: Add new Phi label type and OpPhiSrc/Dst
kraid/nir: Handle nir_phi_instr
kraid: Implement EnumAsU8 for DstLanes
kraid: Rework supported DstLanes queries
kraid: Allow RegRef::word() on subregs
Revert “compiler/rust/bitset: Add an iter_bit_indices() method”
compiler/rust/bitset: Fix a unit test
compiler/rust/bitset: Add a count_set_in_range() method
compiler/rust/bitset: Implement Eq and PartialEq
compiler/rust/bitset: Add a retain() method
kraid: Better RA
kraid/nir: Use correct zero sizes for unused ALU components
kraid/swizzle: Enable Swizzle::swizzle() on word swizzles
kraid/ir,v9: Fix swizzles for the accum source of OpMkVecV2I8I16
kraid/ra: Re-swizzle 64-bit sources that read 32-bit values
kraid/ra: More accurately compute source constraints
kraid/nir: Allow i8v3 ops
kraid: Use a tuple struct for SSARef
kraid: Implement FromIterator for SSARef
kraid: Implement load_ubo
kraid/lower_copy: Use Src::imm_u8() for shifts
kraid: Use a Builder in ParallelCopy
kraid/parallel_copy: Emit small constants directly
kraid/ra: Delete a left-over debug check
kraid/nir: implement nir_op_[ui](add|sub)_sat
kraid/data_type: Add more auto types
kraid: Add a DataType::SR special case
kraid: Add OpLeaBuf and plumb it through
kraid: Add OpTex*
kraid/nir: Plumb through texture ops
kraid: Use flat_map() instead of map().flatten()
kraid/data_type: Handle SR in as_data_type()
kraid/v9: Actually encode OpTexGradient
kraid/v9: Fix src_supports_imm32() for Op[IF]Add
kraid: Add a vec src legalization pass
kraid/model: Add an op_src_supports_mod() query
kraid: Add a word-based copy propagation pass
kraid/widen: Don’t widen messages
kraid/nir: Enable load_global_constant
kraid/v9: Use the right data type for OpShiftLop::src_supports_imm32()
kraid: Add OpAtom* and plumb them through
kraid/v9: Don’t allow src0 swizzles in OpShiftLop::src_supports_imm32()
kraid: Run copy-prop after legalizing_src_swizzles()
kraid/copy-prop: Don’t propagate SSA values with mismatched sizes
kraid/swizzle: Expose the guts of swizzle composition
kraid/copy-prop: Add byte-based copy propagation
kraid: Pass the immediate to Model::op_src_supports_imm32()
kraid/v9: Support immediate buffer/texture handles
kraid/nir: Always use a destination for AtomOp::Xchg
kraid/ra: Handle OpPhiSrc with a swizzle
kraid: Call pan_shader_update_info()
kraid/nir: Respect FLOAT_CONTROLS_ROUNDING_MODE_RTZ
pan/nir: Lower read_invocation to 32 bits
compiler/rust/cfg: Assert that nodes are in a dominance-respecting order
compiler/rust/cfg: Unexpose CFG::from_blocks_edges()
compiler/rust/cfg: Make sorting optional in CFGBuilder::as_cfg()
kraid: Stop re-sorting blocks with CFGBuilder
kraid: Return an Option<RegRef> from Model::preload_reg()
kraid/nir: Use FAURef::user_i32()
kraid: Add special FAUs
kraid: Add OpBarrier and plumb it through
kraid/nir: Respect access flags on loads/store ops
kraid: Plumb TLS size through to pan_shader_info
kraid/nir: Implement load_scratch/shared_base_ptr
pan/nir: Lower scratch and shared to global for Kraid
kraid: Add OpWMask and plumb it through
kraid/nir: Implement load_subgroup_invocation
kraid: Add OpClper and plumb it through
kraid: Don’t report Src::is_zero() with a BNot modifier
kraid/copy-prop: Trivialize zero copies
kraid/ir: Rename the raw src/dst type helpers
kraid: Add a DataType::total_bytes() helper
kraid/validate: Fix source swizzle validation
kraid/ra: Fix W1 widens
kraid: Fix lower_small_constants() for 64-bit sources
kraid: Take a DataType in Opcode::is_valid_variant()
kraid/ra: Also handle OpPhi swizzles in the pre-existing live-out case
kraid/copy-prop: Try to re-type opcodes for more widening
kraid/copy-prop: Fold widen ops into 64-bit sources
kraid/copy-prop: Treat F16ToF32 as a widening copy
compiler/rust/bitset: Improve test_find_aligned_unset_range()
compiler/rust/bitset: Enhance find_aligned_[un]set_range()
nvk/image: Style nits
nvk/image: Rewrite nvk_image_can_compress() to use early returns
nvk/image: Take an nvk_physical_device in can_compress()
nvk: Add an NVK_DEBUG=no_compression flag
vulkan/meta: Use z_off/scale for 2D array images as well
vulkan/meta: Allow resolving a 2D MSAA image to a 3D image
kraid/ra: Fix find_unpinned_bytes() for unaligned ranges
kraid/ra: Relax alignment requirements for staging registers
kraid/nir: Rework mov/vec handling
kraid/nir: Implement nir_op_insert_*
kraid/nir: Implement as_uniform
kraid: Add a Src::fneg_zero() helper
kraid: Use FMA instead of FMUL
kraid/ir: Don’t compare labels in FAU/RegRef.eq()
kraid/nir: Add a special_fau() helper
kraid: Add a new FAUModel
kraid: Merge legalize_immmmediates and legalize_vec_srcs
compiler/rust: Add U8EnumSet::len() and ConstBitSet::len()
kraid/legalize: Add a move_src_to_tmp() helepr
kraid: Legalize FAU sources
kraid: Add OpIDpAdd and plumb it through
nvk: Replace nvk_addr_range with VkDeviceAddressRange
pan: Take a stage parameter to get_nir_shader_compiler_options()
pan: Move PAN_USE_KRAID into pan_compiler.c/h
kraid: Expose our own NIR compiler options
pan: Use Kraid’s NIR options when it’s enabled
kraid/isa,model: Add a op_srs_is_64bit() query
kraid: Align registers based on the new ISA query
kraid/copy-prop: Handle 64-bit OpShiftLop
kraid/nir: Implement 64-bit op_bitfield_select
kraid/nir: Enable more 64-bit ops
nir: Add combined shift-logic ops for panfrost
kraid: Use the new NIR shift+logic ops
kraid: Optimize shift+logic ops
kraid: Document a couple passes
kraid/nir: Implement nir_op_[iu]mul_2x32_64
nvk: Advertise minStorageBufferOffsetAlignment=4 for VKD3D
docs: Add a note about Vulkan implicit sync in the 25.3.0 release notes
compiler/rust/cfg: Remap node edges in remove_unreachable()
compiler/rust/nir: Implement Send+Sync for nir_shader_compiler_options
kraid: Use nir_shader_compiler_options directly
Feelthepain77 (1):
freedreno: add Adreno 613 (Snapdragon 4 Gen 2) to device list
Filip Gawin (5):
r300: avoid UB through implicit conversions on 32bit
r300: use uint32_t instead of long in vertprog
nv30: fix truncated values in line_stipple_pattern
nv30: fix 1 << 31 issues
nv30: fix another left shift cannot be represented in type ‘int’
Francisco Jerez (4):
nir/divergence: Allow local_invocation_id.z to be treated as uniform.
intel/brw: Sort scheduling modes by performance after initial RA failure.
intel/brw/swsb: Omit redundant read-after-read synchronization for back-to-back DPAS.
intel/brw: Add NIR pass to vectorize dot products into DPAS matrix multiplications.
Frank Binns (21):
pvr/ci: drop two tests from bxs-4-64-{fails,flakes}
pvr: re-enable {EXT,KHR}_index_type_uint8
pvr/ci: add AXE-1-16M nightly Vulkan CTS testing
pvr/ci: skip timing out VK reconvergence test for AXE-1-16M
pvr/ci: add some timing out tests on AXE-1-16M to skips list
pvr: drop unused struct member from pvr_render_pass_attachment
pvr: drop unused pvr_descriptor struct
pvr: enable KHR_external_semaphore{,_fd} unconditionally
pvr: define PVR_USE_WSI_PLATFORM for xcb and xlib
pvr: move PVR_USE_WSI_PLATFORM_DISPLAY into a header
pvr: advertise VK_EXT_display_surface_counter
pvr: advertise VK_EXT_display_control
pvr: advertise VK_EXT_direct_mode_display
pvr: advertise VK_{KHR,EXT}_surface_maintenance1
pvr: advertise VK_{KHR,EXT}_swapchain_maintenance1
pvr: advertise VK_EXT_swapchain_colorspace
pvr: advertise support for VK_EXT_acquire_drm_display
pvr: advertise VK_KHR_unified_image_layouts
pvr: rearrange some functions in pvr_arch_border.c
pvr: setup all format fields for custom border color entries
zink: gate some EXT_descriptor_indexing related code
Frank Bouwer (3):
pvr: Fix for depth stencil 2d array writes.
Revert “pvr: Fix for depth stencil 2d array writes.”
pvr: Fix for depth stencil 2d array writes.
Fyodor Kyslov (1):
mesa3d: gfxstream: Add P210 format support
GKraats (2):
hasvk: unbreak assert format != ISL_FORMAT_UNSUPPORTED
crocus: Fix shader precompilation on Gen6 and higher
Ganesh Belgur Ramachandra (8):
amd: import gfx11.7 addrlib
amd: add initial common code for gfx11.7
radeonsi: add gfx11.7
radv: add gfx11.7
amd: use gfx_level instead of family_id to choose addrlib
amd/llvm: fix target feature setting (DumpCode -> dumpcode)
amd/llvm: fix LLVM asserts for signed integer constants
amd/llvm: truncate const intergers to bitwidth
Georg Lehmann (118):
nir: remove nir_link_xfb_varyings
radv: allow input attachment to use pixel coord optimization
radv: move per-primitive fixup closer to radv_nir_lower_io
radv: move fs view_index handling after lowering io
radv: remove unused vs/tes num_outputs from shader info
radv: never call nir_assign_io_var_locations
radv: remove draw_id from mesh shader a bit later
radv: export multi view index as layer after lowering io
radv: remove radv_graphics_shaders_link
nir: disable fp class analysis for 64bit transcendentals
intel/nir_opt_peephole_ffma: fix fp_math_ctlr for modifiers
nir/instr_set: allow cse with fp_math_ctrl mismatches for intrinsics
nir/opt_varyings: back propagate signed zero information to outputs
nir/opt_varyings: do no_signed_zero linking even for non removable stores
nir/opt_algebraic: add more fmulz pattern
ac/nir/lower_tex_coord: fix moving wqm coordinates
nir: fix fp_math_ctrl in fisnan
nir/opt_peephole_select: do not count fmul towards the limit when only used by fadd
nir/loop_analyze: do not count fmul towards the limit when only used by fadd
nir,amd: reassociate fadd to create more fma/mad
radv/ci: update restricted trace checksums
radv: fix amount of sample shading with required sample shaded inputs
ac/nir/lower_tex_coords: fix optimizing cube txd to tex
aco: add tests for cube txd to tex opt
nir/opt_uniform_subgroup: preserve divergence during optimization
tgsi: delete unused lowering pass
aco/tests: use explicit lod in sparse texture test
spirv: always preserve infinities for FMin, FMax and FClamp
radv: use radv_get_sampled_image_desc_size instead of open coding it
radv: add radv_force_64_byte_sampled_image dri conf option
radv: enable radv_force_64_byte_sampled_image for Forza Horizon 6
aco/optimizer: only create v_fma_legacy_f32 when denorms are disabled
nir: seperate ffmaz from has_fmulz
ac/llvm: don’t assert on 32bit ffma before gfx9
ac/llvm: never create ffmaz for broken llvm
radv: support VK_KHR_shader_fma
aco/gfx8: fix 16bit nir_op_ffma
nir/deref: consider atomics that store derefs as complex use
aco/gfx6: fix fp64 floor lowering
radv: don’t lower dfloor in NIR
aco/gfx6: fix fceil lowering
aco/gfx6: use shorter lowering for ftrunc
aco: add rtne pseudo opcodes for fp64 add and fract
aco/gfx6: always use rtne for floor/ceil lowering
aco/gfx6: fix fround_even(-0.0)
aco/gfx6: always use rtne adds for fround_even lowering
radv: enable fp64 float controls on gfx6-7
aco/isel: never manually flush denorms after 32bit fma
nir: preserve infinities and signed zero during atan2
amd/gpu_info: precompute instruction prefetch distance
amd/common: don’t pass radeon_info to ac_align_shader_binary_for_prefetch
amd/common: add helper for INST_PREF_SIZE
radv: remove gfx6 code from ngg emission
radv/gfx11+: program INST_PREF_SIZE for compute
radv/gfx11+: program INST_PREF_SIZE for pixel shaders
aco: add exec_size to prolog/epilog callback
radv/gfx12: program SPI_SHADER_PGM_RSRC4_GS for seperately compiled gs
radv/gfx11+: program INST_PREF_SIZE for NGG and HS
radeonsi: use ac_get_instr_prefetch_size
radeonsi: use exec_size from the aco prolog/epilog callback
aco/ra: fix inline constants with v_dot2c_f32_f16
aco/sched_vopd: fix v_dual_dot2acc_f32_f16 created from VOP2 with inline constant
radv: fix setting inline push constants when only the last one is used
radv: inline 8 and 16bit push constant loads
aco/tests: test v_pk_fmac_f16 and v_dotc_f32_f16 with inline constants
aco/tests: test creating v_dual_dot2acc_f32_f16 from v_dot2c_f32_f16 with inline constant
nir/skip_helpers: fix stores with ACCESS_INCLUDE_HELPERS
nir/skip_helpers: handle vendored store_scratch
nir/skip_helpers: keep descriptors uniform even for stores that skip helpers
nir/skip_helpers: don’t require helpers for non uniform descriptors
aco/assembler: chain branches in emit order
aco/assembler: do not abort when exec is written after position exports
ac/nir/mem_vectorize: never create vec5 stores
aco/assembler: don’t reorder branch insertion block index twice
radv/gfx11+: do not use s[0:1] for unused scratch VA in compute shaders
radv: remove some dead compute scratch code
zink/ci: skip unvanquished-ultra trace on van gogh too
panfrost/lower_bool_to_bitsize: do not assume loop phi source order
nir/phi_builder: do not sort predecessors for phi sources
nir/to_lcssa: do not sort predecessors for phi sources
nir: generalize loop simplification
nir: add pass to optimize shared variables to subgroup operations
nir/opt_algebraic: fix vkd3d-proton pack_half_rtz pattern
aco/isel: emit v_mul_i32_i24 for imul with negative constant
aco/isel: emit v_mul_hi_i32_i24 for imul_high if possible
aco: remove isel setup code for no longer implemented intrinsics
nir,amd: split SGPR input intrinsic to specify workgroup divergence
ac/lower_intrinsics_to_args: use workgroup divergent ttmp intrinsic for subgroup id
nir/divergence: always consider load_ttmp_register_amd uniform
amd: use load_scalar_arg_wg_div_amd for workgroup divergent sgprs
nir/divergence: alyways consider load_scalar_arg_amd uniform
spirv: add option to treat FMax/FMin/FClamp like NMax
radv: add radv_force_nan_preserve_min_max option
radv: enable radv_force_nan_preserve_min_max for DOOM: The Dark Ages
nir: remove explict num_components from nir_def_rewrite_uses_with_alu_src
nir/opt_vectorize: prefer to swizzle vector phis at the destination, not the source
ac/nir: vectorize phis
radv: call nir_opt_phi_precision
nir: clean up weird qsort_r usage
nir/opt_shrink_vectors: restore load_const deduplication
nir/opt_sink: don’t sink comparisons that use ballot(true)
radv: run nir_opt_reassociate_for_fma for VS/GS too
vulkan/nir_lower_heaps: assume no heap addressing can overflow
nir: add num_lsb_zero analysis for 64bit pack and u2u
ac/nir_lower_global_access: assume both addition operands are aligned if one is
nir: add num_lsb intrinsic index for amd arg loads
radv: add dword alignment information to descriptor set/heap pointers
ac/nir/lower_ngg: use workgroup divergence analysis for culling
ac/nir/lower_ngg: allow reuse of workgroup divergent variables even when subgroup ops are used
nir/opt_dead_write_vars: handle atomics as reads
nir: fix divergence for deref_cast
aco/live_var_analysis: make sure shared vgprs are within the encodable vgprs
nir/unsigned_upper_bound: fix float to int conversions
nir: mark some AMD specific shuffles as subgroup ops
aco/optimizer: fix skip_smem_offset_align
nir: support phi sources in nir_rematerialize_deref_in_use_blocks
nir/to_lcssa: fix progress for derefs
nir/to_lcssa: move constants before the loop instead of creating a phi
Gert Wollny (67):
r600/sfn: Add lowering of tess inner and outer default intrinsics
r600: replace TGSI TCS passthrough with NIR version
r600: replace TGSI query shader with nir
r600/sfn: run nir_opt_idiv_const
r600/sfn: Avoid creating group-tagged registers for ALU dests
r600/sfn: signal progress when splitting address loads
r600/sfn: run additional optimization only after successful address split
r600/sfn: Extract some helpers from schedule_alu
r600/sfn: don’t use return parameters in extracted method
r600/sfn: Extract schedule alu groups first
r600/sfn: extract fill_alu_group
r600/sfn: pass reference to group when possible
r600/sfn: Extract group fill failure handling
r600/sfn: extract t-slot allocation when filling ALU groups
r600/sfn: extract idx load state handling in scheduler
r600/sfn: make ALU scheduling return values more meaningful
r600/sfn: collaps no_schedule and scheduled
r600/sfn: simplify ALU scheduling failure handling
r600/sfn: split kcache evaluation into try and commit
r600/sfn: make try_kcache_reservation const
r600/sfn: move tracking of kcache reservation failure to scheduler
r600/sfn: Move tracking of kcache reservation to AluScheduleContext
r600/sfn: extract kcache check out of schedule_alu_to_group_vec
r600/sfn: refactor BlockScheduler::schedule_block
r600/sfn: Move exports emission to helper
r600/sfn: extract check and report for unscheduled instructions
r600/sfn: deduplicate some code in DCE
r600/sfn: deduplicate optimizer logging code
r600/sfn: deduplicate fixpoint loop for optimizers
r600/sfn: refactor CopyPropFwdVisitor::propagate_to
r600/sfn: refactor CopyPropFwdVisitor::visit(AluInsr*)
r600/sfn: extract logging from CopyPropFwdVisitor::visit(AluInstr*)
r600/sfn: refactor CopyPropBackVisitor::visit(AluInstr*)
r600/sfn: Fix typo with AssemberVisitor
r600/sfn: Make some member variables references
r600/sfn: Refactor AssemblerVisitor emit_alu_op
r600/sfn: de-duplicate emit_wait_ack
r600/sfn: use c++ pattern for zero-init of structs
r600/sfn: extract some byte code emission from assembler
r600/sfn: Drop index register handler in assembler
r600/sfn: simplify fill bytecode
r600/sfn: extract emitting the bytecode of Rat Instr too
r600/sfn: extract and decouple ALU post-emit state update
r600/sfn: return LDS opcode properties as tuple
r600/sfn: use opcode switch in ALU post-emit update
r600/sfn: use local opcode consistently in emit_alu_op
r600/sfn: Move lds_queue_read decrement out of prepare_alu_src to caller
r600/sfn: Move copy_src to sfn_fill_bytecode.cpp, rename to fill_alu_src
r600/sfn: Move prepare_alu_src to sfn_fill_bytecode.cpp, rename to fill_alu_src_operands
r600/sfn: Move prepare_alu_dst/copy_dst to sfn_fill_bytecode.cpp, rename to fill_alu_dst
r600/sfn: drop unused literals tracking in assembler
r600/sfn: minor reordering of operations in assembler
r600/sfn: Move last_addr handling out of fill_alu_dst
r600/sfn: Simplify m_last_addr tracking in emit_alu_op
r600/sfn: Validate ALU dst writes in emit_alu_op
r600/sfn: Extract ALU bytecode emission helper
r600/sfn: Handle dst write checks before mova setup split
r600/sfn: Extract ALU dst state update into AssemblerVisitor
r600/sfn: Move LDS ALU emission to fill_bytecode
r600/sfn: Make emit_alu_op return success status
r600/sfn: Move opcode_map to fill_bytecode, pass EAluOp to emit_bytecode_alu
r600/sfn: Move ds_opcode_map ownership to fill_bytecode
r600/sfn: Drop unused AssemblerVisitor members
r600/sfn: Add pin_to_chan method to Register and use it
r600/sfn: rename pin_dest_to_chan to pin_registers
r600/sfn: Pin alu sources as well when registers are pinned
r600/sfn: Drop assertions when emitting IF asm instruction
Gleb Mazovetskiy (1):
os_misc.c: add missing include for mach_host_self()
Gleb Popov (1):
Rename the CACHE_LINE_SIZE define to MESA_CACHE_LINE_SIZE
Grant Nichol (1):
ethosu: Fix -Werror=format build error on 32-bit
Gu, Wangfeng (3):
radv/sqtt: add instruction timing SE mask controls
radv/sqtt: emit pending barrier end before API markers
ac/spm: clamp cache miss counts in derived counters
Gurchetan Singh (14):
gfxstream: fix string array marshalling
gfxstream: emit global state wrapped decoding for vkCmdEvent
subprojects: update libc-rs to 0.2.185
subprojects: update to rustix 1.1.4 + downstream patches
freedreno: fix ignored qualifier
tu: fix -Wmissing-prototypes errors
tu: fix implicit fallthrough
tu: kgsl: fix -Wgnu-alignof-expression warning with Clang
freedreno: explicitly declare required depend_files, part 1
freedreno: explicitly declare required depend_files, part 2
util: rust: sync error handling fixes from downstream
util: rust: minor fixups
virtio: add magma-gpu-rs subdirectory
docs: fix references to moved crates
Han, Mike (3):
amd/vpelib: complete 16bpc RGBA format mapping for 10/12bpc msb/lsb support
amd/vpelib: add format support check
amd/vpelib: Add missing argb variant support
Hans-Kristian Arntzen (21):
wsi/common: Report correct time domain in VkPresentTimingInfo.
loader: Separate out X11 specific screen queries from dri_helper.h.
loader: Clear screen resources struct on init.
wsi/x11: Setup screen resources on x11_connection creation.
wsi/x11: Add helper to find appropriate screen resources for a window.
wsi/x11: Set up screen resources on swapchain creation.
wsi/x11: Add helper to compute xrandr rate estimate.
wsi/x11: Update xrandr refresh estimate on geometry change.
wsi/x11: Update refresh rate estimate based on MSC feedback.
wsi/x11: Implement main body of present timing.
wsi/x11: Add Xwl support for present timing.
wsi/x11: Only accept VRR refresh rates when we’re flipping.
wsi/common: Prefer host query resets when available.
wsi/common: Pass along requested timing feedback as well.
wsi/x11: Avoid non-causal present timings when not flipping.
wsi/common: Refactor out the search for a present_timing struct.
wsi/common: Ensure that google display timing results propagate.
wsi/common: Always ensure that we can get a GPU done timestamp.
wsi/x11: Be more adaptive in how much the sleep is pulled back.
radv: Consider VkImageView usage rather than VkImage usage in feedback.
radv: Only consider default feedback loops for appropriate layouts.
Hsieh, Mike (2):
amd/vpelib: add optional __stdcall calling convention via build option
amd/vpelib: add indirect shaper config support
Hyunjun Ko (13):
anv/video: fix up H.264/H.265 encode session parameters to match advertised caps
anv/video: fix to set the upper bound of the bitstream of h265.
anv/video: Add to check size mismatch during motion field estimation.
anv/video: define ANV_VIDEO_AV1_MAX_DPB_SLOTS
anv/video: Change size of the cached array of recently decoded AV1 frames.
intel/genxml: update VDENC commands for gen125
anv/video: Add h264 vdenc tables from media-driver
anv/video: Make H264 encoder work on Gen125
anv/video: fix to set valid coded size for the source pictures.
anv/video: Add h265 vdenc tables from media-driver
anv/video: Make H265 encoder work on Gen125
anv/video: Enable video encoding on gen125
anv/video: Support H265 10-bit encoding
Iago Toral Quiroga (2):
pan/bi: TEX_GRADIENT may need helper invocations
CODEOWNERS: update broadcom maintainers
Ian Romanick (21):
brw: Lower all phis to scalar
brw: Don’t lower phis involved in DPAS instructions to scalar
brw: Calcuate divergence before brw_from_nir
nir/opt_constant_folding: Don’t fight with nir_lower_bit_size
nir: Use nir_instr_remove_v in nir_def_replace
nir/opt_if: use nir_def_replace() instead of nir_def_rewrite_uses()
nir/opt_if: Merge if-statements with inverted conditions
nir/algebraic: Convert bcsel of addition to addition of b2i or b2f
nir/opt_shrink_stores: Don’t shrink ivec2 stores to int64 images
brw: Use nir_opt_shrink_stores
brw: Use nir_opt_shrink_vectors
brw: Add functions to calculate flags usage without a brw_inst
brw: Replace logical operations with predication
brw: Use nir_opt_uub
brw: Use nir_opt_fp_math_ctrl
elk: Use nir_opt_uub
elk: Use nir_opt_fp_math_ctrl
nir/divergence: Handle SYSTEM_VALUE_INSTANCE_INDEX
brw/predicate: Add missing test with farther_flags
brw: Handle empty top block in brw_nir_move_interpolation_to_top
brw/validate: Gfx11 can’t have accumulator src0 in 3-src instructions
Icenowy Zheng (36):
pvr: follow other drivers’ practice for copying build ID
pvr: skip emitting query program when copy result / reset with 0 queries
isaspec: decode: manually print the sign when printing NaN float values
pvr: wait for graphics jobs in CopyQueryPoolResults
pvr: increase maxPerStageResources for new maxPerStageDescriptorStorageBuffers
pvr: do not setup deferred RTA clear for active render targets
pvr: properly handle deferred RTA clears for 2D array view of 3D image
pvr: add deferred RTA clear command to list after checking it’s not NULL
pvr: record deferred RTA clears for secondary cmdbuf subcmds
pvr: ignore DS attachment’s D or S when it’s unused in dynamic rendering
dri: try to enable GL_ARB_compatiblity when supported GL core version is 3.1
pvr: setup viewindex if the shader wants it even when multiview disabled
pvr: prohibit clang-format from touching the dri options list
pvr: add dri options used by common WSI code
pvr: fix handling of invalid attachment info in pvr_init_fs_outputs_mrt
pvr: copy sub_cmd flags except owned when executing subcmds out of pass
pvr: stop to derive rt datasets based on geometry_terminate
pvr: add a structure containing data kept for suspended renderpasses
pvr: preserve and pass more data for suspending render passes
pvr: remove dEQP-VK.pipeline.monolithic.misc.no_rendering from fail list
pvr: return FORMAT_NOT_SUPPORTED for unknown image types
pvr: prevent direct access to VkImageSubresourceLayers::layerCount
pvr: implement CmdBindIndexBuffer2
pvr: implement GetRenderingAreaGranularity
pvr: implement GetImageSubresourceLayout2
pvr: implement GetDeviceImageSubresourceLayout
pvr: advertise VK_KHR_maintenance5
zink: move maint5 to gl21_baseline capabilities set
docs/zink: add maint5 to the list of required extensions
llvmpipe: stub other functions inside compute shaders for ORCJIT
pvr: bump conformance version to 1.4.3.3
Revert “pipe-loader: fallback to zink instead of kmsro for render nodes”
pipe-loader: use zink for powervr device nodes
zink: check Z/S aspect before creating Z/S image view
pvr: apply the culling everything viewport shift for only triangles
vulkan: update spec to 1.4.354
Iván Briano (15):
anv: silence warning
intel/brw: add load_coverage_mask_intel intrinsic
intel/brw: add load_msaa_rate_intel intrinsic
intel/brw: add load_frag_shading_rate_intel
anv/brw: add conservative raster on/off to FS_CONFIG
anv/brw: handle FullyCoveredEXT
anv: add and use a drirc option to enable FullyCovered for vkd3d
anv: fix return of cmd_buffer_set_indirect_stride() function
anv, iris: fix MOCS Index setting of EXECUTE_INDIRECT_* commands
intel/dev: ARL-H supports EXECUTE_INDIRECT_*
anv: don’t try to clear d/s attachments not backed by an image
brw/rt: fix max_t selection on intersection report
brw/rt: split HitAttribute area in pending/committed
brw/rt, anv: reduce maxRayHitAttributeSize
anv: fix 2d-array to 3d blits
Jaakko Jokinen (1):
nir: Add cases to nir_get_io_offset_src_number()
JaeHoon Lee (25):
v3d: release the texture reference if shadow resource creation fails
v3d: drop the tiled temporary when bailing on unsupported blits
v3d: free the cache buffer when loading a corrupt disk cache entry
v3d: create the compute job after the zero-sized dispatch check
v3dv: only report 16-bit float formats as blendable at 32/64 bpp
vc4: fix last_layer selection in the blit sampler view
v3d: fix slot and input indexing in v3d_set_global_binding
v3dv: report maxDrawIndirectCount of 1 without multiDrawIndirect
v3dv: honor wait dependencies for job-less submissions
v3d: clamp transform feedback offset to buffer size
v3dv: report the correct dynamic storage buffer UAB limit
v3dv: fix blake3 key truncated to 20 bytes in pipeline cache
v3d: fix blake3 key truncated to 20 bytes in shader cache
vc4: fix incorrect resource unref in vc4_flush_resource
vc4: free vertex and constant buffers on context destroy
broadcom/compiler: really enable GFXH-1625 TMUWT validation
broadcom/compiler: validate magic waddr writes
broadcom/qpu: remove empty qpu_validate.c
v3dv: make room in the descriptor map for the no-sampler entries
v3dv: use the binning VS variant for the binning VPM config
v3dv: record the multiview geometry shader with its Vulkan stage bit
v3dv: record the no-op fragment shader with its Vulkan stage bit
nvk: free copy_memory_indirect_temps on command buffer destroy
v3dv: avoid restoring stale descriptor state after a meta op
nvk: report fills from memory correctly
Jaishankar Rajendran (2):
vulkan/runtime: enable parametrization of ASTC software decode
anv: tune parameters of the ASTC software decoding
Jakob Sinclair (13):
panvk: Enable scissor_mode for draws
panvk: Remove unnecessary functions
vulkan/meta: Don’t issue a full drawcall for clears
pan: Support lowering D24X8 to D24
gallium: fix type size in z24_unorm_packed_pack_z_32unorm
pan/va: Decode support for ARSHIFT_OR on Valhall
pan: Add G52 skip for xlib wsi failure
pan/compiler: fix spilling for 64-bit values
pan: Add missing v14 primitive flag
panvk/draw: Separate build from prepare functions
panvk/csf: Use RUN_FULLSCREEN for cmd_draw_rects
panvk/csf: Use RUN_FULLSCREEN for cmd_draw_volume
pan/crc: Fix CRC check for sparse AFBC images
Jan Meisel (3):
nir/range_analysis: handle msad_4x8 in unsigned upper bound
radeonsi/vcn: fail feedback for truncated encodes
radv: fix RADV_PERFTEST=nircache enablement
Janne Grunau (7):
nir/gather_info: clear interpolation qualifiers only in fragment stage
asahi: nir: lower flrp64
asahi: ci: Drop no longer failing VK.wsi.xcb.present_timing test
panfrost: ci: Drop no longer failing VK.wsi.xcb.present_timing test
asahi: ci: Add failing b10g11r11 and e5b9g9r9 copy tests
hk: xfb: Avoid assertions in nir_slot_num_components
poly: Fix comment after moving passthrough_gs
Jason Macnak (6):
gfxstream: Override VkDeviceDeviceMemoryReportCreateInfoEXT vk.xml
virtgpu_kumquat_ffi: replace mutex.get_mut() with mutex.lock()
gfxstream: support testing d32 s8
gfxstream: kumquat: validate device dmabuf support before use
gfxstream: route vkGet*ProcAddr to VkDecoderGlobalState
gfxstream: Avoid transfering VkAllocationCallbacks between guest and host
Jeremy Gebben (5):
kk: Implement VK_KHR_dynamic_rendering_local_read
kk: Refactor encoder state
kk: Set availability for extra multiview queries in vkCmdEndQuery()
kk: Implement VK_QUERY_TYPE_TIMESTAMP
kk: Fix Vulkan to Metal stage translation for timestamps
Jeremy Huddleston (38):
bin/install_megadrivers: Bail out if libname suffix is never reached
gallium/targets: Use libname_suffix for installed driver names
glx/apple: Convert K&R-style declarations to ANSI prototypes
glx/apple: Switch logging to os_log on macOS 10.12+
glx/apple: Replace apple_glx_diagnostic with apple_glx_log_*
glx: free visinfo on BadMatch in glXCreateWindow’s AppleGL path
glx: Fix stale end-comment on __glXInitialize direct-rendering block
glx: drop redundant __glXErrorString forward declaration
glx: fix DRI3-not-available diagnostic skip on macOS
glx: NULL-check frontend_screen in glXCreateContextAttribsARB
glx: simplify FBConfig wire decode
glx: free glx_drawable on CreateDRIDrawable failure
glx: bail bind_extensions on screens without frontend_screen
glx: drop dead AppleGL glXGetProcAddressARB fallback
glx/apple: Add create_context_attribs entry to the applegl_screen_vtable
zink: fix GLX_USE_APPLE typo (should be GLX_USE_APPLEGL)
glx: drop dead GLX_USE_APPLE check inside glXSwapBuffers
glx: Guard declaration of glx_accel and kopper to match use
glx/apple: free gc in applegl_destroy_context
glx: fix per-display drawHash / zombieGLXDrawable / dri2Hash leak on GLX_USE_APPLE builds
glx/apple: silence OpenGL deprecation warnings
glx/apple: return CGLError from apple_visual_create_pfobj instead of aborting
glx: route copy_context through a vtable slot
glx: route swap_buffers through a vtable slot
glx: extract drawable lifecycle into a vtable
glx/apple: allow selection between AppleGL and Gallium at runtime for GLX_USE_APPLE=1 builds
glx/apple: skip AppleGL election on macOS 26 and newer
glx: remove GLX_USE_APPLE and collapse the guards it gated
glx: Fold __glXGetDrawableAttribute and __glXQueryDrawable together
glx: Fold CreatePbuffer/DestroyPbuffer into CreateDrawable/DestroyDrawable
zink: Add missing link against libxcb-present
zink: Address libvulkan.1.dylib dlopen failure on macOS
glx/apple: honor the client-requested GLX context version and profile
llvmpipe: link all LLVM targets on Apple to fix build failure when using static LLVM libraries
dri/st: Fall back to Z32_FLOAT depth configs when Z32_UNORM is unsupported
glx/apple: only skip AppleGL election on macOS 26.0 through 26.5
llvmpipe: don’t create a screen when the process is not allowed to JIT
glx/apple: silence OpenGL deprecation warnings in libglx
Jesse Natalie (23):
d3d12: Handle THREAD_SAFE maps and use them for async query results
microsoft/compiler: Back-propagate interpolator modes from FS
wgl: Use an hwnd xor hdc for framebuffers
d3d12: add screen pending-free list plumbing
d3d12: clear stale per-context BO state at context destroy
d3d12: transfer batch local_bos refs to screen at submit
d3d12: transfer batch->bos refs to screen at submit
d3d12: reclaim in-flight BO memory on allocation failure
d3d12: implement pb_fence vtbl for cache/slab reuse
d3d12: drop peer-batch peeking in resource_is_busy / wait_idle
d3d12: proactively trim completed pending-free entries
nir_lower_non_uniform_access: Add ASSERTED for assert-only var
va: Wrap assert-only code in NDEBUG
microsoft/compiler: Don’t assume phi ordering
util: Fix u_math on MSVC arm64
mesa/st: PBO memory barriers imply image barrier if PBO download goes through compute
d3d12: Use enhanced barriers for memory_barrier when we can
d3d12: Fix transition_array_size for 3D textures
d3d12: Fix WARP version detection for broken int64
ci/windows: Update WARP to 1.0.20
wgl: Move sub-8bpc pixel formats to extended format list to match other Windows drivers
d3d12: Disable vao fast path for AMD
mesa: Fix shared state bookkeeping for dynamic share list changes (wglShareLists)
Jhanani Thiagarajan (1):
intel/mda: Change the default output directory
Jianfeng Liu (1):
freedreno/drm: Fix uninitialized read of BO metadata on import
Jianxun Zhang (1):
intel/decoder: Print more information in shader’s headline
Jiyu Yang (3):
nir/loop_analyze: Use pass_flags for memoization in is_only_uniform_src
panfrost: cleanup precomp_cache on screen destroy
egl/dri2: exclude >8bpc configs from GLES1 renderable/conformant bits
Job Noorman (60):
ir3/ra: fix killed src detection while spilling
ir3/shared_ra: fix live-out reload after src reload
nir/get_io_offset_src_number: support @load/store_global_ir3
ir3/isa: use same src for ldg.a OFF field on a6xx/a7xx
ir3: always use byte offset for @load/store_global_ir3
nir/opt_offsets: add support for @load/store_global_ir3
ir3: move feature check down in ir3_nir_max_imm_offset
ir3: enable opt_offsets for load/store_global_offset
ir3: mark __alias_n as UNUSED in foreach_src_in_alias_group_n
ir3/cf: fix rewriting uses with different dst types
ir3/shared_ra: use ir3_cursor instead of instr in reload helpers
ir3/shared_ra: insert reloads before tied dst pcopies
ir3/cp: support propagating const vecs
ir3: allow const src0 for ldg.a/stg.a/ray_intersection
ir3: don’t cache driver param instructions
ir3: allow (ss) on all cat7 instructions
freedreno/computerator: fix UAV view size
ir3/spill: extract child intervals for live-in reloads
ir3/ra: add ir3_ra_src_is_killed helper
ir3/ra: fix killed src detection for spillall min limit
ir3: use a1.x addressing for ldg.k with dst 256
ir3: don’t use bitfields in ir3_shader_output
ir3: don’t store shader_options in the cache
freedreno/drm-shim: allow chip selection by chip_id
tu: use chip_id instead of gpu_id for the cache UUID
tu: add option to override the build ID
vulkan: add vk_shader_module_hash helper
vulkan: use consistent module hashing for pipeline stages
nir/lower_undef_to_zero: add filter argument
ir3: lower undef booleans to zero
nir/get_io_index_src_number: support @load_ssbo_address
nir/lower_ssbo: take offset_shift into account
nir/lower_ssbo: add option to only lower large SSBOs
nir/lower_ssbo: add option to insert bounds checks
tu: Add option to raise the maximum SSBO size
ir3: fix possible signed overflow in ir3_link_add
ir3/opt_prefetch_descriptors: rematerialize defs at preamble start
nir/lower_vars_to_scratch_global: make callback deterministic
ir3/lower_vars_to_scratch_global: use stable sort for variables
nir: add nir_shader_deref_pass
nir: add nir_src_as_{alu,tex,phi}_src helpers
nir: add nir_convert_address_format pass
nir/lower_explicit_io: add support for 64bit_global_32bit_offset vars
nir/lower_explicit_io: support shifting non-const array index
nir: add load/store_global_offset intrinsics
nir/lower_explicit_io: add support for load/store_global_offset
nir/set_io_offset: add support for adjusting BASE
nir/lower_explicit_io: support offset_shift for 64bit_global_32bit_offset
rusticl/kernel: add support for 64bit_global_32bit_offset
ir3: add support load/store_global_offset
ir3: enable opt_offsets for load/store_global_offset
tu,ir3: use 64bit_global_32bit_offset for global memory
ir3/lower_tess: use load/store_global_offset
ir3/lower_shader_clock: use load_global_offset
ir3: don’t manually lower load/store_global
tu/lower_ray_query: use load_global_offset
tu,ir3/analyze_ubo_ranges: use load_global_offset
tu/lower_ssbo_address_size: use load/store_global_offset
nir,ir3: remove load/store_global_ir3
ir3/opt_preamble: lower load_global_offset to preamble
Joe Wang (2):
ac/spm: bump AC_SPM_MAX_COUNTERS_PER_GROUP to 16
radv,ac/spm: add user-defined raw counter collection
Jon Turney (7):
ddebug: Fix use of alloca() without #include “c99_alloca.h”
glx/windows: Avoid shadowing ‘type’ parameter of driwindowsCreateDrawable()
glx/windows: Add stdbool.h include to ‘direct GLX via WGL’ implementation
glx/windows: Fix compilation of driwindows_glx after driscreen changed from pointer to member
glx/windows: Fix compliation after code motion to put event base in ‘dri’ context
glx/windows: Add GLX_USE_WINDOWSGL in new places it’s needed to build libGL
glx/windows: Drop static from driwindowsCreateScreen()
Jordan Justen (48):
brw: Don’t set header_size at init since it will be re-set in later code
brw/compact: Precompact using 2src fields on 3src instructions
intel/gen: Add gen 9 through Xe2 instruction formats in JSON
intel/gen: Add gen_inst_info.py script to generate C++ headers
intel/gen: Create gen_info_util.h
intel/gen: Make use of generated instruction info
intel/gen/compact: Add compact tables from brw/brw_eu_compact.c
intel/gen: Add gen_raw_compact_inst type
intel/gen: Add gen_compact_accessor for compact/uncompact
intel/gen: Implement compact support
intel/gen: Split out type decode functions for use with uncompact
intel/gen: Implement uncompact support
intel/gen: Account for compact nop pad instruction in gen_scan_raw_layout()
intel/gen: Support declaring ISA fields with disconnected bits
intel/gen: Support accessing fields & sub-fields with disconnected bits
intel/gen: Merge THREE_SRC0_VSTRIDE HI/LO into a gen_split_range
intel/gen/xe: Merge THREE_SRC1_VSTRIDE HI/LO into a gen_split_range
intel/gen/xe: Merge BFN_FUNC_CONTROL HI/LO into a gen_split_range
intel/gen: Merge uncompat control bits into a gen_split_range
intel/gen: Merge uncompat datatype bits into a gen_split_range
intel/gen: Merge uncompat subreg bits into a gen_split_range
intel/gen: Merge uncompat src0 bits into a gen_split_range
intel/gen: Merge uncompat src1 bits into a gen_split_range
intel/gen: Merge uncompat 3src control bits into a gen_split_range
intel/gen: Merge uncompat 3src source bits into a gen_split_range
intel/gen/xe: Merge uncompat 3src subreg bits into a gen_split_range
intel/gen: Merge SRC_A16_SWIZZLE HI/LO ranges
intel/gen/xe: Merge Xe2 DATATYPE_INDEX HI2/LO3 into a gen_split_range
intel/gen/xe: Merge Xe2 compact 3src subreg HI2/LO3 into a gen_split_range
intel/gen: Assert that the gen opcode is supported by this platform
intel/gen: Start Xe3P support
intel/gen: Add gen_byte_stride()
intel/gen: Add Xe3P validation for src1 byte stride matching dst
intel/gen/validation: Start enabling Xe3P tests, but skip for now
intel/gen/validation: Update tests for new Xe3P src1 restriction
intel/gen/validation: Drop WA 22016140776 on Xe3P
intel/gen/validation: Enable running validation tests for Xe3P
intel/gen: Remove mac/mach/macl on Xe3P
intel/gen: Add mullh instruction for Xe3P
intel/gen/xe: Rename decode/encode_type_3src to decode/encode_type_short
intel/gen: Disable compact on Xe3P for now
intel/gen: Add Xe3P two source encoding changes
intel/gen/tests/basic: Add 2src round-trip tests covering nvl src1 changes
intel/gen: Add Xe3P three source src1 encoding changes
intel/gen/tests/basic: Add 3src round-trip tests covering nvl src1 changes
intel/gen/compact: Split datatype into 1src / 2src versions
intel/gen: Add Xe3P compact support
intel/executor: Enable Xe3P
Jose Maria Casanova Crespo (59):
broadcom/compiler: Add V3D 7.1 v8dot dot product QPU instructions
broadcom/compiler: hardware-accelerated 4x8-bit dot products on V3D 7.1+
broadcom/compiler: Add v8dot and setnnmode scheduler dependencies.
broadcom/compiler: Eliminate redundant setnnmode instructions
v3dv: Expose hardware-accelerated integer dot products on V3D 7.1+
broadcom/compiler: move nir_lower_undef_to_zero out of optimization loop
v3dv: bump maxComputeSharedMemorySize to 32 KB
v3d/v3dv: Use new V3D_MAX_CSD_WG_SIZE = 256
v3dv: lower oversized compute workgroups to 256 invocations
v3dv: include mem_offset in vkCmdFillBuffer destination
v3dv: Enable KHR_shader_subgroup_extended_types
v3dv: expose maxFragmentOutputAttachments as max_rts
v3dv: avoid duplicate bo_handles between cpu_job and CSD lists
v3dv: assert timestamp pool BO is disjoint from dst buffer BO
broadcom/ci: skip SSBO tests close to the 60s threshold on rpi4
v3dv: avoid 16F TLB usage for B10G11R11_UFLOAT copies
v3dv: advertise VK_EXT_scalar_block_layout on V3D 7.1+
v3dv: Enable meta_copy_buffer with TFU for V3D 7.1
v3dv: move destroy_update_buffer_cb to a generic helper
v3dv: use TFU copy with stride-0 for vkCmdFillBuffer
v3dv: extract TFU helpers for format-plane and slice-stride args
v3dv: rename copy_buffer_to_image_tfu to copy_buffer_image_tfu
v3dv: implement TFU image-to-buffer copy on V3D 7.1
v3dv: relax buffer padding in TFU buffer<->image copy
v3dv: share zero-fill TFU staging BO at device level
v3dv: expose the full simulator memory to applications
broadcom/qpu: support output pack on itof/utof
v3d: move nir_lower_frexp after nir_lower_bit_size
v3dv: lower flrp16 for consistency with flrp32
v3d: widen sub-32-bit subgroup arithmetic and vote ops
v3d: improve liveness analysis for packed partial writes
broadcom/qpu: expose V3D 7.1 packed-f16 instructions
v3d: emit packed-f16 ALU ops natively on V3D 7.1
v3dv: enable lowered shaderFloat16/Int16/Int8 + VK_KHR_shader_float16_int8
broadcom/compiler: fix payload-register liveness condition
v3d: Enables GL_ARB_clip_control for v71+
v3d: use NO_GUARDBAND clipper for near-zero viewport Z scale
v3dv: set non-zero array stride in null texture descriptor state
broadcom: add and use max_render_targets to devinfo
broadcom: raise framebuffer size to 7680 on V3D 7.1
v3dv: gate Dawn-required limits and features behind V3D_WEBGPU_OVERRIDE
ci: igalia farm maintenance
v3dv: allow TFU readahead padding above maxMemoryAllocationSize
v3dv: route blending of UNORM16/SNORM16 RTs through software lowering
v3dv: rename format_plane unorm/snorm flags to sw_unorm/sw_snorm
v3dv: fix crash on device creation failure before meta initialization
v3dv: close the primary node fd on physical device destruction
broadcom/compiler: reduce the compile-strategy fallback ladder
broadcom/compiler: split v3d_nir_to_vir_finish out of v3d_nir_to_vir
broadcom/compiler: support probing a compile’s pre-spill register pressure
broadcom/compiler: consolidate the compile-strategy logging
broadcom/compiler: move the 2-thread strategies to compile_2t_strategies
broadcom/compiler: pick the 2-thread compile strategy by register pressure
broadcom/compiler: don’t leak the compile on assembly allocation failure
broadcom/compiler: drop V3D_DEBUG=opt_compile_time
broadcom/ci: unskip CTS tests that are no longer slow
broadcom/compiler: abort 2-thread spill loops over the best result so far
vc4: save the fragment constant buffer around the YUV blit
vc4: unbind the textures around the blitter clears
José Roberto de Souza (40):
anv: Change fill_inline_params() first parameter from struct GENX(COMPUTE_WALKER_BODY) to uint32_t *
anv: Move VMA heaps init and finish of vma heaps to anv_va.c
anv: Move init and finish of state pools to its own functions
anv: Move code to load color border to memory to a function
intel/brw: Explicitly upcast UB to UW for SHR with vector immediates
intel/tools: Fix parse of ‘[HWCTX].replay_*’ in aubinator_error_decode_xe
intel: Sync xe_drm.h
intel: Add support for madvise purgeable VMAs in Xe KMD
iris: Improve and standardize the behavior of madvice in i915
intel/brw: Fix nir_intrinsic_load_inline_data_intel register offset calculation
intel/dev: Remove unused intel_get_device_info_for_build() function
intel/dev: Add URB max entries values
intel/dev: Use URB mesh/task min/max values in intel_device_info
intel/dev: Add a Xe2+ table of URB min and max entries
anv: Add assert to make sure we don’t push more than max_push_regs to push constants
anv: Replace most parameters of fill_inline_param() by a struct
anv: Add function to get each anv_state_pool
anv: Use anv_device_get_general_state_pool()
anv: Use anv_device_get_aux_tt_pool()
anv: Use anv_device_get_dynamic_state_pool()
anv: Use anv_device_get_binding_table_pool()
anv: Use anv_device_get_scratch_surface_state_pool()
anv: Use anv_device_get_internal_surface_state_pool()
anv: Use anv_device_get_bindless_surface_state_pool()
anv: Use anv_device_get_indirect_push_descriptor_pool()
anv: Use anv_device_get_push_descriptor_buffer_pool()
anv: Replace va.bindless_surface_state_pool access with a function
anv: Replace va.indirect_descriptor_pool access with a function
anv: Replace va.dynamic_visible_pool access with a function
anv: Replace va.internal_surface_state_pool access with a function
anv: Replace va.dynamic_state_pool access with a function
anv: Replace va.indirect_push_descriptor_pool access with a function
anv: Replace va.push_descriptor_buffer_pool access with a function
anv: Replace va.aux_tt_pool access with a function
anv: Replace va.binding_table_pool access with a function
anv: Replace va.scratch_surface_state_pool access with a function
anv: Fix memcpy overflows around sampler state
anv: Support sampler state of different sizes
anv: Replace anv_descriptor_set_binding_layout::descriptor_data_sampler_size by a local variable
anv: Fix copy of sampler state for bindless
Juan A. Suarez Romero (36):
v3d: mark mapped BO as initialized for valgrind
v3d/ci: add OpenCL regressions
ci: igalia farm maintenance
Revert “ci: igalia farm maintenance”
broadcom/ci: update kernel for nightly runs
vc4/ci: update expected results
loader: check if the kernel driver is amdgpu
broadcom/simulator: V3D is always 4.2 or above
v3dv: allow device with only render node
v3d/drm-shim: add GPU selection
v3dv: disable threadeded submissions under drm-shim
broadcom/ci: update kernel for nightly jobs
Revert “ci: igalia farm maintenance”
mesa: allow GL_TEXTURE_COMPARE_{MODE,FUN} with EXT_shadow_samplers
v3dv: increase max push constants size
Revert “people: update Marek’s email”
broadcom/ci: upgrade kernel in DuTs
v3d/ci: update expected results and document failures
v3dv: fix assertion on push constants
st/mesa: release sampler view
rusticl: fix leak in `util_queue`
v3d: initialize value in query info
v3d: free vertex and constant buffers on context destroy
v3d: add more blitter ops for saving resources
vc4: mark mapped BO as initialized for valgrind
vc4: initialize value in query info
vc4: move util_copy_constant_buffer to the function beginning
vc4: add blitter operations
doc/features.txt: enable VK_KHR_shader_float16_int8 for v3dv
v3dv: fix buffer creation usage flags validation
doc/features.txt: fix VK_KHR_shader_float16_int8 for v3dv
v3dv: enable VK_KHR_shader_quad_control / clustered subgroups
v3dv: enable VK_KHR_shader_subgroup_rotate / rotate subgroups
v3dv: enable VK_KHR_shader_maximal_reconvergence
v3d: add support for array of textures blit with TFU
v3d: save fragment constants on sand8/sand30 blit
Julia Zhang (15):
radv: add new option RADV_DEBUG=notmz
radv: allocate encrypted rings BOs
radv: create encrypted BOs for protected cmd_buffers
radv: enable surface protected capability
radv: set TMZ bit in sdma_copy packet
radv: save protected queue and non-protected queue seperately
radv: advertise VK_EXT_pipeline_protected_access
vulkan/wsi: return image compression properties for surface formats
vulkan/wsi: pass compression control when filtering DRM modifiers
vulkan/wsi: copy swapchain compression fixed-rate flags
radv: reject DCC modifiers when image compression is disabled
radv: advertise EXT_image_compression_control_swapchain
radv/ci: skip compression_control cases
radv: implement bo_wait_for_idle
radeonsi: avoid unmatched Perfetto events
Julien Schueller (4):
glx: avoid crash on glXBindTexImageEXT when no texture target set
egl: fix _EGL_NATIVE_PLATFORM fallback for unrecognized native displays
drisw_glx: handle XGetGeometry failure in get_drawable_geometry
st/drawpixels: tile images larger than max texture size instead of clamping
Kajal Kajal (1):
freedreno/blitter: copy full depth of src box in resource_copy_region
Karmjit Mahil (34):
freedreno/decode,ir3: Mark decoded dwords as const
freedreno/decode: Fix error() in script.c
freedreno: Don’t set UCHE_CLIENT_PF
gbm: Remove unused ARRAY_SIZE macro
gbm: Replace VER_MIN with common MIN2
docs: Fix struct redefinition errors
util: Add heap_memory_percent driconf option
vulkan: Add heap budget helper function
hk: Add heap_memory_percent driconf support
asahi: Add heap_memory_percent driconf support
v3dv: Add heap_memory_percent driconf support
v3d: Add heap_memory_percent driconf support
tu: Add heap_memory_percent driconf support
freedreno: Add heap_memory_percent driconf support
panvk: Add heap_memory_percent driconf support
panfrost: Add heap_memory_percent driconf support
nvk: Add heap_memory_percent driconf support
crocus: Add heap_memory_percent driconf support
pvr: Add heap_memory_percent driconf support
vc4: Use os_get_gpu_heap_size()
freedreno/computerator: Remove VLA giving a build warning
util/u_trace: Fix copy_func indentation in generated code
util/u_trace: Fix indentation of generated _trace function
util/u_trace: Evaluate copy_func expression in python
util/u_trace: Refactor TracepointArgStruct
util/u_trace: Commonize emitted print functions
util/u_trace: Move copyright into a Mako def
util/u_trace: Commonize trace function header with a Mako def
util/u_trace: Avoid sprintf when we already have a string
util/u_trace: Add ArgBlob for storing structs in traces
util/u_trace: Add u_trace_backend_type
tu: Emit bin info in Perfetto render_pass
drm-shim/freedreno: Fix fprintf format specifier
android_stub: Replace __ANDROID_API_V__ with 35
Karol Herbst (195):
nak: the MS location comes last in TLD, same spot as depth compare in TEX
mesa/st: do not advertise CL subgroup features on the GL side
radeonsi: advertise support for subgroup rotate
iris: advertise support for subgroup rotate
nak/lower_cf: remove single src phis
nak: call nir_opt_fp_math_ctrl
nak: call nir_opt_algebraic_distribute_src_mods
ci: install libstdc++-static on fedora
rusticl: link the C++ runtime statically
softfloat: make sign bit an unsigned int
nir: add fmul_rtz
nir: handle fmul_rtz in a couple of places
nak: handle nir_op_fmul_rtz
nak: use fmul_rtz for NAK_INTERP_MODE_PERSPECTIVE
nir: add fmul_rtz optimizations
nir/lower_cl_images: call nir_progress on every function
gallivm/nir/soa: use uint for booleans
llvmpipe: never pass a NULL function name to LLVMAddFunction
nvk: Use nvk_cmd_fill_memory in CmdResetQueryPool when possible
include: update CL headers
rusticl/program: handle CL_INVALID_CONTEXT for clCompileProgram and clLinkProgram
rusticl/kernel: update error code handling for clSetKernelExecInfo
rusticl/kernel: return CL_INVALID_WORK_GROUP_SIZE in clEnqueueNDRangeKernel for an explicit 0 workgroup
rusticl: start implementing CL 3.1 support
rusticl: implement CL 3.1 platform features
rusticl: implement CL 3.1 device features
docs/features: add OpenCL 3.1 section
bin/gen_release_notes: add paragraph on OpenCL support
rusticl: update names of types now core in 3.1
ci: update OpenCL 3.1 piglit fails
clc: do not use std::filesystem
Revert “rusticl: link the C++ runtime statically”
glsl/softfp: rename ffma to fmad
lima: rename ppir_op_ffma to ppir_op_fmad
nir: rename ffma to ffma_old
nir: rename nir_fmad to nir_fmad_old
nir: add new float multiply-add opcodes
nir: validate new float_mul_add options
nir/tests: handle new multadd opcodes
nir/opt_algebraic: add fmad and ffma_weak lowering rules
nir: handle new multadd opcodes in lowerings and opts
nir: handle new multadd opcodes in helpers
nir: duplicate old ffma opts where necessary for new multadd ones
nir/tests: use ffma_weak
ntt: use ffma_weak
llvmpipe: port over to ffma_weak
softpipe: keep weak_ffmas around
intel/elk: port over to nir_op_ffma
intel/jay: support nir_op_ffma
intel/brw: port over to nir_op_ffma
i915: support nir_op_fmad
nv50/ir: port over to new multadd opcodes
nv30: advertize new float multadd options
nak: port over to nir_op_ffma
zink: port over to nir_op_ffma_weak
ir3: port to nir_op_fmad
freedreno/ir2: use nir_op_fmad
tu: use nir_op_ffma_weak in lowering
ac: handle new float multadd opcodes
ac: use nir_op_ffma_weak
ac/llvm: support new multadd opcodes
aco: support new multadd opcodes
radv: use nir_op_ffma_weak
radeonsi: advertize new float multadd options
r600,sfn: support new multadd opcodes
r300: port over to nir_op_fmad
agx: port over to nir_op_ffma
kk: support nir_op_ffma
bitfrost: support nir_op_ffma
microsoft/compiler: support nir_op_ffma
d3d12: use nir_op_ffma_weak
etnaviv: port over to nir_op_fmad
pco: port over to nir_op_ffma
pvr: use ffma_weak for lowering
lima: support nir_op_fmad
svga: use weak_ffma
virgl: advertise new muladd options
nir: add fmad_or_ffma helpers and use it in lower_double_ops
nir: update ffma helpers to use new opcodes
nir: make lowering use new ffma opcodes
mesa: use ffma_weak
vulkan/meta: use nir_op_ffma_weak
tgsi_to_nir: translate MAD as ffma_weak
glsl: translate fma as fma_weak
vtn/glsl: translate fma as ffma_weak
vtn: handle OpFmaKHR
vtn: use ffma_weak
vtn/opencl: map mad to ffma_weak and fma to ffma
vtn_bindgen2: keep ffma_weak
nir: remove ffma_old
ci: update traces due to ffma rework
ci/windows: add dEQP-VK.glsl.builtin.precision_double.mix.compute.vec3 fail
zink: keep ffma_weak and use GLSLstd450Fma for it
zink: support nir_op_ffma
nir: add nir_intrinsic_cmat_load_shared_nv to nir_get_io_offset_src_number
nak/sm70: add helper for memory load store addresses
nak: wire up UGPR Ld/St/Atom encoding
nir: add uniform address to nvidia IO intrinsics
nak: add UGPR/GPR lowering for load/store/atom instructions
nak: optimize iadds with an uniform operand in iadds of address calculations
zink: proper advertise keep_weak_ffma for fp16
rusticl/kernel: handle nir shader compilation failures gracefully
rusticl: more intel compat stuff
rusticl/spirv: add SPIRVToNirOptions type
rusticl/spirv: silence GenericPointer cap warning
rusticl/spirv: properly set float execution mode at spirv_to_nir time
gallium: add fp16_no_denorms cap
nvk: enable VK_KHR_shader_fma
nir/opt_algebraic: add missing fmadz lowering for lower_fmulz_with_abs_min
nir/opt_dead_write_vars: cache is_entrypoint of the function
nir/lower_alu: fix lower_fminmax_signed_zero for denorms
asahi: fix dst range in buffer copy region
asahi: fix compute blitter for float16 image copies
asahi: move batch flushing into agx_launch_internal
asahi: fix fdiv lowering
meson: enable more rust 2024 lints
rusticl/util: add Traits to help with usage of CString
rusticl/kernel: store kernel names as CString
rusticl/program: store log as a CString
rusticl/program: wrap compiler option parsing
rusticl/util: fix rustc-1.95 compilation error
vtn/opencl: convert libclc workaround handling to a switch statement
vtn/opencl: fix edge case behavior for cospi
vtn/opencl: fix edge case behavior for sinpi
vtn/opencl: fix edge case behavior for tanpi
rusticl/program: print compiler output as Rust string
rusticl/util: add CStrExt trait
rusticl/util: add CStrExt::from_ptr_or_empty
rusticl/program: add CompileOptions::get_clang_args
rusticl/program: construct __OPENCL_VERSION__ inside CompileOptions::get_clang_args
rusticl/program: turn iter map into loop inside CompileOptions::new
rusticl/program: handle -create-library inside CompileOptions::new
rusticl/program: move -cl-std handling inside CompileOptions::get_clang_args
rusticl/program: set __OPENCL_C_VERSION__ ourselves
rusticl/program: store build options as CString
rusticl/program: implement CL_PROGRAM_BUILD_OPTIONS without a copy
rusticl/program: implement CL_PROGRAM_BUILD_LOG without a copy
spirv: set num_components for OpAtomicFlagTestAndSet
rusticl/kernel: override libclc shader config helpers
nak/instr_sched_prepass: Take predicate spilling into account when scheduling instrucitons
nak: normalize lop3 constant sources
nak: convert base to iadd for non-uniform ldcx lowering
nak: run nir_opt_constant_folding after nak_nir_lower_load_store
nir/opt_phi_precision: bail on load_const conversions between float and ints
rusticl: move the worker queue into the Platform
Reapply “rusticl: fix leak in `util_queue`”
rusticl/kernel: updated dim_threads in Kernel::suggest_local_size
rusticl/kernel: extract impl of suggest_local_size
rusticl/kernel: adjust grid at the end of suggest_local_size_impl
rusticl/kernel: add suggest_local_size tests
rusticl/kernel: add suggest_local_size_impl_gcd
rusticl/kernel: remove code to fill non full subgroups
rusticl/kernel: rework block size selection
gallium: remove PIPE_BARRIER_GLOBAL_BUFFER
rusticl/device: fix long vector_width queries on devices without int64 support
nak: implement shfr
nir: rework float compare late algebraic opts
nir: enable more opts for unordered and neo float compares
nak: implement and enable has_fneo_fcmpu
nak: implement ford and funord
nir/algebraic: pattern-match manual iadd64
brw: advertise fp64 fma on hw with fp64 support
anv: enable VK_KHR_shader_fma
anv: fix wrong rebase conflict resolution from VK_KHR_shader_fma MR
nak/sm20: fix immediate encoding for F2I and F2F
nak/hw_tests: add F2I test for NaN behavior
nir: use function foreach helpers inside nir_cleanup_functions
nir: add nir_shader_fully_linked helper
rusticl: return Result instead of Option from convert_spirv_to_nir
rusticl: abort compilation if the nir shader is not fully linked
gallium: add pipe_caps::hw_clear_buffer_sizes
rusticl/util: implement Debug for CLVec
rusticl/kernel: convert Queue parameter to Device in launch
rusticl/kernel: add interface to launch kernel with arguments without binding them
rusticl/kernel: add offset to buffer bindings
rusticl/meta: add builtin kernel support
rusticl/meta: add builtin kernels for buffer fills
rusticl/mem: make Image::fill return a closure
rusticl/mem: use meta for clEnqueueFillBuffer and clEnqueueSVMMemFill
rusticl/mem: implement 1Dbuffer fills on top of a plain buffer fill
asahi: update agx_get_cl_cts_version for submission 471
mesa_clc: support 32 bit targets
meson/rusticl: fix typo in depfile for builtin shaders
rusticl/meta: mark builtin kernels SPIR-V as 64 bit
rusticl/mesa: compile 32 bit version
rusticl/program: add Program::from_spirv_with_devs
rusticl/meta: support 32 bit devices
rusticl/meta: split out ulong kernels
rusticl/memory: return 0 for CL_IMAGE_SLICE_PITCH also for 2d images
rusticl/kernel: add libclc source hash to kernel shader keys
vtn/opencl: fix libclc needing fp16 lowering to fp32
nouveau: Fix return of dangling pointer in nouveau_fence_new
clc: make libclc optional for configs not needing it
clc: use our downstream fork of libclc
rusticl: warn if we load not our own libclc fork
Ken Cunningham (1):
llvmpipe: fix arch of LLVM JIT when cross compiling on Apple
Ken Xue (1):
radv: remove checking on the gralloc handle->numFds
Kenneth Graunke (113):
jay: Add missing ROR case
jay: Don’t forget UACCUM!
iris: Implement force_dual_color_blend_by_location via NIR
iris: Call elk_nir_lower_fs_outputs for Gen8 RT reads, not brw
nir: Set FRAG_RESULT_DUAL_SRC_BLEND in outputs_written when lowering
brw: Switch FS outputs to semantic IO and FRAG_RESULT_DUAL_SRC_BLEND
brw: Set prog_data::dual_src_blend from NIR outputs written bitfield
brw: Drop dead code from dispatch limit check for dual source blending
brw: Limit SIMD width based on NIR rather than first backend compile
nir: Allow bias for nir_texop_sparse_residency_intel
nir: Lower SSBO helper writes too
intel/nir: Only add an explicit LOD 0 when lod/bias don’t already exist
anv: Delete anv_instance::mesh_conv_prim_attrs_to_vert_attrs
anv: Use device->info.has_mesh_shading in key->mesh_input check
jay: Include depth and stencil on all MRT stores
jay: Add a TODO for coarse pixel shading
jay: Gripe more clearly about dual source blending
brw: Lower sample_pos for non-per-sample shaders in NIR
jay: Move render target store payload/descriptor construction to backend
jay: Implement fragment shader stencil writes
jay: Implement sample mask writes
jay: Add comments summarizing the PS thread payload layout
jay: Set Dispatch GRF Start Register in jay_setup_payload()
jay: Add a GPR_FROM_UGPRS opcode
jay: Implement sample position
jay, nir: Make a dispatch_mask_intel intrinsic
jay: Implement coverage mask
jay: Implement load_fs_config_intel
jay: Prohibit JAY_STRIDE_8 for EXPAND_QUAD
jay: Call constant folding before collecting FS outputs
jay: add a hack until we munge barycentrics dynamically
jay: Don’t skip sampler payload copies for 2 or fewer sources
anv: Drop TES dispatch mode asserts
brw: Fix URB read length for tessellation evaluation shaders
brw: Ensure entire input load fits in push data
brw: Refactor urb_read_length setting for TES
brw: Fix mistake in brw_nir_lower_deferred_urb_writes
brw: Fold constants after nir_lower_io for VS/GS/TES outputs
jay: Add URB load support
jay: Fix scratch surface address save/restore
jay: Remember sp_delta_B when rematerializing stack pointer lane 0
jay: Generalize EXTRACT_LAYER to take an arbitrary mask
jay: Handle facing that differs across subspans
jay: Implement viewport index FS input
jay: Fix null render target writes
jay: Drop render target stores with unconditional discards
jay: Implement dual color blending (but require SIMD16)
jay: Don’t swap FS interpolation .yz deltas
jay: Implement fragment shader barycentrics
jay: Pass proper simd_width to brw_nir_apply_key for fragment shaders
jay: Add tessellation evaluation shader support
jay: Add an INTEL_JAY=all option
jay: Unroll loops before lowering deferred URB writes
jay: Assert FS input deltas exist
jay: Fix hard coded number of FS inputs
jay: Ignore RT store condition if there are no outputs
jay: Improve unconditional discard removal
jay: Fix rewrite_without_flags for SEL with other flag sources
anv: Fix shader stats when using jay for non-compute stages
jay: Still predicate Null RT store if everything is discarded
jay: Implement load_subgroup_size
intel/nir: Improve address reuse in brw_nir_lower_immediate_offsets
intel/nir: Turn load_global_constant into load_global_intel too
jay: Call intel_nir_lower_shading_rate_output earlier
jay: implement load_frag_shading_rate
jay: Store the FS config def
jay: Store a test of the dynamic “is coarse?” FS config bit.
jay: Set prog_data->uses_fs_config when coarse pixel shading is dynamic
jay: Set the coarse pixel render target descriptor bit
jay: Don’t run the entire optimization loop before prog data
jay: Use nir_lower_frag_coord_to_pixel_coord
jay: Lower to pixel_coord_intel and frag_coord_w_rcp after prog data
jay: fix frag coord .z lowering with coarse pixel shading
jay: Allow BFN on U16 types
jay: Make a builder local in setup_fragment_payload
jay: Implement coarse pixel coordinate calculations
jay: Fix stack smashing with more than 16 FS inputs
jay: Rewrite FS output gathering
jay: Run brw_nir_lower_alpha_to_coverage earlier
jay: Optimize out noop samplemask writes
jay: Add missing HF conversion stride restrictions
jay: Tighten mixed stride restrictions
jay: Make a jay_clobbers_address_reg() helper
jay: Add a new VECTOR_EXTRACT opcode for indirect moves
jay: Implement indirect push constant loads for 32-bit
jay: Implement indirect push constant loads for 8/16-bit sizes
jay: Move brw_nir_apply_key call to be shared among stages
nir: Early out in nir_opt_shrink_vectors if all components are read
nir: Don’t shrink intrinsics and undefs to vec5s
jay: Speed up shuffles and vector extracts with uniform offsets
nir: Add an option for whether TCS invocation_id should be uniform
intel/compiler: Set nir_divergence_tcs_invocation_id_uniform
jay: Emit a noop URB write for EOT if there isn’t one to reuse
jay: Increase JAY_NUM_LAST_USE_BITS to 64
jay: Implement tessellation control shaders
jay: Use LOOP_ONCE if a loop ends in HALT too, not just BREAK
brw: Fix GS EOTs to not have an empty channel mask on LSC platforms
brw: Update comment that’s so old it makes no sense
brw: Inline brw_do_emit_fb_writes
brw: Assert that repclears aren’t used on Gfx12+
brw: Drop brw_compile_fs_params::allow_spilling
jay: Use INTEL_SIMD_DEBUG=cs for compute shaders, not fs
intel: Drop INTEL_DEBUG=no{8,16,32} flags
intel: Fix multipolygon flags in INTEL_SIMD_DEBUG default handling
intel: Refactor SIMD selection’s debug flag handling
intel: Replace INTEL_DEBUG=do32 with INTEL_SIMD_DEBUG
brw: Switch to INTEL_SIMD_FORCE for multipolygon modes
brw: Respect subgroup size requirements even with INTEL_SIMD_DEBUG
brw: Allow spilling and other poor decisions when using INTEL_SIMD_DEBUG
brw: Fix INTEL_SIMD_DEBUG=fs32 to work at all
brw: Rework FS SIMD selection to follow requirements over debug flags
jay: Add u16 and f16 support to CSEL
intel: Temporarily disable madvise on iris on xe.ko
Koch, Pawel (1):
Update docs regarding anv shader dumps Reviewed-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com>
Konstantin (3):
vulkan/cmd_queue: Handle struct copies that are not pointers
lavapipe: Re-emit push constants if the size changed
lavapipe: Fix push_constant_size for shader objects
Konstantin Seurer (64):
vulkan/radix_sort: Add support for 96-bit keys
vulkan: Rename radix_sort to radix_sort_u64
vulkan: Rename key_id_pair to key32_id_pair
vulkan: Implement 64-bit morton codes
radv/rt: Use 64-bit keys for gfx11-
util/u_trace: Add an option to emit additional code
util/u_trace: Rework resource management
util/u_trace: Print tracepoints with indentation
vulkan: Fixes for a spec update
vulkan,spirv: Update spec to 1.4.352
radv: Move debug options to radv_instance.h
radv: Move a whole bunch of debug/profiling related into a subdir
radv/tools: Rename radv_debug to radv_debug_hang
radv: Move radv_find_memory_index to radv_debug.c
radv: Add and use helpers for managing internal allocations
llvmpipe: Use i1 for sparce residency and expand as needed
llvmpipe: Implement sparse residency feedback for buffers
llvmpipe: Fix sparse binding large areas
lavapipe: Bump maxRayDispatchInvocationCount to the min requirement
nir: Duplicate the name in nir_def_set_name
llvmpipe: Remove lp_llvm_descriptor_base
lavapipe: Reduce descriptor sizes even further
lavapipe: Implement VK_KHR_shader_untyped_pointers
lavapipe: Add lvp_nir_lower_push_constants
llvmpipe: Fix memory leak when allocating sample functions
lavapipe: Perform shader object compatibility check early
lavapipe: Ignore src_plane for samplers
tools: Update imgui to the docking branch and add backends
meson: Add some include directories
vulkan/bvh: Add defines for acceleration structure types
radv: Use a separate BLAS pointer copy pass for BVH4
radv: Rename copy_blas_addrs to copy_addrs
radv: Rename serialization fields in radv_accel_struct_header
util: Add RTI file format definitions
radv/tools: Add RTI file dumping
rti: Initial commit
vulkan: Handle arbitrary build flag counts in vk_build_stage
vulkan: Make vk_build_stage non-static
vulkan: Filter for updates in vk_build_stage
vulkan: Move build_flags to vk_build_config
vulkan: Add vk_accel_struct_cmd_begin_debug_marker
vulkan: Use vk_build_stage for encode/update passes
vulkan: Shuffle around bvh build code
meson: Add mesa_python_path
vulkan: Move capture_key_pressed to vk_device
util/u_trace: Add an option for accumulating tracepoint ranges
util/u_trace: Do not generate empty structs
radv: Add u_trace support
util/u_trace: Include payload in the range accumulation key
radv: Include build_flags in the range key
util/u_trace: Release memory for reused timestamps
radv: Ignore entrypoints inside meta OPs for utrace
radv,anv: Enable BVH updates
tool: Rename RTI to gamma
radv: Fix generating ray history code
gamma: Set the window title to gamma
u_trace: Initialize fuzzy_* callbacks correctly
vulkan: Fix ROOT_FLAGS_OFFSET_ID offset
radv: Store root_flags for BVH8
radv: Use 64bit keys on GFX12
gamma: Do not render minimized viewports
gamma/radv: Display the ray launch ID
radv: Delay lowering printf
radv/bvh: Fix updating acceleration structures containing AABBs
Kovac, Krunoslav (3):
amd/vpelib: fix custom color space handling
amd/vpelib: Enable VPE cap for 3DLUT
amd/vpelib: Fixes for external lut compound
Lakshman Chandu Kondreddy (3):
zink: Query external memory handle type compatibility
freedreno: Add support for A704
zink: Set can_do_invalid_linear_modifier workaround for QCOM blob driver
Lars-Ivar Hesselberg Simonsen (9):
pan/genxml: Print shader hex in trace for Valhall
pan/va/disasm: Print 64 bit src/dest regs as reg pairs
pan/va/disasm: Align FAU printing
pan/va/disasm: Align indentation
panvk: Fix debug flag overlap
panvk/v10+: Align allocations >= 64k to 64k
panvk: Ensure 64k alignment for sparse images
pan/format: Prefer 16X16_BLOCK_U over INTERLEAVED_64K
panvk/v10+: Fix size gt -> gte for 64k alignment
Leandro Dorileo (1):
intel/executor: inform oa not available if that’s the case
Leder, Brendan Steve (Brendan) (1):
radeonsi/vpe: Update DCC API and programming
Lei Huang (1):
amd/virtio: enable Android amdgpu-virtio build option
Leon Perianu (1):
pvr: enable VK_EXT_device_memory_report
Lin, Ricky (1):
amd/vpelib: DPM detect first frame action
LingMan (6):
rusticl: Drop custom `addr` implementation
mr-label-maker: Add `Rust` label for `src/compiler/rust`
mr-label-maker: Drop rule applying `Rust` to all .rs files
mr-label-maker: Apply `Rust` label to `clippy.toml` and `build-rust.sh`
mr-label-maker: Apply `Rust` label to `rustfmt.toml`
mr-label-maker: Apply `Rust` label whenever crate dependencies are changed
Lionel Landwerlin (160):
intel/dev: fixup intel_needs_workaround() macro
anv: avoid C23
anv: fix compute push constant allocations on pre Gfx12.5 platforms
anv: fix invalid value for push block index
anv: fix debug printfs on hang
anv: fixup compute queue detection
anv: rework debug flag
anv: switch from INTEL_DEBUG to ANV_DEBUG for shader-print
anv: remove unused defines
anv: fix relocations into internal shaders
anv: simplify inline uniform descriptor loads
nir: expose nir_opt_dce_impl
anv: run a single impl loop for apply_pipeline_layout
anv/apply_layout: move some helpers around
ci/zink/intel: disable TGL demo-v2 trace
brw: track push constants shader stats
anv: promote push constant pointers to push buffers
anv: add a pass to realign global loads on DX CBV resources
intel/ci: update expectation for RPL
vulkan: add tracking for VK_EXT_primitive_restart_index
anv: implement VK_EXT_primitive_restart_index
anv: expose VK_KHR_shader_constant_data
anv: fix null pointer access
anv: stop using queue priority KHR aliases
anv: remove a bunch of KHR alias uses
anv/docs: update environment variable docs
anv: reorder debug options
anv: add a shader-dump debug option
imgui: update copy and port all tools using it
intel/tools: add eu stall viewer
brw/lower_texel_address: add heap support
anv: split sampler state packing from API object creation
brw: add heap support to brw_lower_storage_image
intel: add resource intrinsic support for heaps
anv: add lowering of descriptor heap intrinsics
anv: implement EXT_descriptor_heap entry points
anv: add descriptor heap binding support
docs: document ANV_DEBUG=desc-dirty
anv: enable EXT_descriptor_heap
anv: fix arc artifacts on Farming simulator 2022
anv: print out the content of the printf buffer at vkDestroyDevice
anv/brw/nir: fix wa_18019110168
anv: expose non binding-table/push-pointer flushing
anv: expose RT state flushing
anv: enable compute state flushing with indirect state
anv: move a bunch of structures to anv_types.h
anv/intel: add device generated commands shaders
anv/apply_layout: use the resource index to compute descriptor buffer addresses
anv: add apply_layout support for device bindable shaders/pipelines
anv: add a helper to flush the descriptors for indirect compute execution
anv: program relative push set offset for descriptor buffers device bindable shaders
vulkan: add pipeline helper to retrieve scratch-size/ray-queries
anv: add support for indirect execution set
anv: add indirect command layout support
anv: add unspecified internal kernel send count support
anv: allow simple shader spilling for complex ones
anv: enable generation shader calls
anv: handle descriptor binding with DGC
anv: implement generated preprocess & execute
anv: add barrier flags handling for preprocess buffers
anv: handle preprocess buffer creation on <= Gfx12.0
anv: track generated commands work with perfetto
anv: expose VK_EXT_device_generated_commands by default on Gfx12.5+
anv: add a device generated command debug option
anv: add Gfx9 support VK_EXT_device_generated_commands
anv: expose VK_KHR_maintenance11
anv: group all performance drirc together
anv/iris: stop using 3DSTATE_PUSH_CONSTANT_ALLOC_PS on Gfx12.5
anv: rename push constant allocation helper
anv: add an option to disable push constant space reallocation
vulkan/runtime: fix invalid address flags value for CmdCopyBufferToImage2
anv: fixup null address check
anv: implement VK_KHR_device_address_commands
anv: remove old entrypoints
anv: implement missing device image property compression filtering
vulkan/wsi: write VkImageCompressionControlEXT from swapchain to image creation
anv: enable VK_EXT_swapchain_compression_control when possible
blorp: stop requesting the fp64 shader for ELK
blorp: only request fp64 shader on when required
anv: sweep the NIR fp64 shader before keeping it on the device
anv: only load fp64 software shader when needed
anv: add an option to disable allocation over subscription
brw: simplify VF component packing code
anv: add SIMD32 requirement heuristic for Dragon Dogma 2
brw/jay: move some coarse lowering to NIR
brw/jay: move sample_mask_in handling to NIR
docs/features: updates for Anv
anv: temporarily reenable scratch page by default
anv: bump max compute workgroup count
anv: further optimize dirty state after secondary emission
anv: only reprogram line-stipple if enabled
util: add a script to auto-generate a drirc infrascture per driver
util/drirc_gen: enable validation for a specific driver
hasvk: rename a couple of drirc options
hasvk: add a driver section for drirc
drirc: remove non Anv option in the Anv section
anv: use the new generation script for drirc
anv: fix missing bindless flag hashing
anv: fix render target remapping tracking at the beginning of render passes
brw: avoid requiring a valid render target for empty fragment shaders
spirv: fixup infinite recursion with shader replacement
anv: use shader source hash rather than cmd_buffer fields
intel: switch shader hash to 64bit value
mi_builder: mi_umax2 tests
anv: rename drirc script
anv: move fake_sparse drirc to feature category
anv: move compression control drirc to feature section
anv: fake VK_EXT_image_compression_control on Xe2+
iris: only call brw_nir_fs_needs_null_rt() with no render targets
brw: fix null render target decision
anv: fix assert/crash in import of compressed local memory on xe2+
anv: align storage texel buffer support on image support
anv: add missing condition to update 3DSTATE_RASTER
anv: fix 3DSTATE_SF line width programming with Bresenham lines
anv: don’t forget dataport flush for ANV_DEBUG=dgc-dump
elk: assert always/never on some of the FS config flags
brw: remove always true condition
brw: remove interpolator coarse bit setting
brw/jay: track usage of fs_config by backend
brw: only check for shader_info::fs.uses_sample_shading
anv/brw/jay: de-dynamify per-sample interpolation
anv: hash binding tables for EXT_descriptor_heap too
brw: add shader key to enable robust SLM accesses
anv: add hitman2 workaround for SLM load vectorization
anv: fix descriptor heap indexing of YCbCr embedded samplers
vulkan/runtime: fixup group building with shaders from libraries
spirv: add parsing of vkd3d-proton shader hashes
drirc: add a callback mechanism do deal with shader hash & options
anv: enable VK_EXT_descriptor_heap by default
vulkan: condition cmd_queue initialization to driver need
anv: fix push constant address emission for gfx commands
anv: add missing handling of push pointers in gfx dgc
anv: use vkd3d-proton provided shader hashes if available
anv: add infrastructure to deal with missing barriers in applications
anv/brw: fixup 64bit array image accesses
anv: fix 64bit image atomic emulation with EXT_descriptor_heap
anv: more dgc push constant fix
anv: fix push pointer optimization with DGC
anv: add memory heap budget tracking across VkInstance
anv/brw: limit push constant promotion in vertex shaders
anv: introduce an option to disable disk cache
anv: fixup RT building barrier
anv: fix compression control reporting on xe2+
anv/ci: turn on astc emulation testing
anv: fill min_array_element with indirect descriptors
anv: fixup max push data delivered to shaders
anv: add workaround for atomics on R11G11B10 images
anv: remove previous Horizon Forbidden West workaround
anv: flush accumulated barriers for top of TOP_OF_PIPE
anv: fix Wa_18040903259
vulkan/runtime: fixup vk_shader leak on RT group recompile
anv: fix push buffer descriptor address relocation
brw: add missing INTEL_FS_CONFIG_PER_PRIMITIVE_REMAPPING handling
brw: fix wa_18019110168 lowering
anv: fix leak in RT binding point
anv: fix barrier for Wa_1508744258 / Wa_14024015672
iris: fix barrier for Wa_1508744258 / Wa_14024015672
jay: copy resource_intel surface handle value
anv: fixup the logic dealing with STATE_BYTE_STRIDE
anv: only consider active view-capable queues for image views
Lishin (3):
mesa/st: relax shader_has_one_variant checks for GLES2
broadcom/qpu: add V3D 7.1 disasm tests
v3d/v3dv: use common compute limits
Liu, Mengyang (2):
aco: fix broken VGPRs reservation for 64-bit attributes in VS prologs
amd: disable reset_filter_cam for mec
Lone_Wolf (2):
ac/llvm: fix build with LLVM 23 (MCSubtargetInfo)
clc: fix build with LLVM23 (TargetRegistry::lookupTarget)
Lorenzo Rossi (104):
nir: Extract float_is_half tests in common code
nir/opt_algebraic: optimize fadd/fmul with 16-bit source and constant
pan/compiler: Allow 16-bit alpha for atest_pan
pan/compiler: Fix WaRaR hazard in pressure scheduler
pan/compiler: Lower unaligned scratch memory accesses
pan/compiler: Handle ssbo_atomics in lower_vs_atomics
nir/lower_point_size: Handle 16-bit point sizes
nir/opt_sink: Add pan-specific load_input
panfrost: Constant-fold io locations after lowering
pan/compiler: Sort preprocess
panvk/jm: Fix tls_size overwrite in indirect draws
pan/compiler: Rework scratch memory strategy
panvk,panfrost: Pass inputs and info to postprocess
panvk: Remove pan_optimize_nir call
pan/compiler: Rename bifrost_optimize_nir
pan/compiler: Sort postprocess
pan/compiler: Collect nopersp varyings in lower_noperspective_fs
pan/compiler: Add better documentation for second lower_int64
pan/compiler/lower_fs_inputs: Do not trust slot->alu_type
panfrost: Split default key creation in helper function
panfrost: Plumb VS varying_layout in FS
pan/bi: Vectorize f2f16 on v10 and earlier
pan/bi: Switch old license texts to SPDX
pan/mid/fuse_io_cvt: Disable fusion on highp
pan/bi: Add nir_fuse_io pass
pan/bi: Add a printing helper for pan_varying_layout
pan/valhall: fuse_cmp skip when fusing the same instruction
pan/bifrost: Make CSE independent of liveliness labels
nir/opt_algebraic: Optimize mediump fadd/fmul done in highp
pan/bifrost: Fix 16-bit demote_if
kraid: Fix out-of-tree build issue
kraid/tests: Edit meson to help rust-analyzer provide IDE suggetsions
panfrost: Separate the compiler from libpanfrost
panfrost: Reorder meson definitions
compiler/rust/lower_bounded: Add FromIterator impl
kraid/swizzle: Add fold_u64
kraid/ir: Add SrcMod::fold_u64
kraid/ir: Add FauRef UserPage creation utilities
kraid: Add alloc_vec utility
kraid: Fix FauRef Display bug
kraid/ops: Add a small crate documentation for conventions
kraid: Add OpIMul
kraid/model: Ensure dyn Model is Send + Sync
kraid: Add hw_runner
kraid: Add basic hw_tests
kraid: Add Foldable and initial tests for OpShiftLop
kraid/hw_runner: Unmap buffers on drop
kraid: Sort opcodes and keep them sorted
kraid: Replace alloc_vec with alloc_ref
compiler/rust/float16: Implement total_cmp as present in f32
kraid/encode_v9: Implement accumulator ops for FCmp
kraid/ops: Fix Display for OpFCmp
kraid: Add tests for OpFCmp
kraid: Add Foldable impl for OpCSel
kraid/encode_v9: Implement accumulator ops in ICmp
kraid: Add tests for OpICmp
kraid: Add tests for OpIAdd
kraid: Add tests for OpIMul
kraid: Add OpISub with tests
kraid: Add OpClz with tests
kraid: Add OpIToF32
kraid/nir: Add IMul
kraid: Add ufind_msb
kraid: Add ShaderInfo
kraid: Add register preloading
kraid: Add load_push_constant
kraid/hw_tests: Use preloaded registers
pan,nir: Add Panfrost image intrinsics
pan/bi: Lower image load/store/lea in NIR and fix OOB access
pan/bi: Remove unused backend lowering
panfrost/model: Add var,cvt,sfu rates for Valhall architectures
panfrost/compiler: Properly compute Valhall ALU bound
drm-shim.py: Add more panfrost models
panfrost/drm-shim: Fix assertion on v12+
kraid: Legalize immediates
kraid: Add OpIAbs and plumb it through
kraid/nir: Fix nir_op_extract*16
kraid: Add BitRev and wire it up
kraid: Add OpPopCount and wire it up
kraid: Add support for clamp and round to OpFAdd
pan/bi: Add kraid-specific algebraic rules
kraid: Add F32ToI32 and wire it up
kraid/algebraic: Lower b2i conversions
kraid/algebraic: Lower nir_op_pack_uvecX_to_uint
kraid: Wire up fneg
kraid: Add OpFma and wire it up
kraid: Add OpFlush
kraid: Add FRound and wire it up
kraid: Add Frexp and wire it up
kraid: Add OpFMin/OpFMax and wire them up
kraid: Add OpLdExp and wire it up
kraid: Add a pass macro for validation and debug
kraid: dead-code elimination pass
pan/bi: Fix f2f16(a@16) in shader-db run on v13
pan/bi: Move pan_nir_fuse_io in bi_optimize_late
pan/nir_fuse_io_cvt: Add texture cvt fusion
pan/compiler: Don’t widen unaligned push-constants too much
panfrost: Unify FAU constants and relocation handling
panfrost: Promote constants to FAU
panfrost/midgard: Fix fau max not initialized
panfrost: Fix wrong layout reuse in user clip planes
pan/nir: Fix header static inline function
pan/compiler/stats: Fix ALU not being used in instruction bounds
pan/bi: Propagate swizzle in bi_optimizer_result_type
Louis Montagne (2):
zink: relax build-id length assertion for Mach-O
meson: allow DRI on darwin to enable Zink + EGL builds
Loïc Molinari (17):
pan/crc: Restrict CRC buffer creation to 1st RT mipmap level
pan/crc: Introduce pan_fb_info_is_fully_covered()
pan/crc: Check AFBC renderblock size on v5 and v6 too
pan/crc: Check CRC buffer validity and coverage on v5 and v6 too
pan/crc: Simplify CRC buffer selection logic
pan/crc: Use RT selection loop in single RT case
pan/crc: Check CRC requirements in dedicated function
pan/crc: Disallow CRC on sparse AFBC images
pan/crc: Simplify CRC buffer initialization
pan/crc: Cache temporary CRC info
pan/crc: allow setting a NULL pointer to the CRC validity state
pan/crc: Store CRC state in a struct
pan/crc: Enable CRC for multiple RTs on v6
pan/crc: Disable CRC on v4
pan/crc: Enable Empty Tile Elimination
pan/crc: Optimize clear color hashing
panfrost/ci: Mark “spec@!opengl 1.4@copy-pixels” as flake
Lucas Francisco Fryzek (1):
util/u_trace: Don’t use empty initializer list
Lucas Fryzek (1):
Modify x11_xcb_display_supports_xshm to get xshm opcode
Lucas Stach (4):
etnaviv: clean up index buffer handling code a bit
etnaviv: move index buffer handling in draw_vbo after derived state handling
etnaviv: reserve state emission space early in draw_vbo
etnaviv: move sampler source update before draw space reservation
Luigi Santivetti (3):
pvr: de-dup strncmp in pvrsrvkm winsys
pvr: add missing multi-arch support for pipeline exec and stats
pvr: re-use texture state words for each load op
Lukas Zapolskas (2):
pan/pps: Move PanfrostDevice to a separate file
pps: Add the Primitive, Instruction, Pixel and Fragment unit types
Maaz Mombasawala (3):
Revert “ci: vmware farm is offline, stop using it”
Revert “ci-farms/vmware: Disable vmware tests for now”
svga: Update CI expectations.
Marc Alcala Prieto (41):
pan/genxml: Add performance-trilinear enum values
pan/genxml: Add missing enum values on v9-v13
pan/genxml: Add v14 definition
pan/genxml: Implement RUN_FRAGMENT2
pan/decode: Remove progress-related decoding logic
pan/genxml: Build libpanfrost_decode for v14
pan/clc: Build for v14
pan/fb: Implement pan_emit_fb_desc for v14+
pan/desc: Implement pan_emit_fbd for v14+
pan/texture: Add v14+ YUV pipe format mappings
pan/format: Add v14+ YUV pipe format mappings
pan/afbc: Add v14+ AFBC YUV compression mappings
pan/afrc: Add v14+ AFRC YUV compression mappings
pan/lib: Build for v14
panvk: Implement RUN_FRAGMENT2
panvk: Handle provoking vertex and simultaneous reuse on v14
panvk: Build for v14
pan: Add v14 support
pan/va: Fix packing test for LdVarBufImmF16 on v11
pan/bi,va: Use dedicated LD_VAR_BUF_FLAT* opcodes on v14+
panfrost: Implement RUN_FRAGMENT2 on the Gallium driver
panfrost: Build the Gallium driver for v14
panfrost: Advertize Mali-G1-Pro support
docs/panfrost: Advertize Mali-G1-Pro support
pan/decode: Support INTERLEAVED_64K Z/S target dumps
pan: Layer offset is not longer available starting on v14
panvk/csf: Allow 256 layers per tiler descriptor on v14+
panfrost: Advertise Mali-G1-Premium and Mali-G1-Ultra support
panfrost: Remove duplicated flushes before RUN_FRAGMENT[2]
panvk/csf: Emit fragment layer state just before RUN_FRAGMENT2
panvk/csf: Implement incremental rendering on v14+
pan/csf: Fix incremental rendering on v14+
pan/bi: Load vertex view index from preload on v14+
panvk: Fix multiview support on v14+
pan: Add helper for max multiview view count and rise it to 16 on v14+
pan/compiler: Rename multiview to per_view_outputs
pan/va: Fix serialization of atomic operations using BI_ATOM_OPC_AUMIN
pan/va: Unit test BI_ATOM_OPC_AUMIN
pan/ci: Remove GLES shader image load/store atomic flake
panvk: Fix DRM format modifiers for multi-planar YUV formats
panvk/csf: Avoid poisoning read-only fragment SRs
Marek Olšák (167):
nir: add back color0/1 system values and VARYING_SLOT_PARAM_GEN_AMD
ac/nir: add ac_nir_get_io_driver_location as replacement for IO bases
ac,radeonsi: don’t use nir_intrinsic_base for FS outputs
radeonsi: don’t recompute IO bases for FS outputs
radeonsi: stop setting si_shader_info::output_semantic for FS
radeonsi: stop using si_shader_info::output_semantic for passthrough TCS
radeonsi: stop using output_semantic[] for LS outputs passed via VGPRs
radeonsi: remove si_shader_info::output_semantic[]
radeonsi: remove si_shader_info::num_outputs
ac,radeonsi: stop using nir_intrinsic_base for TCS inputs passed via VGPRs
ac/llvm: correctly load 16-bit TCS inputs from VGPRs and simplify
ac/llvm: reorder/remove variables in visit_load_input
radeonsi: update shader info in si_nir_lower_color_flatshade_twoside
ac,radv,radeonsi: don’t use nir_intrinsic_base for FS inputs
radv: remove radv_recompute_fs_input_bases
radeonsi: compute si_shader_info::color_attr_index without input_semantic[]
radeonsi: compute si_shader_info::inputs_read without input_semantic[]
radeonsi: remove si_shader_info::input_semantic[]
radeonsi: don’t call nir_recompute_io_bases for FS
radeonsi: set num_vs_inputs from nir->num_inputs and use it more
radeonsi: just get si_shader_info::num_inputs from NIR
amd: remove unnecessary and transitive #includes
ac/nir: add ac_nir_assign_fs_input_locations to set PS input locations in stone
nir/opt_licm: add a private state structure for the pass
nir/opt_licm: use nir_metadata_control_flow
nir/opt_licm: hoist instructions across multiple levels of nested loops
radeonsi/ci: remove the fixed XFB test from fails/flakes
radeonsi/ci/build: also fetch video decode/encode sample for VK CTS
nir/opt_dce: factor out dead instruction removal into a helper
nir/opt_dce: add shader_info::assert_inputs_not_dead
ac/nir: factor out ac_nir_lower_tex_coords from ac_nir_lower_image_tex
ac/nir/lower_tex_coords: move input loads instead of cloning them
aco/tests: update ACO tests for ac_nir_lower_tex_coords refactoring
glsl,gallium: add pipe_caps::glsl_bindless_handles_are_32bit
radeonsi: set glsl_bindless_handles_are_32bit
nir: add frag_coord_xy
nir/lower_wpos_ytransform: handle frag_coord_xy
nir/opt_frag_coord_to_pixel_coord: handle frag_coord_xy
nir: add direct lowered frag_coord building to replace lowering passes
nir: use nir_build_frag_coord everywhere
amd: add a tool that prints tiling layouts for all shim devices
winsys/amdgpu: revert invalid changes from CS functions
winsys/amdgpu: fix memory leaks when amdgpu_cs_create fails
amd/tools: rewrite ac_print_tiling_layouts to print all layouts, including XORs
nir/opt_licm: add filter callback
nir/tests: add nir_opt_licm tests
nir/licm: allow speculative hoisting across terminate if the filter is set
nir: add an option to ignore INTERP_MODE_NONE in nir_shader_gather_info
radeonsi: fix a typo in si_shader_update_spi_shader_formats
ac,radeonsi: add a helper to print PS input VGPR layout
ac,radeonsi: add helpers to print SPI_SHADER_COL/Z_FORMAT
aco,radeonsi: use enums for color barycentrics instead of input VGPR indices
radeonsi: remove dead get_frag_coord_from_pixel_coord optimization
radeonsi: use shader_info::fs::uses_sample_shading for ac_nir_lower_ps_early
ac: add ac_shader_args::line_stipple_tex_ena
radeonsi: move SI_SPI_PS_INPUT_ADDR_FOR_PROLOG into a helper function
aco,radeonsi: don’t forward LINE_STIPPLE_TEX_ENA VGPR from the PS prolog
aco,radeonsi: declare prolog CENTROID VGPRs only if used
radeonsi: declare prolog ANCILLARY & SAMPLE_COVERAGE VGPRs only if used
radeonsi: declare prolog LINE_STIPPLE_TEX_ENA VGPR only if needed
radeonsi: declare prolog LINEAR_SAMPLE/CENTER VGPRs only if used
radeonsi: simplify get_interp_info_from_input_load
radeonsi: stop using TGSI definitions for interpolation
radeonsi: handle any size of shader args in the LLVM PS prolog
nir/opt_move_to_top: add an option to exclude moving at_offset/at_sample loads
nir: generalize nir_vertex_divergence_analysis -> nir_custom_divergence_analysis
nir/opt_varyings: use workgroup divergence to identify convergent mesh outputs
util/set: add helper _mesa_set_equal
nir/opt_varyings: rewrite elimination of duplicated outputs
nir/tests: don’t leave “namespace {” unclosed in nir_opt_varyings_tests.h
nir/tests: test new output deduplication cases
nir/opt_varyings: always report progress when calling nir_remove_varying
nir/tests: use ASSERT_EQ instead of ASSERT_TRUE in nir_opt_varyings tests
nir: add missing SYSTEM_VALUE_FRAG_COORD_W_RCP
nir: handle load_frag_coord_w_rcp in multiple passes, same as non-rcp
nir: change nir_frag_coord_form options to a bitmask
nir: add nir_frag_coord_use_pixel_coord for OpenGL
radv: switch to nir_frag_coord_xy_z_w_separate with w_rcp
radeonsi: switch to nir_frag_coord_xy_z_w_separate with w_rcp
radeonsi: enable nir_frag_coord_use_pixel_coord
radeonsi: don’t treat sample_pos as using frag_coord
ac,radeonsi: remove all frag_coord_xy code
nir/opt_frag_coord_to_pixel_coord: factor out helper nir_all_uses_of_float_are_integer
radv: add a pass that selects either frag_coord_xy or pixel_coord, but not both
radv: remove dead load_sample_pos code
radv: move SPI_PS_INPUT_ENA emission into radv_emit_ps_state
radv: select frag_coord_xy and pixel_coord conditionally based on dynamic state
nir/opt_algebraic: add more ffract/ffloor/ftrunc/f2u/f2i patterns
radeonsi/tests: add an ordered append bandwidth test
radv: ignore color attachment samples for ps_iter_samples
ac/nir/lower_ps_early: remove obsolete comment
ac/nir/lower_ps_early: assume frag_coord_is_center is always true
ac/nir: add a new pass ac_nir_lower_sample_mask_in
radv: switch to ac_nir_lower_sample_mask_in
radv: enable SAMPLE_COVERAGE PS VGPR dynamically
radv: fix an inefficiency where the ANCILLARY PS VGPR was enabled but unused
radeonsi: use ac_nir_lower_sample_mask_in
ac/nir/lower_ps_early: remove now-unused lowering of sample_mask_in
radv: make RAST_SAMPLES_STATE dirty in CmdBeginRendering only on gfx12+
radv: emit_rast_samples_state uses uses_vrs_attachment only on gfx11+
radv: don’t leave SPI_PS_INPUT_ENA uninitialized with NULL PS to fix a hang
ac/surface: print the modifier in ac_surface_print_info
ac: add basic HTILE dword printing
radv: bump the sparse alignment requirement to 64K
radv: fix VK_MEMORY_PROPERTY_DEVICE_COHERENT_BIT_AMD with sparse buffers
radv,radeonsi: disallow VRS flat shading if SubgroupInvocationID is used
radv: rename vrs_coarse_shading -> vrs_flat_shading
nir/opt_idiv_const: a / uint_max -> b2i(a == uint_max)
radv: stop using set_sh_reg_idx(3) to reduce CP overhead
radv: fix setting COMPUTE_DISPATCH_INTERLEAVE on the gfx queue
radeonsi: remove unnecessary and indirect #includes
radeonsi/ci: allow glcts to be in the cts directory
radeonsi/ci: change DEQP_TARGET to default for Wayland
ac,radeonsi: remove uses_kernel_cu_mask and associated code
nir/opt_varyings: rewrite indirect IO tracking and dead IO elimination
nir/opt_varyings: split tidy_up_indirect_varyings
nir/opt_varyings: shrink pathological varying arrays to 1 element
nir: fix ibfe handling in ssa_def_bits_used
nir: extend ssa_def_bits_used to allow getting bits for any src component
nir: add a comp parameter into nir_def_bits_used
nir: change nir_all_uses_of_float_are_integer to return type masks and bits used
nir: change nir_def_bits_used to accept nir_scalar
nir/opt_idiv_const: a / b (where b > uint_max / 2) -> b2i(a >= b)
radv/lower_opt_fs_frag_pos: optimize f2u32(frag_coord_xy) & 0x1 to subgroup ops
radv: use quad_pos for pixel_coord conditionally based on dynamic state
radv: disable AMD_device_coherent_memory on gfx12 due to out of order behavior
ac/nir: fix incorrect upper bound for view_index
radv: lower view_index to a user SGPR instead of layer_id
radv: use PKT3_SET_SH_REG_PAIRS for setting multiple view_index SGPRs on gfx12
nir/tests: restructure opt_varyings_tests_bicm_sysval
nir/opt_varyings: move (c ? interp_input0 : interp_input1) into the prev shader
nir: add shader_info::sample_mask_in_declared because it has side effects
radv: fix a rare crash with NULL PS and force_vrs_per_vertex
radv: use radeon_opt_set_context_reg for PA_CL_VRS_CNTL to fix random behavior
radv: use VRS flat shading even if other VRS state is enabled
radv: remove the VRS rate output if VRS flat shading overrides it
radv: cosmetic VRS changes
radv: don’t use PS_ITER_SAMPLE to force VRS 1x1, use SC/DB VRS override instead
radv: remove the VRS rate output if VRS is force-disabled by FS
radv: don’t execute pre-rast shader info code for FS
radv: remove no-op code from radv_consider_force_vrs for POPS
radv: disallow force_vrs_per_vertex with FragCoord when using GPL & ESO
radv: move force_vrs_per_vertex to emit_fsr_state to make it robust (rewrite)
radv: remove redundant PA_CL_VRS_CNTL setting from the initial state
radv: reduce duplication in gfx103_emit_vrs_override_state
radv: disable the VRS image on gfx11.x if the VRS rate is overridden
radv: fold gfx103_pipeline_vrs_flat_shading into its only use
radv: always set EN_VRS_RATE=1 because GE_VRS_RATE can also disable it
radv: inline radv_is_vrs_enabled
radv: use shader_info::fs::sample_mask_in_declared
radv: ignore VRS state for sample_mask_in lowering and optimizations
radv: take sample_mask_in_declared into account when lowering to pixel_coord
radv: set key.ps.force_vrs_enabled and key.vrs_may_be_enabled more accurately
bin/drm-shim: forward the error code from the command to the user
radv: fix low pixel throughput with NULL DS on GFX11.x
radv: don’t set DB_Z_INFO.NUM_SAMPLES = 3 on gfx12
radv/nir_trim_fs_color_exports: use a state structure to pass parameters
radv/nir_trim_fs_color_exports: remove mrt0.w if alpha_to_one makes it dead
ac: fix a GPU hang with LLVM due to incorrect VGPRS decoding of LLVM output
ac/llvm: rename ac_parse_shader_binary_config -> ac_parse_llvm_binary_config
radeonsi: use wave64_vgpr_encode_granularity
radv: don’t expose memory types from AMD_device_coherent_memory without the ext
radv: fix determining the raster prim for guardband
radv: fix determining the raster prim for line mode
radv: fix determining the dynamic raster prim for FS barycentrics
radv: fix determining the static raster prim for FS barycentrics and front_face
nir/opt_varyings: fix incorrect counting of emit_vertex within a block
Mario Kleiner (22):
wsi/display: Expose VK_FORMAT_B8G8R8A8_UNORM before VK_FORMAT_B8G8R8A8_SRGB
wsi/display: Improve connector->last_nsec timestamping.
wsi/display: Add workaround for all-zero valued pageflip events.
wsi/display: Deal with vblank-less systems for VK_EXT_present_timing.
wsi/common: Small compliance fixes for VK_EXT_present_timing.
wsi/common: Allow VK_EXT_present_timing present without presentStageQueries.
wsi/common: Allow to return queue_done_time in host time domain.
wsi/wayland: Unconditionally assign present_timing.time_domain.
wsi/common: Add VK_GOOGLE_display_timing support for KHR_display.
wsi: Don’t try to create a timestamp query pool without driver support.
vulkan/wsi: Optionally expose VK_GOOGLE_display_timing on wsi wayland+x11.
wsi/wayland: Always use clock monotonic domain for GOOGLE_display_timing.
vulkan/wsi: Add hk, nvk as VK_GOOGLE_display_timing supported drivers.
docs/features: Add missing VK_EXT_present_timing enabled for X11.
wsi/display: Actually fix vblank-less systems for VK_EXT_present_timing.
pvr: Expose VK_KHR_present_id and VK_KHR_present_wait.
hasvk: Expose VK_KHR_present_id2 and VK_KHR_present_wait2.
hasvk: Expose VK_KHR_calibrated_timestamps.
hasvk: Expose VK_EXT_present_timing and VK_GOOGLE_display_timing.
wsi/display: Don’t update connector last_frame/nsec in vkGetSwapchainCounterEXT.
wsi/x11: Skip next_present_ust_lower_bound assignment in certain FRR mode.
wsi/x11: Refine VRR vs. FRR detection a bit.
Martin Roukala (né Peres) (18):
zink/ci: mark blender-demo-cube_diorama as flaky on gfx1201
turnip/ci: document recent flakes
ci: disable the valve-kws farm
Revert “ci: disable the valve-kws farm”
freedreno/ci: reduce the parallelism of the a750-vk job
freedreno/ci: document more failures for the a750-gl-cl job
radv/ci: reduce parallelism for radv-gfx1201-vkcts
radv/ci: document more flakes
radeonsi/ci: document new flakes
amd/ci: tighten the timeouts of the Valve jobs
zink/ci: document a recent regression on navi10
zink/ci: document more flakes
zink/ci: tighten the timeouts of the valve jobs
nvk/ci: tighten the timeouts of the valve jobs
radv/ci: bump the timeout of the valve vkd3d-asan jobs
ci: allow controlling which hw test jobs to create at pipeline creation
panfrost/ci: turn bifrost / valhall rules into per-kernel driver
radv/ci: document another WSI flake in radv-renoir-vkcts-full
Mary Guillemard (45):
nvk: Use SET_REFERENCE in nvk_CmdResetQueryPool
nvk: use MME shadow RAM in nvk_meta begin/end
nvk: Move nv_push closer to their uses in nvk_cmd_begin_end_query
nvk: Clear counters at the begin of a query
nvk: Remove delta handling from query pool
nvk: Conditionally enable counters when needed
nvk: Move report offset to reports_start for nvk_CmdCopyQueryPoolResults
nvk: Handle zero queries in CmdCopyQueryPoolResults and CmdResetQueryPool
nvk: Store available and timestamps packed together
nir/lower_bit_size: Preserve float controls when lowering alu ops
nvk: Handle foreign queue dependencies
nvk: Handle host accesses barrier
nvk: Multiply by local_size for CS invocations in DGC codepath
nak: Allow YY swizzle for SM20 and SM32 asserts
nir/nir_format_convert: Add missing u2f32 in nir_format_unpack_r9g9b9e5
nir,nak: Add match_any_nv
nak: Add a lowering pass for shared memory atomics in mesh stages
nvk: Prepare nvk_shader for GS header upload for mesh shaders
nvk: Prepare cbuf for mesh shader support
nvk: Add support for mesh and task shader binding
nvk: Implement mesh draw commands
nak: Implement mesh and task shader stages
nvk: Do not set lower_cs_local_index_to_id
nvk: Only lower shared memory for compute shaders
nvk: Lower mesh and task shaders
nvk: Advertises VK_EXT_mesh_shader
docs/nvk: Add some notes about mesh shading and ISBE layout
nvk: Do not report task and mesh stages as supported on pre-Turing
nvk/nvkmd: Do not merge bind operations across VA mappings
nvk: Implement support for non graphics timestamp
nouveau/mme: Add some simple MME shadow RAM dumper
nouveau/mme: Add a test for MME Shadow RAM behavior
nvk: Increase maxStorageBufferRange and maxBufferSize
nvk: Default to output primitives as lines for tesselation parameters
nvk/ci: Update expectations and document failures
nvk: Use I2M in CmdUpdateBuffer when possible
panvk: Split cmd_prepare_push_uniforms logic
drm-shim/nouveau: Report proper values in DRM_NOUVEAU_GET_ZCULL_INFO
drm-shim/nouveau: Stop using nouveau gallium names for classes
drm-shim/nouveau: Add Ada A to Blackwell B support
nvk: add a build option to override the build ID
nvk: Only increment CS counters when query is active in CmdDispatchBase
nvk: Do not enable remap in nvk_copy_indirect
nvk: Do not take base into account when lowering emulated attributes
nvk: Reenable compression support on Turing with nouveau 1.4.3
Matt Turner (15):
intel/elk: Remove some dead code
intel/elk: Remove dead TXL_LZ/TXF_LZ opcodes
radv: fix UB in radv_format_pack_clear_color for snorm formats
radv/perfcounter: guard select1 access in radv_emit_select
radv/perfcounter: add GFX11 performance counter selectors
radv: expose VK_KHR_performance_query on GFX11
util, llvmpipe: flush subnormals to zero on ARM/AArch64
nir: fix dedup_entry memcmp on structs with padding
gallivm: fix lp_build_round on altivec/VSX
gallivm: fix small_unorm -> unorm8 fetch path on big-endian
nir/tests: allow relative error in compare_inexact
nir: fix f2u/f2i constant folding to poison NaN and out-of-range inputs
nir: use i2f32 for patterns with signed-extraction opcodes
nir: fix pack_uvec4_to_uint to mask input components to 8 bits
nir/tests: fall back to integer comparison when float interpretation is NaN
Matthieu Oechslin (7):
r600: Fix crash on R600/R700 with custom border color
r600: Improve and document R600_TRACE
r600: Stop emitting relocs with virtual address enabled
r600: Workaround GPU hang with compute shaders when VA is enbaled
r600: Calculate address at emit time for SSBOs
r600: Fix MSAA 2D view from array with VA enabled
r600: Document RADEON_VA and remove SB options references
Mauro Rossi (5):
radv: Fix gnu-empty-initializer errors in 480a94fb
radv: Fix gnu-empty-initializer errors in 8c10eab1
radv: Fix gnu-empty-initializer errors in ca9191a8
intel/common: remove fallthrough annotation in unreachable code
pan/perf: fix building error due to ‘Mali G1.xml’ file name with space
Maíra Canal (2):
etnaviv/ml: derive stride-2 destriding offsets from padding
v3dv: Drop legacy comments about single-sync support
Mel Henning (36):
nak: Use shader_info->var_copies_lowered
nak: Use NIR_LOOP_PASS
nvk: Split out nvk_cmd_fill_memory
nvk: Allocate a zcull save region in fewer cases
nvk: Zero zcull data in layout transition
nvk: Don’t LOAD_ZCULL w/ VK_RENDERING_RESUMING_BIT
nvk: Re-enable zcull save/restore
nvk: Add a wfi for blackwell in CmdDispatchIndirect
nvk: Disable compression on Turing
compiler/rust: Fix inline wrapper include dir
nak/nvdisasm_tests: Fix expected value of F16v2
nak: Fix encoding of f16x2 min/max on sm90+
vk/meta: Move get_uint_format_for_blk_size to common
nvk: Make nvk_cmd_buffer_queue_flags non-static
nvk: Split out aligned_for_linear_attachment
nvk: Use meta for image copies where possible
nil: Pass ImageDim to Tiling::choose()
nil: Pick tiling params closer to proprietary
nvk: Interp frag_coord at centroid for min_sample_shading
nvk: Fix DGC localsize computation
nvk: Serialize shaders with asm
nvk/meta: Rename begin/end with a _gfx suffix
nvk/meta: Implement save/restore for compute
nvk/meta: Add save_generic helpers
nvk: Use compute meta for some vkCmdCopyImage2
nvk: Use meta for vkCmdCopyImageToBuffer2
nvk: Use meta for vkCmdCopyBufferToImage2
nvk: Use meta for vkCmdCopyBuffer2
nvk: Add _ce suffix to nvk_cmd_fill_memory
nvk: Use meta for vkCmdFillBuffer
nvk: Handle large indirect stride pre-Turing
nvk: Prepare indirect draws for 64-bit stride
nvk: Convert draw/dispatch to device_address_commands
nvk: Don’t re-align ssbo size/address
nvk: Move ssbo_4b_align to drirc
nvk: Use ?: in nvk_physical_device_compiler_flags
Michael Cheng (11):
intel/ds: Add end_event_dyn() and CREATE_DUAL_EVENT_CALLBACK_DYN macro
intel/ds: Label compute events with dispatch dimensions in Perfetto
intel/ds: Label selected draw events with vertex count
brw: Fix ordered dependency exec_all handling on Xe2+
intel/brw: allow baking more SBID dependencies into instructions on Xe2+
intel/brw: Don’t bake a long-pipe RegDist with an SBID dependency
intel/brw: Factor out combinable_ordered_pipe() helper
nir/opt_gcm: add option to keep texture ops in large loops
intel/brw: keep texture ops in large loops
intel: Fix DEBUG_FS_SIMD mask
intel: Fix operator precedence in intel_simd_overridden
Michal Krol (7):
gallium: add pipe_sampler_view::min_lod_clamp
lavapipe: implement VK_EXT_image_view_min_lod with fractional minLod
gallivm/llvmpipe: fix VK_EXT_image_view_min_lod via texture handle path
lavapipe: lower array-deref-of-vec for mesh shader outputs
lavapipe: fix format properties for R10X6G10X6B10X6A10X6_UNORM_4PACK16
gallivm: honour exec mask in EmitMeshTasksEXT
gallivm: don’t deref a NULL buffer descriptor with an empty exec mask
Michel Dänzer (14):
winsys/amdgpu: Use render node only as fallback
mr-label-maker: Label src/gallium/winsys/amdgpu as radeonsi
mr-label-maker: Label src/gallium/winsys/radeon as r300, r600 & radeonsi
egl/gbm: Do not destroy BO of current front buffer
egl/gbm: Use local variable for better readability
egl/gbm: Eliminate max_age local variable
egl/gbm: Ignore buffers with no BO for destroying excess BOs
egl/gbm: Ignore current front buffer in get_back_bo
egl/gbm: Use local variable for better readability in get_back_bo
egl/gbm: Eliminate local variable “age” in get_back_bo
egl/gbm: Use continue instead of nested block
egl/gbm: Eliminate local variable “max_age” in get_back_bo
dri3: Increment draw->send_sbc after waiting for last presentation
dri3: Simplify target_msc calculation in loader_dri3_swap_buffers_msc
Mike Blumenkrantz (104):
lavapipe: KHR_device_address_commands
radv: add RADV_QUEUE_DISABLE env var for selectively disabling queues
llvmpipe: fix min_samples + A2C
lavapipe: fix indirect memory copies
lavapipe: fix pushconst data updating
lavapipe: null out local var to avoid uninit warning
util/format: support 256-bit formats in util_format_get_tilesize()
lavapipe: use the right type for DGC mesh draws
lavapipe: rework immutable samplers
lavapipe: allow fbfetch with shader objects
vk/cmd_queue: always ceil() param lens
vulkan: update spec to 1.4.350
lavapipe: maintenance11
llvmpipe: always set view_index for linear rasterizer
llvmpipe: unify setting raster_state for thread data
lavapipe: update cbuf count when remapping attachments
lavapipe: unset attachment remap state if pColorAttachmentLocations==NULL
lavapipe: fix setting colormasks when attachments get remapped
ci: stop skipping HIC tests on lavapipe
zink: use maintenance5 to more effectively set storage texel usage for bufferviews
zink: delete zink_resource_object::storage_buffer
zink: remove remaining maint5 checks
aux/trace: silence -Waddress warnings in macros
zink: delete unused descriptor variable
zink: fix mixing of mesh descriptor bindings with gfx bindings
meson: fix renderdoc integration define
vulkan: move vk_shader_stages_from_bind_point() to vk_util
zink: disable implicit sync handling for qcom proprietary
zink: rework custom sample locations
lavapipe: enable some forgotten ds3 states
zink: fix unbinding vertex buffers from null VS state
zink: add another anv/adl flake
zink: create views for samplers lazily
lavapipe: correctly disable depth/stencil in secondaries
vk/cmd_queue: simplify gross struct duplication
zink: use custom sample locations to (mostly) handle multisample=disabled
zink: link up vs COLx vars -> fs BFCx
zink: be more conservative about query pool sizing
lavapipe: stop using pipeline layouts in some places
lavapipe: Implement VK_EXT_descriptor_heap
zink: handle uint wrapping with batch submit count
zink/bo: reduce wasted memory due to the size tolerance in pb_cache
zink/bo: add an enum to disambiguate bo types
zink/bo: stop using pb_buffer vtable for destroy
zink/bo: use only a single layer of slabs
zink/bo: use pb_buffer_lean to save a little mem
zink/bo: check for usage before completion when reclaiming bos
zink: use maint11 for sso shader object compile
zink/clear: fix full_clear condition in texture clear
zink/clear: handle texture clears on current fb texture
llvmpipe: create a zeroed payload for use without task shaders
zink: always return DMA_BUF type handles from resource_get_handle
zink: tag tc info update in a few more places
util/tc: iterate the rp info more accurately during batch execution
aux/tc: enforce strict resolve semantics
vulkan/wsi: pass VkSurfaceCapabilities2KHR to get_capabilities
vulkan/wsi: add VK_IMAGE_CREATE_MULTISAMPLED_RENDER_TO_SINGLE_SAMPLED_BIT_EXT where supported
lavapipe: EXT_multisampled_render_to_swapchain
zink: fix import2d sampler view creation
zink: when triggering zink_blit_barriers() for src==dst, apply separate barriers
zink: stop forcing barriers if previous access was write
zink: properly invalidate fb attachments on dontcare stores
zink: proactively apply transfer sync when tracking renderpasses
zink: don’t invalidate cbufs without inlined resolve
zink: add some ci flakes
tu: handle partially set resolve attachment info without crashing
util/tc: store resolve geometry to rp info
zink: use tc info to handle partial resolves
util/tc: unset TC_RESOLVE_STRICT
zink: set NO_TASK_SHADER for pipeline layouts with shader objects
zink: a618 ci updates
zink: always use src stages when flushing glMemoryBarrier calls
zink: always flush specified memory access for glMemoryBarrier calls
zink: reset usage following SHADER_WRITE access
zink: drop imageless_framebuffer requirement
zink: add a vb param to vertex buffer binding
zink: move vb binding out of c++
zink: hook up VK_KHR_device_address_commands
zink: use DAC for vertex binding
zink: fix a missing case of zink_batch_submit_count_diff()
zink: stop unsetting resource usage on batch reset
zink: free nir if cs program create fails
zink: enable signed vbs
st/pbo_compute: account for drivers failing to create cs shaders
zink: split more read/write barriers
zink: stop adding usage with last-ref tracking
zink: unset unordered access on ordered transfer ops
zink: use bigger hammer to force sync between unordered->main cmdbufs
zink: add api for disabling reordered read/write
zink: unset ordered_access_is_copied when disabling unordered access
zink: noop per-resource synchronization for unordered->ordered access
gallium/cso: make unbind_context an explicit call
lavapipe: handle depth blit aspect masking
st/context: unbind gs shader before deleting hw select gs shaders
zink: always un-suspend queries on end
lavapipe: advertise dynamicRenderingLocalReadDepthStencilAttachments
zink: fix the fix for ZINK_RENDERDOC=all
zink: don’t increment unique_id for reused bos
zink: translate depth write ALWAYS to GE/LE if possible
util/blitter: fix blitting multiple array layers
zink: start ZINK_DEBUG=perfinfo
zink: revert cached mem handling for staging uploads
zink: add anv ci flake
zink/ci: switch zink/anv jobs to surfaceless+i915
Mohamed Ahmed (8):
nil/modifiers: Clarify drm_format_mods_for_format rejecting modifiers for unsupported color formats
nvk: Calculate and stash the plane offset and alignment at create time
nvk: Extend tiled_shadow to be multiplanar
nvk: Defer tiled shadow plane memory allocation to draw time
nvk: Enable multiplanar YCbCr linear modifiers
nvk: Use the pre-calculated offsets for sparse binds
nvk: Remove nvk_image_plane_size_align_B()
nil: enable PLC for compressed data
Nanley Chery (23):
intel/blorp: Halve max bpp for some redescribed blits
anv: Add transfer_src usage for ANDROID_external_format_resolve
anv: Improve the fast clear layout perf-warn
anv: Improve the CCS_E-incompatible perf-warn
anv: Avoid aux-disabling paths for block-compression
anv: Allow CCS on more storage images for gfx12.5
anv: Move storage check out of CCS-compat helper
anv: Flush previous aux-mode changes
intel/isl: Define a CMF for ASTC formats
intel/isl: Fix the initial state HiZ state for Xe2+
anv: Dedent a closing curly brace
anv: Skip some CCS performance warnings on gfx9-11
anv: Allow partial depth fast clears on gfx12+
anv: Set TRANSFER_DST_BIT for HiZ operations
intel: Add and use ISL_AUX_USAGE_ZCS
hasvk: Delete enum anv_depth_reg_mode
iris: Rework HiZ plane optimization disabling
iris: Disable HiZ planes for some read-only tests
anv: Track the depth buffer aux usage
anv: Enable overriding HiZ in depth stencil state
anv: Rework HiZ plane optimization disabling
anv: Bypass HiZ planes for read-only depth tests
anv: Drop the anv_disable_hiz drirc option
Natalie Vock (9):
radv/rt: Don’t overwrite bvh_base at the start of the traversal loop
radv: Dump printf buffer after detecting a GPU hang
radv/rt: Cache stack sizes of ahit/isec shaders from imported NIR
radv: Work around midpoint sorting issues instead of disabling
radv: Fix destroying address binding reports
mailmap: Update my email
nir/opt_loop: Don’t peel header blocks that jump
radv/nir: Clean up descriptor index lowering
radv: Expose mutable acceleration structure descriptors
Nataraj Deshpande (1):
intel/perf: map ray tracing counters to RAYTRACING group
Neha Bhende (1):
svga: fix shared memory index for svga driver
Nemallapudi, Jaikrishna (1):
intel/dev: fix timebase_scale ticks-to-ns precision loss across 2^32
Nick Hamilton (5):
pco: fix clamping the array index when shaderImageGatherExtended is enabled
pvr: Enable shaderImageGatherExtended
pvr: Revert don’t csb emit multi-layer clear attachments without rta support
pvr: Fix load-op shader when loading from a 2d image view of a 3d image
glthread: fix check for unroll draws using user VBOs when the ctx supports GLES
Okenczyc, Andrzej (1):
amd/vpelib: Report unsupported status if streams target rect equals 0
Olivia Lee (23):
pan/bi: fix memory access alignment
pan/genxml: add definitions for adjacency draw modes
panvk: add support for adjacency primitive topologies
panvk: fix executable properties handling for IDVS varying shaders
pan/csf: rename immediate CS add builder functions
pan/v13: add CS builder functions for reg/reg add and sub instructions
pan/v13: add CS builder functions for shift instructions
pan/v13: implement constant integer multiplication CS helper
pan/v13: implement CS udiv
panvk: remove redundant invalid primitive topology cases
panvk/csf: allow SYNC_WAIT-style synchronization in launch_gfx_cs
panvk: add create_shader helper for compiling full meta shaders
panvk: return both gpu and cpu pointers from cmd_prepare_*_push_uniforms
poly: allow VS outputs with <32 bits per component
poly: refactor GS lowering to store output and selected variables together
poly: preserve src_type in GS rast shader outputs
poly: preserve output component counts in GS
poly: move hk passthrough GS code to libpoly
poly: allow specifying output types in passthrough GS
nir: add nir_slot_num_components helper
poly: allow specifying component count in passthrough GS key
poly: clarify assertion failure message in lower_store_to_var
panvk/csf: flush primitives generated query writes from CSF
Omar Rashwan (2):
intel: Fix bit width of int literal in eu stall viewer
intel: define type for std::max in eu stall viewer
Patrick Lerda (19):
r600: refactor r600_shader_buffer_info_sel
r600: refactor eg_setup_buffer_constants
r600: rename sh_txs_cube_array_comp to sh_resinfo_via_uniform
r600: cypress resinfo buffer size workaround
r600: fix alpha-to-coverage and alpha-to-one used together
r600: add sample_lz and sample_c_lz opcodes compatibility
r600: update r600 nir for sample_lz and sample_c_lz
r600: enable EXT_texture_shadow_lod
docs/features: add GL_EXT_texture_shadow_lod
r600: remove r600_get_hw_atomic_count
r600: fix atomic buffer offset
r600: implement tes and tcs instanced gl_PrimitiveID support
r600: make r600_copy_region_with_blit global
r600: implement msaa 2d view from array
r600: update vertex emit_varying_pos
r600: fix atomic_counter_post_dec
r600: update memory barrier operations
i915: fix emit_hw_vertex() unbounded memory access
r600: update muladd support configuration
Paulo Zanoni (30):
intel/isl: fix assert when surf->size_B is > UINT_MAX
intel/isl: warn about excessive num_elements only once
anv: don’t silently convert view ranges from u64 to u32 then u64
docs/envvars: remove ANV_SPARSE and ANV_SPARSE_USE_TRTT
docs/envvars: document ANV_SYS_MEM_LIMIT
docs/envvars: update the ANV_DEBUG documentation
intel/nir: fix sparse shadow comparison for BRW
anv/sparse: bring back our (limited) support for depth/stencil
brw: evict memory for workgroup scope in Xe2 and newer
intel/mi_builder: add mi_ixor()
intel/mi_builder: add mi_umax2()
intel/blorp: prepare for usage of mi_builder.h
libcl/vk: add aligned(4) to VkCopyMemoryIndirectCommandKHR
libcl/vk: add VkCopyMemoryToImageIndirectCommandKHR and its members
anv: implement VK_KHR_copy_memory_indirect
intel/tools: fix stall_csv_filename maybe-unitialized error
intel/brw: move cache_mode assignment to after send->sfid choice
brw: split cache mode selection into atomic, load and store modes
brw: have a single if-ladder to pick cache_modes
brw: control cache_mode through bypass_{l1,l3} variables
intel/blorp: don’t silently ignore compilation failures
intel/blorp: fix blorp base key initialization
intel/blorp: move struct blorp_blit_prog_key to blorp_blit.c
intel/blorp: don’t include “util/format_rgb9e5.h”
anv: give anv_ensure_fp64_shader() a chance to be called
brw: don’t preprocess software doubles if opts->softfp64 is not set
anv: don’t put clear colors for aliased images in private bindings
intel/blorp: rearrange struct blorp_blit_prog_key
intel/blorp: pack every blorp key struct
intel/blorp: memset(0) blorp keys during initialization
Pavel Ondračka (111):
r300,i915/ci: update expectations
r300/ci: update expectations
i915/ci: update expectations
r300: fix MSAA resolve COLORPITCH tiling after pipe_surface de-pointerization
r300: dirty VS state when switching variants
dri3: add big-endian 8888 fourccs to dri3_cpp_for_fourcc
dri: add big-endian 8888 entries to dri2_format_table
dri: add big-endian 8888 entries to driImageFormatToSizedInternalGLFormat
nir: fix partial loop unroll OOB check for loops not starting at 0
nir/tests: add helpers for counting used/unused instructions
nir/tests: add partial unroll OOB tests
r300: drop unused input arrays from ntr
r300: drop multiple ubo support from ntr
r300: drop GS/tess and load_draw_id support from ntr
r300: drop framebuffer fetch handling from ntr
r300: drop GLSL 4.x texture ops from ntr
r300: drop unsupported sampler dimensions from ntr
r300: drop the i915g vertex_id/instance_id U2F branch from ntr
r300: drop GLSL 4.x interpolation intrinsics from ntr
r300: drop opcode paths lowered before emission from ntr
r300: drop TEX2/TXB2/TXL2 dead path from ntr
r300/ci: run EGL deqp tests
r300/ci: update expectations
i915/ci: update expectations
r300: remove unused LIT opcode
r300: pack immediates more aggressively to avoid running out of constant slots
r300: reuse positive and negative immediate values
r300: remove extra newline for compiler errors
r300: remove the redundant control flow checks
r300: stop dumping TGSI
r300: always convert to NIR and move ntr later
r300: collect input/output info directly from NIR
r300: move compiler init earlier and to a helper
r300: fix use-after-free of remap_table in rc_remove_unused_constants
r300: build the dummy fragment shader directly as NIR
r300: emit one full vec4 immediate per NIR load_const
r300: move r300_transform_*_trig_input out of nir_to_rc
r300: extract TGSI->RC translation helpers into nir_to_rc.h
r300: emit RC instructions directly from nir_to_rc
r300: drop the TGSI opcode middle step from ntr
r300: allocate FS outputs from NIR locations
r300: use NIR varying locations directly in ntr
r300: stop declaring samplers with ureg in ntr
r300: lower sysvals to varyings
r300: use backend texture targets directly in ntr
r300: drop ureg shader properties in ntr
r300: drop ureg_DECL_temporary in ntr
r300: drop ureg_DECL_vs_input in ntr
r300: drop dead tg4_offsets and query_levels paths in ntr
r300: drop ureg_DECL_address
r300: get rid of user_src and ureg_dst
r300: stop using ureg_dst_undef in ntr
r300: get rid of ureg_program in ntr
r300: get rid of ureg_src_undef in ntr
r300: get rid of ntr_emit_load_output
r300: get rid of the precise modifier
r300: use RC registers directly in nir_to_rc
r300: drop more dead ntr code
nir/algebraic: prevent ffract optimization on lowered ffloor
r300: fix R300_VAP_TCL_BYPASS state leak
r300: fix vs->first leak in swtcl delete path
r300: add NIR pass to append the wpos output
r300: add NIR pass to add required color outputs
r300: convert swtcl vertex shader setup to NIR
r300: remove dead first-time build path from r300_pick_vertex_shader
r300: clean up some dead draw/TGSI leftovers
r300: remove draw support on big endian
meson: require r300 LLVM draw only on x86
gallivm: add NIR pass to lower load_ubo_vec4 to load_ubo
gallivm: add no_integers intrinsic fixup pass
gallivm: add NIR pass to lower float if conditions
gallivm: add algebraic NIR pass for no_integers
gallivm: pass base NIR ALU types to cast_type
r300: enable VS instance ID in draw
r300: prepare for the the draw NIR path
draw: use gallivm NIR for no_integers vertex shaders
i915: use gallivm NIR for vertex shaders
r300: use signed index offset for index translation
r300: always use 32-bit indices on big endian
r300: use R32_FLOAT as 32-bit dummy vertex format
r300: fix occlusion query results on big endian
r300: fix BE 8888 render-to-texture endian state
r300: fix BE RGB565/RGB5 render-to-texture formats
r300: fix BE constant blend color for colorbuffer formats
r300: fix BE depth/stencil raw transfer endian state
r300: clean up endian swap selection
r300: add NIR LICM pass to enable removal of residual loops
i915: add NIR LICM pass to enable removal of residual loops
r300: remove dead state constants
r300: add private NIR state constant plumbing
r300: handle texture destination constraints in nir_to_rc translation
r300: lower backend texture coordinate handling in NIR
r300: lower fragment position in NIR
r300: lower VS seq/sne in NIR
r300: lower FS alpha-to-one in NIR instead of backend
r300: handle FS depth output channel mapping at nir_to_rc time
r300: lower r300 FS derivative stubs in r300_optimize_nir
r300: move r500 FS derivative fixup to nir_to_rc translation
r300: extend the wined3d A0 rounding pattern recognition
r300: some post-int/bool lowering optimizations
r300: don’t split ALU instructions on R5xx
r300: run rgb alpha conversion after RC optimize
r300: penalize presubtract NOP in pair scheduling
r300: run late CSE after lowering vectors to registers
r300: fix swtcl per-vertex point size
r300: remove redundant FACE input workaround
r300: keep NIR output count consistent
r300: unify WPOS output handling between swtcl and hwtcl
r300: share vertex shader variant handling in hwtcl and swtcl paths
r300: emulate gl_FrontFacing on R3xx/R4xx
nv50_ir_ra: align B96 spill slots to vec4
Peng_Lx (1):
turnip/kgsl: close the dma-buf fd of our own allocations
Peyton Lee (14):
amd/vpelib: add alpha fill support check
amd/vpelib: Support vpe 2.0
amd/gmlib: add tm_generate_formatted_3DLut
radeonsi/vpe: add VPE 2.0 support
frontends/va: add ABGR format mappings
amd: validate and expose VPE 2.0.0
radeonsi: gate format and rotate/flip support by VPE version
amd/vpelib: support vpe 2.2
ac/gpu_info: add VPE_2_2 support
radeonsi/vpe: adjust message
amd/vpelib: tighten external LUT compound color pipeline updates
amd/vpelib: refine coding style
amd/vpelib: Fix Color Corruption AV1 Issue
amd/vpelib: Replace hardcoded bg format table size with sizeof
Philipp Zabel (2):
etnaviv/isa: Fix Meson warning about etnaviv_isa_rs dummy library
meson: add rusticl to with_driver_using_cl
Physics Enthusiast (1):
venus: allow to use vtest as a fallback for virtgpu
Pierre-Eric Pelloux-Prayer (40):
radeonsi: clamp cp prefetch size
ac/info: add gfx12.1 identification
radeonsi/tests: update expectations
amd/virtio: use AMDGPU_VA_MGR_RESERVE_HALF_VA_FOR_PRT
amd/virtio: fix amdgpu_sw_info_address_prt_wa_control_bit handling
radeonsi: handle NULL return value from amdgpu_cs
gallium/vl: only release created sampler views
radeonsi: delay aux context initialization to first use
radeonsi: add has_gfx_compute property to si_screen
radeonsi: don’t use staging texture when we can’t blit
radeonsi/vce: deal with has_gfx_compute being false
radeonsi: create a mm subfolder for multimedia code
radeonsi: add si_init_screen_nir_options
radeonsi: add gfx subfolder
radeonsi: move shader cache code to new file
radeonsi: extract si_init_gfx_caps from si_init_screen_caps
radeonsi: add si_resource_copy_buffer
radeonsi/gfx: add si_gfx_screen.c
radeonsi/gfx: move code from si_get to si_gfx_screen
radeonsi: add si_gfx_context.c and move code from si_pipe.c
radeonsi: add si_context.c
radeonsi: move all multimedia files to mm
radeonsi: move more code to gfx subfolder
radeonsi: move function prototypes from si_pipe.h to si_gfx.h
radeonsi/gfx: remove unnecessary u_stub usage
radeonsi/gfx: move static inline helpers to si_gfx.h
radeonsi: add tests subfolder and move AMD_TEST code inside
gallium/u_blitter: remove unused skip_viewport_restore
radeonsi: fix sqtt setup
radeonsi/sqtt: hash only the relevant part of the shader key
radv, radeonsi: do sqtt buffer_size calc using uint64
ac/sqtt: add ac_sqtt_update_bo_size
radeonsi: fix sdma copy for gfx10
radeonsi: consolidate aux context creation into si_get_aux_context
radeonsi: use aux context locks in si_destroy_screen
ac/parse_ib: initialize data variables to 0
radeonsi: delay si_disk_create_cache call
radv/rra: bump rt_driver_interface_version
radv: restore RRA capture support
radeonsi: fix typo in si_copy_from_staging_texture
Pohsiang (John) Hsu (15):
mediafoundation: periodic clang-format
mediafoundation: code clean up
meson: Make with_gfx_compute depend on video encode support (mediafoundation)
d3d12: add av1 handling to d3d12_video_encoder_get_encode_headers and d3d12_video_encoder_update_current_encoder_config_state_av1
mediafoundation: extract code to ProcessDX12EncodeContext
mediafoundation: fix a few minor variant bool handling
mediafoundation: detach xThreadProc frame processing from apiLock to unblock concurrentt ProcessOutput calls
d3d12: fix infinite gop handling in d3d12_video_enc_av1.cpp
mediafoundation: define AVC_LOG2_MAX_FRAME_NUM_MINUS4, HEVC_LOG2_MAX_PIC_ORDER_CNT_LSB_MINUS4 instead of using number.
mediafoundation: preserve low latency ping pong behavior between ProcessInput and ProcessOutput
mediafoundation: change default value for HEVC_LOG2_MAX_PIC_ORDER_CNT_LSB_MINUS4 to 12
mediafoundation: initial av1 dx12 hmft prototype
d3d12: add support to output temporal delimiter for AV1 via raw_header
mediafoundation: ask for temporal delimiter for AV1 via raw header
d3d12: fix msvc build warning C4819
Qiang Yu (2):
ac,radeonsi,radv: fix print IB assertion fail for reserved fields
ac,radeonsi,radv: use V_581A_* engine sel for non-pws acquire_mem packet
QwertyChouskie (3):
docs/features: Remove mentions of r300 and nv30
docs/features: Fix typo
docs/features: Mark VK_EXT_descriptor_heap done for anv
Radu Costas (6):
pco: Set register classes for vec refs
pco: Move preproc_vecs out of loop
pco: Add debug variables for RA
pco: Move RA context handling to state-based
pco: Try allocating with optimal temp registers
pvr, ci: Update axe and bxs failure list
Rahul Mahantappa Bhadrashette (1):
radeonsi: emit MESHLET registers for mesh shaders in gfx10_emit_shader_ngg
Raviraj Uppal (2):
driconf: disable allow_rgb16_configs for SPECviewperf
radv: app workaround implemented using internal layers for GFXBench 5.0
Rhys Perry (122):
ac/nir_lower_global_access: perform range analysis if useful
ac: add gfx11.7 enums
aco/gfx11.7: add opcode numbers
aco: adjust some gfx_level checks for gfx11.7
aco/gfx11.7: don’t create v_dot2c_f32_f16
aco/gfx11.7: don’t use v_pack_b32_f16 in do_pack_2x16
aco/gfx11.7: allow any src VGPR for VOPD with two v_dual_mov_B32
ac/gpu_info/gfx11.7: enable has_point_sample_accel
aco/gfx11.7: claim support
radv/gfx11.7: take GFX12 paths in radv_nir_lower_cooperative_matrix
radv/gfx11.7: enable float8
radv/gfx11.7: enable shaderMixedFloatDotProductFloat8AccFloat32
radv/gfx11.7: don’t advertise shaderImageFloat32AtomicMinMax
aco: refactor spiller to use spills_needed variable
aco: prefer spilling smaller temporaries if it finishes spilling
aco: use RegisterDemand::operator[] more
radv: move ac_nir_lower_indirect_derefs to end of radv_shader_spirv_to_nir
radv: lower indirect derefs after linking
radv: don’t use radv_optimize_nir after lowering indirect derefs for RT
ac: move lds_size_per_workgroup to ac_compiler_info
ac: move has_cs_regalloc_hang_bug to ac_compiler_info
radv: assert there is no padding in cache keys
radv: initialize nir_shader_compiler_options directly in compiler info
radv: move load_grid_size_from_user_sgpr to radv_physical_device
radv: move use_llvm to radv_compiler_info::key
radv: move fields to radv_compiler_info::key
radv: add fields to radv_compiler_info from radv_physical_device_cache_key
radv: remove radv_compiler_info::cache_key
radv: hash radv_compiler_info::key into the cache key
radv: remove most fields from radv_physical_device_cache_key
radv: remove radv_physical_device_cache_key
radv: remove radv_device_cache_key
ac/llvm: fix isub image atomic
radv: inline shader_compile()
radv: remove radv_aco_convert_opts
radv: replace radv_nir_compiler_options with a LLVM one
radv: split radv_compiler_info’s family into debug::family and key::family
aco/gfx11.7: fix v_pk_min_f16/v_pk_max_f16 opcode numbers
nir/search: fix nir_algebraic_automaton after constant folding op(bcsel)
nir: rename nir_src_parent_instr to nir_src_use_instr
radv,ac: make rembrandt and vangogh cache compatible
radv: don’t pass GPU name to disk_cache_create
radv: remove family from cache key
aco/validate: fix some RA validator error messages
aco/ra: fix v3b VALU at byte>0
aco/ra: test the register file in get_reg_specified() when necessary
aco: add helpers to get instruction subdword capabilities
aco: rework subdword definition RA validation a bit
aco/ra: fix compact_relocate_vars path for get_reg_for_operand
aco/ra: fix fill() with certain subdword cases
aco: fix regclasses for spill/reload subdword temporaries
aco/ra: don’t rename phi operands in get_reg_phi()
aco/ra: use phi_dummy instead of is_phi()
aco/ra: remove precolored checks in get_reg_impl()
nir/algebraic: optimize ishl(iadd(ishl, ishl))
nir/algebraic: optimize ishl(iadd(iadd(iadd(a, #b), c), d), #e)
nir: make cmat_muladd_amd a subgroup intrinsic
nir: add load_deref_transpose_amd
nir: add load_global_transpose_amd
nir,ac/nir,aco: add load_global_tr_amd
radv: track cooperative matrix robustness
radv: use load_deref_transpose_amd for transposed cooperative matrix loads
radv: fix usage of radv_nir_cmat_length
aco: add cost estimation of s_barrier
aco: don’t emit workgroup-scope p_barrier for single-wave workgroups
aco/waitcnt: always use uint32_t for event masks
aco: fix printing of primitive exports
aco: optimize redundant s_wait_alu vm_vsrc(0) during waitcnt insertion
aco: only assume load/store with semantic_atomic is atomic
aco: don’t emit waitcnts before subgroup-scope execution barriers
aco: add split barrier instructions
aco: use split barrier instructions
aco: schedule split barriers
ac/gpu_info: add has_smem_partial_oob_access_bug
radv: workaround has_smem_partial_oob_access_bug
nir/opt_undef: fix prefer_nan
aco: don’t increase barrier exec scope to subgroup
radv/bvh: use atomic load/store in update_gfx12.comp
ac/lower_global_access: set cursor earlier
ac/lower_global_access: rewrite try_extract_additions
ac/lower_global_access: parse u2u64 even if *out_offset!=NULL
ac/lower_global_access: extract constants after ishl/imul
ac/lower_global_access: combine multiple 32-bit offsets
radv: use radv_shader_stage_key::keep_{statistic,executable}_info more
radv: move nir_debug_info from debug to key
radv: don’t create nir_string if dump_shader=true
radv: simplify radv_declare_shader_args parameters
radv: add radv_shader_stage_key::keep_shader_arg_info
radv: cache shader IR, asm and spir-v
radv: parse stats from binary in radv_parse_binary_debug_info
radv: make raytracing radv_shader_stage_key array per-shader
radv: make raytracing radv_shader_stage_key initialization per-shader
radv: merge radv_shader_stage_key for combined ahit/isec shaders
radv: rework creation of traveral radv_shader_stage_key
radv: inline some helpers used in radv_pipeline_get_shader_key
radv: use nir_opt_uub
radv: repeat loop in radv_optimize_nir_algebraic_early more
radv: do nir_opt_algebraic last in radv_optimize_nir_algebraic_early
radv: fix barriers in decompress shaders
drm-shim: implement most readlink() without initializing the shim
drm-shim: skip init_shim() if drm_shim_fd_lookup() would be NULL
nir: add NIR_MEMORY_CONTROL_ARRIVE and NIR_MEMORY_CONTROL_WAIT
nir: add nir_lower_disordered_control_barriers
aco: implement NIR_MEMORY_CONTROL_ARRIVE and NIR_MEMORY_CONTROL_WAIT
vtn: implement SplitBarrierEXT
radv: implement VK_EXT_shader_split_barrier
radv: do radv_parse_binary_debug_info in radv_shader_dump_asm
ac/nir: skip SMEM fixup for more 32-bit load_global addresses
radv: set RADV_CMD_DIRTY_GFX12_HIZ_WA_STATE around attachment clears
vtn,nir: print shader filename and spec constants
radv: include debug information in bvh shaders
radv: shorten internal shader names
vulkan/bvh: fix to_emulated_float(-0.0)
vulkan/bvh: don’t update min/max_bounds with inactive nodes
vulkan/bvh: enable SignedZeroInfNanPreserve in leaf.h
vulkan/bvh: limit valid 32-bit node keys to 0xfffffe00
ac/nir/ngg: don’t shrink device-scope memory barriers
ac/nir/ngg: track uniformity with multiple set_vertex_and_primitive_count
nir/load_store_vectorize: rework barriers to use entries
nir/load_store_vectorize: recreate entry key when adding from predecessor
vtn: don’t fail at uniformity decorations on variables and OpBufferPointer
radv: drop support for cooperativeMatrixRobustBufferAccess
Rob Clark (65):
tu: Remove use of fd_perfcntr_type
freedreno: Remove use of fd_perfcntr_type/result_type
freedreno/perfcntr: Remove type and result_type
freedreno/registers: Add json to describe perfctr groups
freedreno/registers: Generate perfcntr tables
freedreno/perfcntrs: Switch to generated perfcntr tables
freedreno/registers: Small reg32 vs reg64 fixes
freedreno/registers: Sync back xml changes from kernel
freedreno/registers: Add pipe to perfcntr group
freedreno/registers: Add gen8 perfcntr support
freedreno/registers: Correct register name
freedreno/registers: Add gen8 perfcntrs
pps: Re-emit time clock_sync more regularly
freedreno/ds: Use gpu timestamps
freedreno/ds: Split a6xx/a7xx counters out
tu: Fix preemption latency selector values
freedreno/a6xx: Expose subgroup ops
rusticl: Support/ignore -qcom-accelerate-16-bit
freedreno/common: Fix X2-90, add X2-85
freedreno/registers: Skip deprecated warns for kernel
freedreno/registers: Add a6xx CMP counter group
freedreno/registers: Gen8 perfcntr fixes
drm-uapi: Sync msm_drm.h
freedreno/common: Add ioctl ptr helpers
freedreno/fdperf: Move where we setup counter groups
freedreno/fdperf: Prepare for partial-counter usage
freedreno/fdperf: Add PERFCNTR_CONFIG support
freedreno/ds: PERFCNTR_CONFIG support
freedreno/ds: Add a8xx derived counters
freedreno/perfcntrs: Add helpers to resolve group and countable
freedreno/perfcntrs: Add helper to assign counters
tu: Use counter allocation helper
freedreno/a6xx: Use counter allocation helper
freedreno/perfcntrs: Refactor derived counter setup
freedreno/perfcntrs: Use helper for derived counters
freedreno: Skip BV perfcntrs
tu: Disable preemption for counters on gen8
tu/gen8: Program slice selector regs
freedreno/a6xx: Program gen8+ slice SEL regs
freedreno/perfcntrs: Expose gen8 counters
freedreno/a6xx: Push RB_A2D_PIXEL_CNTL magic into blitter
freedreno/registers: Improve A2D docs
freedreno/registers: Add RB_RESOLVE_CNTL_0.YUV_PLANE_ID
freedreno/a6xx: Un-open-code RB_A2D_PIXEL_CNTL
tu: Un-open-code RB_A2D_PIXEL_CNTL
perfetto: Add API to flush track events
util/thread: Flush traces at thread exit
util/queue: Flush perfetto before blocking
rusticl: Flush perfetto track events
perfetto: Increase SMB size
perfetto: Use BufferExhaustedPolicy::kStall
freedreno/perfetto: Add non-draw stage
freedreno/perfetto: serialize clk snapshots
freedreno/perfetto: Use sequence-scoped clk
freedreno/crashdec: Update gpu revision parsing
freedreno/crashdec: Add additional HFI queue
freedreno/a6xx: Don’t clamp 32b clear values
freedreno/a6xx: Fallback for conditional blits
freedreno/a6xx: Handle R9G9B9E5 blits as R32_UINT
freedreno/a6xx: Set HALF_PRECISION for R11G11B10_FLOAT
freedreno/a6xx: Don’t forget UBO driver params
freedreno/decode: Fix shader stats in summary mode
freedreno/ci: Update a660-vk-traces-restricted checksums
nir/convert_address_format: Split convert_def into two passes
nir/convert_address_format: Handle non-deref sources
Rob Herring (Arm) (34):
ethosu: Make quantization shift signed
ethosu: Add a common initializer for struct ethosu_operation
ethosu: Store ethosu_tensor struct ptr in feature map
ethosu: Move stride calculation to lowering
ethosu: Fix concatenation OFM scaling
ethosu: Support axis 1 concatention
ethosu: Add fully-connected operation
ethosu: Rename ethosu_lower_add to ethosu_lower_eltwise
ethosu: Support element wise op with constant IFM buffer
teflon: Add multiply operation
ethosu: Add multiply operation support
teflon: Add TANH operation support
ethosu: Add logistic and TANH operations
teflon: Add hard swish operation
ethosu: Add hard swish operation
teflon: Add LeakyRelu operation
ethosu: Add LeakyRelu operation
teflon: Add quantize operation
ethosu: Add quantize operation
ethosu: Add reshape operation
teflon: Add minimum and maximum operations
ethosu: Add minimum and maximum operators
ethosu: Add performance counter debug output
teflon: Ensure all TfLiteRegistration fields are 0
meson: Skip NIR tests with headers-only NIR
teflon/tests: Use reference kernels
ethosu: Compute elementwise broadcasts from OFM shape
ethosu: Fix depthwise conv layout for IFM depth 1
ethosu: Flatten fully connected inputs
ethosu: Fix command dependency tracking
ethosu: Improve equal-cost block selection
ethosu: Use full scale for NOP pooling
ethosu: Preserve spatial dimensions for FC lowering
ethosu: Preserve fused pad extents
Robert Mazur (6):
ci: update firmware tag to ff46ce35
ci: update kernel tag to v6.19-mesa-712d
imagination/ci: Use standard CI-tron gfx-ci/linux kernel
pvr: switch core count mesa_logw() to pvr_finishme()
pvr: introduce PVR_IGNORE_FINISHME_WARNINGS envvar
pvr/ci: enable PVR_IGNORE_FINISHME_WARNINGS
Rohit Athavale (1):
mediafoundation: Test compile steps v/s step , and set build flag
Roland Scheidegger (1):
gallivm: fix subtle filtering issue with different min/mag filter for cube maps
Roman Stratiienko (3):
v3dv/android: Add deferred ANB allocation support
v3dv: move noop_job creation to device scope
v3dv: Emulate multi-queue support via vk_queue for Android
Romaric Jodin (1):
anv: Declare 00-mesa-defaults.conf as an input to anv_dricrc_gen.py
Rouf, Farhan (4):
amd/vpelib: Introduced reset to frontend
amd/vpelib: Changed cmd_info input for background segment
amd/vpelib: Chroma coefficient select for sampled formats corrected
amd/vpelib: Refactoring Reset Pipes Function
Rudraksha Gupta (1):
freedreno: add Adreno 225
Ryan Houdek (1):
turnip: Add an override to uncached memory type
Ryan Mckeever (2):
pan/bi: check if preds are dominated by header in bi_find_loop_blocks
docs: advertise VK_KHR_multiview support for Bifrost
Ryan Neph (1):
anv/xe: prevent WaitIdle optimization for fences with exported sync_fd
Ryan Zhang (4):
panvk: add VK_IMAGE_LAYOUT_DEPTH_READ_ONLY_OPTIMAL to host copy layouts
gfxstream/platform: add missing inc_include to platform_virtgpu build
panvk: set cfg cull status according to primitive topology
panvk: Drop empty SYNC_ONLY bind queue ops
Saeed, Ghamr (1):
amd/vpelib: variable was accumulating size and not reset properly
Sagar Ghuge (49):
anv/rt: Copy 16bytes at once instead of copying 8bytes
anv: Fix Wa_14021821874, Wa_14018813551, Wa_14026600921
brw: Pass write back register for ray query messages
intel/genxml: Update xml for dynamic stack ID control fields
anv: Enable dynamic stack ID control on Xe3+
intel/genxml: Disable compute walker mid-thread preemption
util: Increase array size to 20
intel/genxml: Added dispatch timeout counter extended field
anv: Update values for DispatchTimeoutCounter
anv: Set execution mask based on SIMD size
brw/rt: Commit hit even if we are skipping closest hit shader
brw/rt: Update committed hit leaf type properly
brw/rt: Use BLAS(Object) level to get the ray address
jay: Implement halt
anv/rt: Skip invalid node in child block count
anv: Pass vk_acceleration_structure_build_state as param
anv/rt: Extract common code in separate header
anv/rt: Use constant BVH offset instead of pushing
anv: Track parent-child map for BVH update
anv: Track leaf block offset map
intel: Add debug option to dump out parent-child map
anv: Implement update BVH
intel: Add debug hook to dump out BVH after update
ci-farms/vmware: Disable vmware tests for now
anv: Allocate lookup maps for update based on mode and flag
intel: Add drirc option to write lookup maps unconditionally
Revert “anv: Fix Wa_14021821874, Wa_14018813551, Wa_14026600921”
anv: Workaround game bug for Witcher3
jay: Extend CS payload to handle BTD stack IDs
jay: Handle nir_intrinsic_load_btd_stack_id_intel intrinsic
jay: Factor out RT message header build part
jay: Implement btd_retire instrinsic
jay: Implement BTD Spawn intrinsic
anv: Compile init RT shader with Jay
anv: Bump subgroup size for histogram and prefix shader
vulkan: use center/extent form for instance node AABB transform
jay: Control cache_mode through bypass_{l1,l3} variables
jay: Setup bindless thread payload
jay: Handle nir_intrinsic_load_btd_global_arg_addr_intel
jay: Handle nir_intrinsic_load_btd_local_arg_addr_intel
jay: Set simd width for bindless shader
jay: Handle EOT for RT shader
jay: Init header with zero for all components
jay: Stuff stackIDs for trace_ray message
brw: Track if CS uses fences
intel: Fix async compute thread limit
anv: No need to flush RT cache if we update buffer via CS
jay: Track max stack size for bindless shaders
jay: Add loop_once_halt opcode
Sahitya Kandru (1):
freedreno: Modify reg_size_vec4 for a608 and a612 to 32
Sam James (4):
glx: append extra_ld_args_libgl, not clobber
src: add -Wl,–no-fatal-rwx-sections for two libraries
gallium/dri: fix redundant Meson condition
util: remove bogus const attribute
Samuel Pitoiset (263):
vulkan: add an option to lower SHADER_RECORD_INDEX to non-uniform
radv: lower SHADER_RECORD_INDEX to non-uniform
radv/ci: document some HIC failures since addrlib uprev for GFX11.7
radv: add enable_mrt_output_nan_fixup to the physical cache key
ac/surface: add stencil-only support for host mem->surf copies
radv: add depth+stencil formats support with host image copy
radv: allow depth+stencil formats with host image copy
amd: allow addrlib to enable SIMD if possible
radv: advertise VK_EXT_host_image_copy by default on GFX10.3+
radv/ci: document more HIC regressions on NAVI10
vulkan: refactor vk_pipeline_robustness_state_fill() slightly
vulkan: pre-compute the default robustness state in the device
vulkan,treewide: stop passing vk_device to vk_pipeline_robustness_state_fill()
spirv,treewide: rework specialization constant
radv: fix GPU hangs with PS epilogs and secondaries properly
radv/rt: pass more parameters to radv_rt_nir_to_asm()
radv: add a radv_compiler_info object
radv: use radv_compiler_info everywhere during compilation
spirv: add support for SPV_KHR_constant_data
radv: advertise VK_KHR_shader_constant_data
radv: move queue related cmd buffer state to a new struct
radv: move uses_perf_counters to radv_cmd_buffer_queue_state
radv: move shader_upload_seq to radv_cmd_buffer_queue_state
radv: remove redundant initialization when beginning a cmdbuf
radv: zero-initialize radv_cmd_state only when a cmdbuf is reset
radv: pass radv_compiler_info to radv_pipeline_get_shader_key()
radv: store the number of PS params heuristic to radv_compiler_info
radv: re-introduce DGC+multiview support and enable it for vkd3d-proton only
radv: fix a potential NULL pointer dereference when emitting VBOs
radv: remove an useless check when emitting the index buffer
radv: only emit the “normal” index buffer when needed with DGC
radv: stop dirtying some states after DGC execute
radv: cleanup invalidating vertex draw state
radv: replace use_ngg_streamout by gfx_level checks
ac,radv,radeonsi: replace mesh_fast_launch_2 by gfx_level checks
vulkan: add missing VkMemoryRangeBarriersInfoKHR support
radv: add missing VkMemoryRangeBarriersInfoKHR from DAC
radv: simplify resetting pipeline state for ESO
radv: rename RADV_CMD_DIRTY_PIPELINE to RADV_CMD_DIRTY_GRAPHICS_PIPELINE
radv: stop tracking the last emitted graphics pipeline
radv: add RADV_CMD_DIRTY_COMPUTE_PIPELINE
radv: add RADV_CMD_DIRTY_RAY_TRACING_PIPELINE
radv: remove useless tracking about non-coherent RBs with secondaries
radv: slightly rework initializing the default graphics state
ci: bump libdrm to 2.4.133
meson: bump required libdrm to 2.4.133 for AMDGPU
radv: move suspend_streamout to radv_streamout_state
radv: move streamout bindings to radv_streamout_state
radv: remove unnecessary radv_cmd_state::mesh_shading
radv: move index buffer state to radv_index_buffer_state
radv: cleanup suspending/resuming cond rendering with DGC
radv: move conditional rendering state to radv_cond_render_state
radv: move vertex buffer state to radv_cmd_state
radv: re-organize radv_cmd_state slightly
radv/ci: bump timeouts for radv-{navi21,gfx1201}-vkcts-full
ac/gpu_info: store more addr space info
ac/gpu_info: add has_smem_with_null_prt_bug
ac/gpu_info: query the PRT workaround control bit from libdrm
ac/nir: add a pass to fixup SMEM loads with NULL PRT pages
radv: run the pass to fixup SMEM loads with NULL PRT pages
radv: use the “LOW” address space for UBOs
radv/amdgpu: emulate sparse residency for the SMEM loads with NULL PRT workaround
radv: set RADEON_FLAG_EMULATE_SPARSE_RESIDENCY for sparse SSBO/UBO buffers
radv/ci: update list of skipped tests
ci: uprev vkd3d
radv: fix printing image format with RADV_DEBUG=img
radv/meta: fix expanding HTILE on compute with multisampling
docs: describe the contributions workflow for RADV
radv: bump VkConformanceVersion to 1.4.5.3
radv: fix determining needed dynamic states when rasterization is disabled
radv: make optimalTilingLayoutUUID driver and chip specific
ac/surface: allow to select hybrid/block memcpy path for host copies
radv: take advantage of VK_HOST_IMAGE_COPY_MEMCPY_BIT
vulkan: replace VK_SHADER_CREATE_INDEPENDENT_SETS_BIT_MESA with the maint11 flag
vulkan: stop forcing independent sets for shader object
ci: uprev vkd3d
radv/tests: add tests for global pipeline keys compatibility
radv: allow DGC+multiview by default
radv: fix an assertion with RADV_DEBUG=fullsync on GFX11+
radv: do not fallback to compute for image->buffer copies with emulated formats
spirv: preserve the explicit stride for untyped pointers with matrices
radv: add support for VK_SHADER_CREATE_INDEPENDENT_SETS_BIT_KHR
radv: adjust minImageTransferGranularity for transfer queue
radv: advertise VK_KHR_maintenance11
radv: fix another case of VRS with mipmaps on GFX10.3
radv: remove a TODO about layeredShadingRateAttachments
radv: invalidate command buffer state after executing secondaries
radv/meta: adjust an assertion for HTILE expand on SDMA with compute fallback
radv: clear the follower gang semaphore when a cmdbuf is reset
radv: destroy the gang CS when a cmdbuf is reset
radv: fix copying acceleration structure with DAC
nir: fix shuffling local IDs for quad derivatives with larger workgroup sizes
radv: enable radv_wait_for_vm_map_updates for Forza Horizon 6
radv: advertise VK_EXT_device_fault by default
radv: remove an outdated comment in radv_GetDeviceFaultInfoEXT()
radv: move radv_GetDeviceFaultInfoEXT() to radv_device.c
util: pass a struct to driParseConfigFiles()
util: do not generate drirc options that shouldn’t be parsed
util: fix declaring drirc options as string
radv: rename few drirc options for consistency
radv: use the new generation script for drirc
radv/ci: cleanup list of expected failures
util: add very basic way to validate drirc files
radv: validate drirc option names at compile time
radv: close the local fd immediately after the winsys is created
radv: rename radv_zero_vram to vk_zero_vram
radv: use radv_device::ws directly for quering sync payloads
radv: pre-compute a mask of supported global queue priorities
radv: add a separate function to query allocated/usage for each heap
radv: remove declared but unused create_null_physical_device()
radv/amdgpu: simplify syncobj verifications during submissions
radv: determine supported syncobj types directly in the physical device
radv/amdgpu: fix releasing the mutex for virtio and RADV_PERFTEST=localbos
nir: add new intrinsics for SPV_KHR_abort
spirv: implement SPV_KHR_abort
nir: add nir_lower_abort
radv: close the local fd slightly later when enumerating physical devices
radv: remove useless checks when creating a physical_device
radv: rename master_fd to wsi_master_fd
util: share the DOCTYPE for all driconf files
util: add a separate file for Zink drirc
util: add a separate file for RadeonSI drirc
util: add a separate file for turnip drirc
util: add a separate file for ANV drirc
util: add a separate file for NVK drirc
util: add a separate file for r300 drirc
util: add a separate file for iris drirc
util: add a separate file for asahi drirc
util: add a separate file for asahi vulkan drirc
util: add a separate file for panvk drirc
util: add a separate file for panfrost drirc
util: add a separate file for crocus drirc
util: add a separate file for dozen drirc
util: add a separate file for virgl drirc
util: add a separate file for r600 drirc
util: add a separate file for msm drirc
util: add a separate file for d3d12 drirc
util: add a separate file for vmgfx drirc
util: add a separate file for v3d drirc
util: add a separate file for hasvk drirc
util: remove useless comments in 00-mesa-defaults.conf
util,turnip: move drirc entries with vk_dont_care_as_load to Turnip
util,asahi: move drirc entries with no_fp16 to asahi
util: remove declared but unused drirc options
ci: adjust time-trace.sh to not exceed the limit of 255 chars
radv/ci: fix list of expected failures
radv/amdgpu: rework tracking allocated memory for budget
radv/amdgpu: stop deduplicating winsys
ac/nir,radv: lower task payload to zeroes when the mesh shader has no task
radv: enable radv_force_64_byte_sampled_image for Crimson Desert
radv: cleanup conditional header includes
radv: implement VK_KHR_device_fault
radv: advertise VK_KHR_device_fault
radv: fix DGC with conditional rendering and task+mesh shaders
radv/amdgpu: allow RADV_PERFTEST=localbos with virtio
radv: cleanup occurrences of radeon_info::has_vm_always_valid
radv: return VK_ERROR_INITIALIZATION_FAILED if VM_ALWAYS_VALID isn’t supported
ci: uprev vkd3d
util: remove declared but unused DRIC_CONF_VK_REQUIRE_ASTC
util/drirc_gen: add a function to declare commmon VK options
radv: declare common VK drirc options using the helper
anv: declare common VK drirc options using the helper
turnip: declare common VK drirc options using the helper
radv/rt: fix a memory leak with hash tables
radv/rt: fix a memory leak with the RT prolog NIR
radv/rt: fix a memory leak with ahit/isec group
radv: fix a memory leak with perfcounters
radv/amdgpu: destroy the BO for the NULL PRT workaround earlier
vulkan: Update spec to 1.4.353
ci: uprev vkd3d
ci/vkd3d: add support for running with ASAN
radv/ci: run vkd3d jobs with ASAN by default
radv: add the mesh scratch ring BO to the preambles BO list
radv: handle errors correctly when creating gang waits
aco: emit nir_jump_halt
radv: implement VK_KHR_shader_abort
radv: advertise VK_KHR_shader_abort
radv/amdgpu: defer allocating the NULL PRT BO
radv/ci: skip all WSI tests on GFX1201
util/drirc_gen: change the driconf DTD to not require one app/engine entry
util/drirc_gen: fix generating 64-bit driconf options
util/drirc_gen: prevent generating empty structs
util/drirc_gen: add heap_memory_percent to common VK options
util/drirc_gen: allow to override the defaults VK WSI common options
dzn: use drirc_gen
pvr: use drirc_gen
panvk: use drirc_gen
hk: use drirc_gen
v3dv: use drirc_gen
venus: use drirc_gen
radv,anv: remove useless includes for drirc stuff
util/drirc: remove the driver option in drirc_validate
ci: add a new option called profile in ci_run_n_monitor.py
radv/ci: skip all WSI tests also on NAVI21/NAVI31
util: remove useless entries for Intel hasvk
hasvk: use drirc_gen
ac/video: drop an useless drm_minor check
ac/descriptors: fix setting CB_COLOR_ATTRIB3.RESOURCE_LEVEL
ac/gpu_info: only initialize has_desc_resource_level on GFX10+
nir/print: add a missing UNREACHABLE for unknown jump instructions
nir: add a new nir_jump_abort
nir,aco: use nir_jump_abort instead of nir_jump_halt for abort
radv/ci: add more flakes for RAPHAEL
radv/ci: update the list of expected failures for NAVI10
radv: prevent closing the render node fd twice for AMD_FORCE_VPIPE=1
radv: allow to query GPU info without creating a winsys
radv/amdgpu: add a function to query heap info
radv: query heap info without using the winsys
radv: duplicate the fd used for syncobj with KHR_display
radv: create one winsys for each logical device
radv: fix REPLAYED shader arena blocks not being marked as holes on free
radv/amdgpu: fix padding by one VM page
vulkan: fix lowering untyped accel struct with descriptor heap
spirv: mark UBO/SSBO array accesses as always in-bounds
radv: fix a synchronization issue with taskmesh and pending cache flushes
radv: fix a synchronization bug with DGC preprocess and taskmesh
radv: clear gang cache flushes when the command buffer is reset
vulkan: fix incorrect sType for VkDebugUtilsObjectTagInfoEXT
radv: workaround game bugs with Sniper Elite 5
spirv: allow mapping readonly buffers with struct members
radv: fix clearing the streamout state on GFX12
radv: remove redundant memory initialization on the CPU
vulkan: fix initializing address flags
vulkan: add vk_buffer_usage_flags()
radv: cleanup pCreateInfo uses for VkBuffer
radv: cleanup pCreateInfo uses for VkImage
radv: allow ptr to be NULL in radv_cmd_buffer_upload_alloc()
radv/amdgpu: fix computing allocated VRAM for imported BOs from fd
radv: fix a memleak with embedded samplers and descriptor heap
radv: store copying embedded samplers for heap to the shader layout
radv: enable VK_EXT_descriptor_heap by default
radv: disable VRS with MSAA 8x on GFX11-11.7 to prevent GPU hangs
radv: disable VRS for flat shading with MSAA 8X to prevent GPU hangs on GFX11
radv: disable VRS with MSAA 8x also on GFX10.3
radv: remove the deprecated warning for RADV_FORCE_FAMILY
docs,radv: auto-generate driconf documentation from drirc_gen
util: remove declared but unused driconf vulkan-related options
vulkan,anv,radv: do not crash when querying descriptor size for unsupported type
zink: fix a memleak in zink_init_format_props()
glsl: fix a memleak in link_assign_subroutine_types()
vulkan: search VkImageViewUsageCreateInfo in pNext
vulkan: implement VK_KHR_extended_flags
vulkan/wsi: implement VK_KHR_extended_flags
zink/ci: update lists for RADV
loader: fix a memleak
kopper: fix a memleak
zink: fix a memleak with fences
zink: fix a memleak with the emulated GS NIR shader
zink: fix a memleak with sampler state
radv: use the image view usage for MSRSTT transient iviews
radv: implement VK_KHR_extended_flags
radv: advertise VK_KHR_extended_flags
ci: apply patches to fix memleaks for GL/GLES CTS
zink/ci: add a new job for NAVI31 with ASAN enabled
pipe-loader: fix a global-buffer-overflow ASAN error when getting driconf
radv/meta: fix restoring descriptor heaps
Revert “spirv: allow mapping readonly buffers with struct members”
radv: always consider some outputs as invariant
radv: remove radv_invariant_geom driconf option
zink/ci: update trace checksums
radv: force late-Z with fragment shaders that use fbfetch
spirv: fix handling OpAbortKHR
radv: fix invalid assertions in DGC when queues aren’t enabled
Serdar Kocdemir (14):
gfxstream: add gitignore for generated code
gfxstream: Add VK_EXT_pipeline_protected_access
gfxstream: allow VK_KHR_maintenance extensions
gfxstream: some cleanup on device extension allow list
Set driver ID for gfxstream
gfxstream: allow VK_GOOGLE_display_timing
gfxstream: remove android conditioning for sampler extensions
gfxstream: use VK_DRIVER_ID_MESA_GFXSTREAM as driver id
gfxstream: update codegen for host side vulkan header update to v1.4.350
gfxstream: Fix codegen causing missing vulkan structures
gfxstream: disallow maintenance6 extension due to serialization bugs
gfxstream: correctly ignore timeline semaphore info
gfxstream: allow VK_EXT_border_color_swizzle
gfxstream: check in auto-generated guest code
Sergi Blanch Torne (15):
ci: disable Collabora’s farm due to maintenance
Revert “ci: disable Collabora’s farm due to maintenance”
ci: disable Collabora’s farm due to maintenance
Revert “ci: disable Collabora’s farm due to maintenance”
xfiles: update before uprev
ci: disable Collabora’s farm due to maintenance
Revert “ci: disable Collabora’s farm due to maintenance”
ci: review initial ANGLE flakes
ci,crnm: information from pipeline url
ci,crnm: search MR pipelines in forks
ci,crnm: handle exception when auth fail
xfiles: update before uprev Piglit
xfiles: update expectations based on 2026-7-8 nightly
ci: disable Collabora’s farm due to maintenance
Revert “ci: disable Collabora’s farm due to maintenance”
Sergi Blanch-Torne (1):
ci,crnm: bugfix project default
Sergio Sanchez Valencia (1):
d3d12/wgl: reclaim deferred BOs before ResizeBuffers
Shih, Jude (4):
amd/vpelib: Alpha blending enhancement
amd/vpelib: Refactor DPP function table layout
amd/vpelib: Fix Compiler Warnings
amd/vpelib: Realign DPP callback initialization with the updated interface layout
Sid Pranjale (12):
nvk: Implement VK_EXT_shader_atomic_float
vulkan: implement VK_EXT_debug_marker
v3dv: drop legacy CPU queue fallback paths
nak/nir: lower f16vec2 shared atomics
v3dv: replace single-field options struct with bool
v3dv: directly use v3d_has_feature instead of caps struct
broadcom/common: add multisync helpers
gallium/v3d: use common multisync code
v3dv: use common multisync code
v3dv: implement CPU-side fence merging for queue signaling
v3dv: simplify queue submission
v3dv: remove unused no-op job allocation setup
Silvio Vilerino (18):
d3d12: Create PIPE_BIND_SHARED resources with D3D12_RESOURCE_FLAG_ALLOW_SIMULTANEOUS_ACCESS
mediafoundation: Create readable dpb buffers with PIPE_BIND_RENDER_TARGET and PIPE_BIND_SHARED for DX11 sharing
Revert “d3d12: Video sliced encode: Use same ID3D12Fence/different per slice values as optimization”
d3d12: Flush stale video encode wait registrations when reusing ID3D12Fence objects
d3d12: Support video encode AUTO slice/tile only capable hardware
mediafoundation: check for AUTO slice/tile only capable hardware
d3d12: Use sequential video enc subregion signaling
d3d12: d3d12_create_fence_raw to lazily register fence event on waits
d3d12/video: fix comparison-with-wider-type warnings
d3d12: avoid signed integer overflow in copy staging box setup
util: u_trace.c: Fix error C4189: buffer_count: local variable is initialized but not referenced
d3d12: Fix NULL dereference check in d3d12_video_buffer_destroy
d3d12: Use res_device to import resource from different device via handle
pipe: Expose new fence_wait_multiple operation
d3d12: Implement fence_wait_multiple with SetEventOnMultipleFenceCompletion
mediafoundation: Use eventless fence_wait_multiple instead of WaitForMultipleObjects
d3d12: Remove event cleanup since d3d12_fence now uses lazy SEOC
d3d12: Check fence values before SEOC in d3d12_fence_wait_multiple
Simon Perretta (29):
pco: reserve additional outputs for trilinear sampled coeffs
pco: amend tg4 lowering
pco: track how many tg4/raw sample comps are needed
pvr: consider barriers when calculating compute instances
pvr, pco: add support for spilling shared memory to global memory
pvr, pco: store device runtime info in compiler context
pco: conditionally spill shared memory to global memory
pco, pvr: finish and enable VK_KHR_workgroup_memory_explicit_layout
pco: drop global path for null descriptor checking
pco: add mappings for setl, savl ops
pvr, pco: add “real” basic subgroup support
pco: handle mov offset special regs
pco: add support for read_invocation via shared memory
pco: add subgroup ballot support via shared memory
pvr: advertise VK_EXT_shader_subgroup_ballot and ballot feature
pco: add br.skip_next op
pco: commonize execution mask counter ref helper function
pco: add support for subgroup vote_{all,any} ops
pvr: advertise VK_EXT_shader_subgroup_vote and vote feature
pco: add support for reduce/scan ops with cluster awareness
pvr: advertise subgroup arithmetic and clustered features
pco: add support for subgroup shuffle ops
pvr: advertise subgroup shuffle and shuffle relative features
pvr, pco: add support for VK_KHR_shader_subgroup_rotate
pvr: advertise VK_KHR_shader_subgroup_uniform_control_flow
pvr, pco: advertise support for VK_EXT_subgroup_size_control
pco: allow non-pure integer formats for image xchg atomics
pco: lower sysvals early for fragment shaders
pco: allow fence ops to be legalized if they come last in a block
Skyth (1):
spirv2dxil: Replace UAV_FENCE_THREAD_GROUP usage with UAV_FENCE_GLOBAL.
Sonny Jiang (2):
radeonsi: always set is_format_supported in screen create
radeonsi/vcn: Add vcn_5_0_2 support
Stijn Tintel (1):
rocket: fix mmap leak in buffer map/unmap
Stéphane Cerveau (1):
vulkan/video: Reject interlaced picture layout for H.264 baseline profile
Suresh Guttula (1):
ac: Add vcn_5_3_0 support
Sushma Venkatesh Reddy (2):
intel/perf: Add WCL OA support
intel/dev: Clamp PTL+ CS workgroup threads to 32
Tacodiva (1):
vulkan/runtime: Fix bad assumption in GetPipelineBinaryDataKHR
Tanner Van De Walle (5):
draw: add lower-bound assert on shader_stage
gallium/u_blitter: add lower-bound assert on target
util/format: add lower-bound assert on format
dzn: silence PREfast C33010 warnings
nir/nir_builder: inline dst_bit_size calculation in assert
Tapani Pälli (13):
intel/compiler: implement macl part of Wa_18035690555
drirc: use anv_disable_drm_ccs_modifiers for any GTK version
drirc/anv: add flag to disable VK_EXT_subgroup_size_control
drirc: set anv_disable_subgroup_size_control for bg3
anv: do not use resource barrier with split barriers
intel/dev: update mesa_defs.json from workaround database
iris: use INTEL_NEEDS_WA_14025112257 define for workaround
anv: use INTEL_NEEDS_WA_14025112257 define for workaround
anv: allocate tile sized temporary copy instead of whole size
anv: skip writing xfb buffer if we get null information
iris: align down the max_shader_buffer_size
anv: fix a null pointer access with isl_mod_info
anv: optimization for Wa_14025112257 case
Thomas H.P. Andersen (5):
nvk: set queryResultStatusSupport
nvk: use the new generation script for drirc
nouveau/cubin: use libelf 64 bit instead of gelf
nvk: hide NVX_binary_import behind NVK_EXPERIMENTAL=dlss env var
nvk: add env var to allow backwards compat in dlss
Thong Thai (25):
util: move u_stub to src/util, add u_stub_gfx_compute.h
util: allow for overriding u_stub tail
meson: update default build option for libva subproject
meson: check if video encoding support is to be built
frontends/va: decode only stubs
radeonsi: move si_get video functions to si_video
amd: make ac_ib_parser an amd tool build option
gallium/auxiliary/vl: Fix typo in cs_create_shader pseudo-code comment
pipe: Add PIPE_VIDEO_VPP_BLEND_MODE_PREMULTIPLIED_ALPHA
vl/video_buffer: Set alpha swizzle if format has alpha
gallium/video: Add enabled flag to vpp interface
gallium/vl: Implement compositor shader-based alpha blending
frontends/va: Enable shader-based alpha blending
amd: Build nir files only when with_gfx_compute
radeonsi: Remove ACO dependency for non-GFX/compute builds
nir: Only build NIR headers when with_gfx_compute is false
gallium/auxiliary: Simplify auxiliary for non-gfx/compute builds
meson: Make with_gfx_compute depend on video encode support
meson: Don’t require libelf for radeonsi when with_gfx_compute is false
radeonsi: Allow call to stub’d si_init_gfx_context to continue
radeonsi: Store SQTT cb_id
radeonsi: Handle SQTT timestamps
radeonsi: Store SQTT device_id
radeonsi: Implement SQTT CB_START and CB_END
radeonsi: Setup SQTT sampling clocks
Timothy Arceri (17):
glcpp: update out of date comment
glcpp: fix paste within macro function expansion
amd/radeonsi: dont clamp packed user varyings
mesa: fix typo in validation string
ac/nir/lower_tex_coord: update cursor when moving wqm coordinates
ac/nir/lower_tex_coord: basic lower tex coord test
mesa: flush bitmap cache when scissor box changes
nir: use the correct induction var when guessing loop iterations
glsl: allow uniform block layout qualifiers when SSBO enabled
util/u_range_remap: allow insert to truncate range
glsl: treat temp globals wrappers as roots when resolving function calls
zink: fix swap interval changes being dropped
nir/opt_dead_write_vars: handle memcpy_deref as reads
util: add Blockland workaround for crash
util/mesa: add workaround to zero invalidated buffers
util: add workaround for Riddick using round() in glsl 1.20
llvmpipe: emit FS input vertex attributes in driver location order
Timur Kristóf (10):
nir/divergence: Consider ACCESS_SMEM_AMD divergence across subgroups
nir/divergence: Consider uniformity of read_invocation accross subgroups
nir/divergence: Consider ttmp_register_amd and load_scalar_arg_amd as workgroup divergent
ac/nir: When loading an arg, assert that it’s used
ac/nir: Fix SMEM workaround with emulated RT
radv: Wait for idle after every submission on GFX6-7
radeonsi: Wait for shaders and flush L2 after every submission on GFX6-7
Revert “radv: Mitigate GPU hang on Hawaii in Dota 2 and RotTR”
ac/nir/ngg: Remember if a mesh shader has non-API waves.
ac/nir/ngg: Use workgroup divergence analysis for mesh output counts.
Tomeu Vizoso (5):
teflon/tests: avoid loading build-tree tensorflow-lite stub at runtime
teflon/tests: make tflite stubs fail loudly with diagnostics
teflon: remove synthetic model generation and flatbuffers dependency
ci: Remove flatbuffers from builds
teflon/tests: Remove leftover files from synthetic tests
Toshinari Morikawa (2):
virgl: fix memory leak on shader translation
egl: avoid calling loader_get_driver_for_fd with fd = -1
Trigger Huang (12):
radv: supports protected memory allocation
radv: allow creation of protected queues
radv: support secure submission
radv: add protected type bits for memory requirements
radv: enable protected memory
radv: emulate MSRTSS via implicit MSAA resolve
radv/meta: thread separate src/dst sample counts through gfx copy
radv/meta: derive gfx copy dst sample count from the destination
radv/meta: add MSRTSS attachment replicate helper
radv: replicate MSRTSS attachments on LOAD_OP_LOAD
radv: handle VkSubpassResolvePerformanceQueryEXT
radv: enable VK_EXT_multisampled_render_to_single_sampled
UMU618 (1):
venus: fix typo in vn_queue_submit_2_to_1
UMUTech (1):
wsi: correct the erroneous assertion
Utku Iseri (1):
v3dv: close display_fd on incompatible_driver path
Val Packett (3):
util: rust: align API with real eventfd capabilities
util: rust: Support detecting socket file descriptors
util: rust: Add a way to create a Tube from an existing OwnedFd
Valentine Burley (102):
mr-label-maker: Label Collabora farm with driver tags
ci/zink/intel: Disable flaky TGL canvas_moire-v2 trace
zink/ci: Document recent flakes
anv/ci: Revert ADL VKCTS job to stable 6.17 kernel
zink/ci: Move Turnip flakes to correct list
tu/drm/virtio: Fix tu_wait_fence timeout handling
freedreno/drm/virtio: Fix wait_fence ret ordering
zink/ci: Remove Cezanne job
radv/ci: Add more ASAN VKCTS jobs on Cezanne
vulkan/android: Add deferred image helper
panvk: Use vk_android deferred image helper
vulkan/android: Add vk_android_import_anb_memory helper
vulkan: Query memory requirements in vk_android_import_anb_memory
tu: Implement deferred image creation for ANB and AHB
ci/crosvm: Sanitize CROSVM_RET in crosvm-runner.sh
tu: Fix D16 depth clear rounding mismatch in sysmem mode
panfrost/ci: Update kernel to pick up ZSTD support for ZRAM
venus/ci: Skip more robustness tests on ANV
tu: Move Android extensions into main list
tu: Add shared image support on Android
panfrost/ci: Move t860 jobs to nightly
panfrost/ci: Document recent g610 flake
ci/android: Remove SurfaceFlinger wait in get_surfaceflinger_pid
ci/android: Fix intermittent adb root failures
ci/android: Update Cuttlefish build
ci/deqp: Add Android WSI support
lavapipe/ci: Enable WSI testing on Android
turnip/ci: Enable WSI testing on Android
venus/ci: Enable WSI testing on Android
ci/android: Remove CtsDeqpTestCases from Android CTS
ci/deqp: Backport host_image_copy fix
ci/lava: Reduce LAVA job timeout to 20 minutes for Marge
mr-label-maker: Add rule for new trace replay config files
ci: Add missing rule for new trace replay config files
tu/autotune: Clear active_batches before history objects are freed
ci/deqp: Backport validation error fix
ci/deqp: Backport landed patch
ci/deqp: Rewrite headless Android WSI patch
venus/ci: Skip more even more robustness and synchronization2 tests on ANV
ci: Disable debian-riscv64
panvk/ci: Mark dEQP-VK.subgroups.* as flaky on G925
ci: Bump ci-deb-repo revision to update aapt
ci/android: Update Android CTS to android-cts-16.0_r5
ci/android: Add arm64 support for Android CTS
turnip/ci: Add nightly Android CTS job
tu: Disable -Wmisleading-indentation when compiling with GCC
panvk: Fix ignored qualifier warnings
meson: Add Soong compatibility compiler flags to Vulkan drivers
tu/kgsl: Fix memory type support detection for unsupported flags
tu: Merge tu_image_init and tu_image_update_layout
venus/ci: Widen the ANV skips
tu: Advertise VK_KHR_internally_synchronized_queues
turnip/ci: Update ANGLE trace checksum
panfrost/ci: Switch traces over to gpu-trace-perf
tu: Fix vk_queue leak on submitqueue creation failure
vulkan/queue: Add common queue emulation support
tu: Emulate second graphics queue for skiavk on Android
tu/ci: Add coverage for emulated second graphics queue
ci/android: Update Cuttlefish build
anv/ci: Disable anv-adl-vk job
venus/ci: Move pre-merge ANV coverage from Comet Lake to Alder Lake
venus/ci: Retire Intel Comet Lake runner
venus/ci: Revert ADL jobs to stable 6.17 kernel
intel/gen: Explicitly declare gen_opcodes_private.h dependency
perfetto: Centralize perfetto header include in u_perfetto.h
bin: Expose drm-shim in meson devenv
doc/ci: Add drm-shim CI reproduction guide
zink/ci: Increase zink-lavapipe parallelism
virgl/ci: Retire disabled virgl-iris jobs
virgl/ci: Retire disabled android-virgl-llvmpipe
pipe-loader: Enable null winsys on Android
zink/ci: Remove zink-anv-cml-asan job
intel/ci: Remove nightly CML jobs, retire runner
intel/ci: Increase iris-apl-egl parallelism
zink/ci: Drop old VVL filters
panfrost/ci: Fix typo in .panfrost-vk-manual-panthor-rules template
tu: Fix uninitialized gmem_offset when a GMEM layout is impossible
ci: Update kernel to Linux 7.1.2
panfrost/ci: Use Linux 7.1 kernel for more jobs
tu: Fix capture/replay with sampler custom border color
ci/lava: Uprev lava-job-submitter
drm-shim/freedreno: Add support for Adreno 610
drm-shim/freedreno: Shim perf counter config ioctl
bin/drm-shim: Add more freedreno GPUs
tu: Disable VK_EXT_extended_dynamic_state2 patch control points on A702
zink: Gate tess/geom barrier stages on feature support
ci: Bump ci-deb-repo revision to update vulkan-loader
ci/deqp: Backport landed Android WSI patch
ci/deqp: Update VK CTS to 1.4.6.1
ci/deqp: Backport -frounding-math default for GCC builds
tu: Implement VK_EXT_primitive_restart_index
ir3: Fix ballot_components for subgroups smaller than 32
ir3: Derive max_variable_workgroup_size from device limits
ir3: Compute subgroup_size from threadsize_base
tu: Use computed subgroup size
panvk/ci: Update expectations for g610-vk-asan following VK CTS uprev
panvk: Disable VK_KHR_internally_synchronized_queues on Vulkan 1.0
tu: Report correct maxFragmentInputComponents limits
zink: Use ShaderLayer capability for gl_Layer when available
freedreno: Increase reg_size_vec4 for A702
freedreno: Fix VPC_RAST_STREAM_CNTL register layouts
ci/piglit: Switch all trace jobs to surfaceless+gbm
Vincent Cloutier (2):
etnaviv: use buffer resource accessor for indirect draws
etnaviv: support native bitfield extract/reverse/count ops
Vinson Lee (14):
st/mesa: fix implicit conversion warning in st_atom_framebuffer
vulkan/screenshot-layer: initialize info to NULL
gfxstream: codegen: drop const from let-param scalar cast
mesa/main: cast GLhandleARB to unsigned int in api trace
ethosu/mlw_codec: silence warnings in the vendored Regor encoder
ethosu/mlw_codec: silence -Wunused-const-variable in vendored encoder
ethosu: use FALLTHROUGH macro in ethosu_emit_operation_accesses
radeonsi: remove duplicate ‘.bpp’ initializer in si_sdma_copy_image
gfxstream: link goldfish_address_space against perfetto
util/tests: replace sprintf with snprintf in cache tests
util/tests: fix unused variable warnings in cache List test
vulkan/screenshot-layer: replace itoa/sprintf with snprintf
vulkan/screenshot-layer: fix globalLock mutex leak
util/tests: silence unused iterator warning in sparse_bitset_test
Virgile Bello (3):
microsoft/compiler: sink load_invocation_id in TCS split even for single-use
microsoft/compiler, d3d12: flip tess winding at caller, not in nir_to_dxil
microsoft/compiler, d3d12: preserve TCS outputs and pad TES inputs for cross-stage signature matching
Vishnu Vardan (27):
mesa/st: remove redundant has_stencil_export from st_context
mesa/st: remove redundant astc_void_extents_need_denorm_flush from st_context
mesa/st: remove redundant has_shareable_shaders from st_context
mesa/st: remove redundant needs_texcoord_semantic from st_context
mesa/st: remove emulate_gl_clamp from st_context
mesa/st: remove has_time_elapsed from st_context
mesa/st: remove has_multi_draw_indirect from st_context
mesa/st: remove has_indirect_partial_stride from st_context
mesa/st: remove has_occlusion_query from st_context
mesa/st: remove has_single_pipe_stat from st_context
mesa/st: remove has_pipeline_stat from st_context
mesa/st: remove has_indep_blend_enable from st_context
mesa/st: remove has_indep_blend_func from st_context
mesa/st: remove can_dither from st_context
mesa/st: remove lower_flatshade from st_context
mesa/st: remove lower_alpha_test from st_context
mesa/st: remove lower_two_sided_color from st_context
mesa/st: remove lower_ucp from st_context
mesa/st: remove prefer_real_buffer_in_constbuf0 from st_context
mesa/st: remove has_conditional_render from st_context
mesa/st: remove lower_rect_tex from st_context
mesa/st: remove allow_st_finalize_nir_twice from st_context
mesa/st: remove can_bind_const_buffer_as_vertex from st_context
mesa/st: remove validate_all_dirty_states from st_context
mesa/st: remove can_null_texture from st_context
mesa/st: remove redundant has_hw_atomics from st_context
anv/rt: reorder encode_internal_node to only process valid children
Vlad Zahorodnii (1):
wsi/wayland: Add support for wl_fixes.ack_global_remove
Wig Cheng (2):
rocket: pad weight packing input channels to FEATURE_ATOMIC_SIZE
rocket: compute element-wise ADD requant instead of LUT
Wujian Sun (2):
mesa: Fix clipping order in _mesa_clip_blit()
mesa: Allow GL_SRGB_ALPHA_EXT as color-renderable when EXT_sRGB is supported
Xinju Li (1):
nir: resolve functions: only resolve functions that are reachable from main
Yannis Juglaret (1):
nouveau: fix data race in nouveau_fence_ref
Yiwei Zhang (146):
venus: adopt vk_android_init_deferred_image
venus: adopt vk_android_get_ahb_layout
venus: refactor vn_android_get_wsi_memory to return VkDeviceMemory
venus: adopt common vk_image::anb_memory
venus: adopt common ANB helpers
panvk: adopt common ANB helpers
lvp/android: use common ANB implementations
util/android_stub: drop legacy atrace
panvk: drop panvk_android_create_deferred_image
util/os_misc: use ndk api __system_property_get
egl/android: use ndk api __system_property_get
android_stub: drop cutils/properties dependency
CODEOWNERS: update owners for Android components
util/os_misc: use stable NDK __android_log_write helper
intel: use stable NDK __android_log_print helper
broadcom: remove unused Android log utils
android_stub: purge unused log utils
ci: uprev virglrenderer
android_stub: fix update-android-headers.sh for libbacktrace
android_stub: fix libhardware source include path
android_stub: avoid vending in unused headers
android_stub: sync Android 16 headers
pan/nir/tex: use unsigned type for texture op lod_or_fetch
venus: fix a renderer side queue timeline bound race
panvk: fix to report device memory with heapIndex
tu: fix to report device memory with heapIndex
venus: update create_from_device_memory to take a cmd payload
venus: let resource_create_blob wait for mem alloc
venus: fix unbound malloc leak in vn_ring_get_submits
anv: fix lock scope in anv_ensure_fp64_shader
anv: amend missing shader dump finish upon device destruction
venus: amend roundtrip between fence submit and wait idle
panvk: fix plane indexing for afbc image subres layout
pan: handle downscaling of plane view for multiplanar yuv textures
panvk: reject interleaved_64k for multiplanar yuv
panvk: avoid separate reconstruction filter for YUV texturing
panvk: use vk_component_mapping_to_pipe_swizzle
panvk: apply YUV swizzle to the view swizzle for YUV texturing
panvk: override default chroma siting for YUV texturing
panvk: enforce strict import for yuv images
panvk: add get_pan_image_props helper
panvk: use binding layout textures_per_desc to write image view descs
panvk: add pan_texture_get_payload_alignment to help with tex emit
panvk: add and use panvk_image_get_tex_count helper
panvk: properly set up image and view planes for YUV texturing
panvk: lower YUV texturing to do SW CSC
pan: add 8bit multi-planar 420 and 422 format for Vulkan
panvk: enable 8bit multiplanar YUV formats on v9+ to v13
panvk: add P010 native YUV support
vulkan/android: force linear for mutable format
venus/wsi: skip VkPresentRegionsKHR when pRegions is NULL
venus: avoid touching sfb dst slot upon resume
venus: always check device lost on sfb warn order
venus: rename vn_semaphore_feedback_cmd to vn_sync_feedback_cmd
venus: move sfb helpers into vn_feedback
venus: wrap sfb cmd preparation with vn_sync_feedback_command
venus: extract sync feedback host write and query
venus: move sfb suspend resume handling over to vn_feedback
venus: add vn_sync_feedback_enabled
venus: always check device lost on ffb warn order
venus: migrate ffb to use vn_sync_feedback
venus: recycle fence sfb in post submission
venus: drop vk_xwayland_wait_ready override
v3dv: drop vk_xwayland_wait_ready
hasvk: drop vk_xwayland_wait_ready
vulkan/wsi/util: purge vk_xwayland_wait_ready
venus/virtgpu: amend a missing sim mutex init
venus/virtgpu: drop obsolete SIMULATE_BO_SIZE_FIX
venus/virtgpu: implicit fencing is gone
venus/virtgpu: simplify to drop virtgpu_sync
venus: refactor vn_renderer_submit to only take a single batch
venus/virtgpu: merge SIMULATE_SUBMIT into SIMULATE_SYNCOBJ
venus/virtgpu: drop signaled_fd
venus/virtgpu: drop cpu sync timeout
venus/virtgpu: drop wait available
venus/virtgpu: use STACK_ARRAY for syncobj handles
venus/virtgpu: split out sim_syncobj
venus/virtgpu: refactor sim_submit
venus/virtgpu: use uAPI to signal syncobjs
venus/virtgpu: adopt u_sync_provider
venus/virtgpu: use virtgpu syncobj on supported kernels
venus/virtgpu: avoid pretending timeline sync support
venus/virtgpu: sim_syncobj to implement u_sync_provider
venus/virtgpu: simplify sim_syncobj
venus/virtgpu: refactor virtgpu_submit
venus/virtgpu: flatten all the virtgpu_ioctl_syncobj wrappers
venus: properly clean up driver internal sim syncobj allocs
venus: fix imported sync fence payload reset upon export
venus: avoid renderer semaphore wait upon temp payload export
venus: refactor external fence and semaphore advertisment
venus: drop obsolete zink performance workaround
venus: track can_feedback in struct vn_queue_submission
venus: purge sync feedback for sparse binding
venus: split fence/semaphore/event commands to vn_sync.(c|h)
venus: refactor imported semaphore check and wait
venus/wsi: refactor args of vn_wsi_fence_wait and vn_wsi_flush
venus: queue submit to take a single batch
venus: sparse binding to use vn_queue_submit
venus: simplify vn_queue_submission to handle single batch
venus: simplify feedback cmd setup
venus: simplify vn_queue_submission_alloc_storage
venus: extract pNext chain fixup out from feedback cmds init
venus: prepare to handle sparse binding batch interception
venus: properly drop imported semaphores from submission
venus: relax SYNC_FD semaphore import requirement for WSI
venus/renderer: improve renderer backend init logs
venus: recycle idle sfb cmds only after async wait
venus: properly check sync2 enablement
venus: host image copy to scrub present_src layout if needed
venus: clean up sync2 treatment leftovers
venus: simplify sync feedback tracking
venus: update pnext fix tracking
venus: explicitly track if need to fix batch
venus: flatten sync feedback cmd counting
venus: deprecate fence feedback
venus: split ring submission to vn_queue_submission_do_submit
venus: skip empty batch submission
venus: drop vn_sync_payload_external from submission tracking
venus: extract vn_timeout_to_poll_timeout
venus/virtgpu: only signal non-zero initial value
venus/virtgpu/vtest: drop initial value from sync reset
venus: prepare for VN_SYNC_TYPE_SYNC
venus: migrate fence over to VN_SYNC_TYPE_SYNC
venus: relax SYNC_FD fence export requirement
venus: deprecate queue idle wait workaround
venus: rename to be explicit about sync fd semaphore
venus: vn_semaphore_(is|wait)_sync_fd to support VN_SYNC_TYPE_SYNC
venus: rename existing queue submission wait semaphore tracking
venus: count and prepare storage to scrub SYNC_FD signal semaphore
venus: extract syncs and scrub SYNC_FD signal semaphores
venus: ensure renderer sync fence is submitted between queue batches
venus: migrate SYNC_FD semaphore over to VN_SYNC_TYPE_SYNC
venus: relax SYNC_FD semaphore export requirement
venus/virtgpu: hide has_timeline_sync behind a new perf option
venus/vtest: advertise timeline syncobj support
venus: drop vn_renderer_sync_flags
venus: add VN_SYNC_TYPE_TIMELINE_SYNC
venus: track queue internal array index within vn_device::queues
venus: count and init syncs from timeline semaphores
venus: implement host signal for TIMELINE_SYNC
venus: implement counter query for TIMELINE_SYNC
venus: add vn_wait_semaphores_legacy for legacy wait
venus: implement semaphore wait for TIMELINE_SYNC
venus: migrate timeline semaphore to VN_SYNC_TYPE_TIMELINE_SYNC
venus: document timeline semaphore implementation
venus: ensure cached vn_ring_submit batches are bounded
Yogesh Mohan Marimuthu (5):
ac,radeonsi,radv: add has_desc_resource_level var instead of gfx_level check
radv: Program RESOURCE_LEVEL bit in descriptor for dgc
amd: add initial code for gfx1156
ac: set has_smem_with_null_prt_bug to false for gfx1156
ac: set has_desc_resource_level to true for gfx1156
You, Min-Hsuan (1):
amd/vpelib: fix FROD alignment handling after interface change
Zan Dobersek (5):
fd: add a8xx perfcntr countables
tu: only support userspace-managed perfcounters on a7xx and earlier
tu/a8xx: remove enforced TU_DEBUG_FLUSHALL
tu/kgsl: initialize dump bo state in kgsl_bo_init sooner
fd: lrz_block in fdl6_lrz_layout_init() should not be static
Zeyang Lyu (1):
radv: Add base array layer to htile offset
Zhao, Jiali (2):
amd/vpelib: revert predication fix
amd/vpelib: fix HDR external monitor video black content
ZhengMing (1):
vulkan/wsi/win32: Prefer the more popular surface format on Windows
Zoltán Böszörményi (1):
radv: Advertise msrtss in features.txt
adrian baker (4):
jay: add bfloat16 support
jay: fix jay bf16 comment formatting
nir: remove incorrect algebraic properties from intel mixed bf ops
jay: add geometry shader support.
gyeyoung (2):
panvk: fix flags2-only bit leak in legacy format features
panvk: report DRM format modifiers through List2EXT
gyeyoung baek (1):
rocket: drop wrong assert(input_op_1) in ADD fuse path
hmtheboy154 (15):
pvr: add support for driconf for the Vulkan driver
driconf: Add an option to override Vulkan’s deviceName
anv: driconf: Add an option to override Vulkan’s deviceName
hasvk: driconf: Add an option to override Vulkan’s deviceName
nvk: driconf: Add an option to override Vulkan’s deviceName
radv: driconf: Add an option to override Vulkan’s deviceName
venus: driconf: Add an option to override Vulkan’s deviceName
v3dv: driconf: Add an option to override Vulkan’s deviceName
lvp: add support for driconf
lvp: driconf: Add an option to override Vulkan’s deviceName
tu: driconf: Add an option to override Vulkan’s deviceName
panvk: driconf: Add an option to override Vulkan’s deviceName
pvr: driconf: Add an option to override Vulkan’s deviceName
dzn: driconf: Add an option to override Vulkan’s deviceName
hk: driconf: Add an option to override Vulkan’s deviceName
hwandy (1):
Revert “intel/decoder: make libvulkan_intel to depend on stub decoder when buildtyle=release.”
inspector-ambitious (1):
loader: fix loader_open_render_node_platform_devices result allocation
jglrxavpok (1):
RADV: Add object names inside address binding report and vm_fault
jiajia Qian (6):
rusticl: extract tokenize() and fix UTF-8 handling in compile options
rusticl: add LinkOptions struct with validation
rusticl: validate build/compile options before passing to backend
rusticl: validate input_programs binary type in clLinkProgram
ci/panfrost: add piglit OpenCL testing for G610
rusticl/device: use OpenCL spec minimum for mem_base_addr_align
jinmiliu (2):
mesa/st: Set protected content context flag based on pipe context attributes
radeonsi: enable protected context support for Android
johniyoods (1):
egl/dri2: require valid render fd before advertising EGL_WL_bind_wayland_display
jyotiranjan (1):
radv/sqtt: forward zero-submit-count vkQueueSubmit2 for SQTT capture
ljohnson (1):
venus/wsi: deep copy pRectangles when cloning presentation info
llyyr (2):
radeonsi: don’t init screen state functions twice
vulkan/wsi/wayland: use mtx helpers in wait_for_present2
nyanmisaka (1):
intel/dev: update PTL device names
sergiuferentz (1):
gfxstream: Prevent LINUX_GUEST_BUILD from being added to android platforms
squidbus (89):
kk: Use device limits for buffers and compute shared memory.
kk: Enable VK_AMD_shader_image_load_store_lod
kk: Update dynamic depth stencil state regardless of set attachments.
kk: Add type inference for additional built-in intrinsics.
kk: Fix VK_CULL_MODE_FRONT_AND_BACK with points and lines.
asahi,nir: Move asahi dynamic clipz pass to common.
kk: Add support for VK_EXT_depth_clip_control.
kk: Fix issues with maximal reconvergence
kk: Enable VK_EXT_extended_dynamic_state3
kk: Enable VK_EXT_buffer_device_address
kk: Enable VK_(EXT/KHR)_global_priority and VK_EXT_global_priority_query
kk: Workaround for GPU capture under Rosetta 2.
nir: Only attempt subgroups lower_boolean_reduce for single component.
kk: Expand workaround 3 to cover general use of ballot/vote ops
kk: Fix emitting negative infinity
kk: Support subgroup rotate ops
kk: Enable remaining subgroup operations
kk: Fix geometry unroll for list primitives.
kk: Support VK_(KHR/EXT)_index_type_uint8
kk: Enable VK_EXT_multi_draw
kk: Split per-draw data to separate binding
kk: Fix handling of sample mask and sample rate shading
kk: Support VK_EXT_post_depth_coverage
kk: Support robustBufferAccess2
kk: Support nullDescriptor
kk: Enable VK_(EXT/KHR)_robustness2 and VK_EXT_pipeline_robustness
kk: Enable VK_(EXT/KHR)_line_rasterization
kk: Support shaderCullDistance
kk: Query device for supported sample counts
kk: Complete VK_EXT_memory_budget
kk: Create image layout from vk_image
kk: Implement index buffer robustness for BindIndexBuffer2
kk: Handle accurate OpSMod and default point size requirements
kk: Disable A8_UNORM format
kk: Support device without queue
kk: Fix image copies for depth/stencil<->color and differing subresources
kk: Support new query pool and dynamic rendering flags
kk: Enable maintenance extensions through VK_KHR_maintenance10
kk: Separate linear and GPU optimized image layout properties
kk: Support VK_EXT_host_image_copy
kk: Support VK_KHR_shader_fma
kk: Support attachment feedback loop extensions
kk: Support VK_KHR_unified_image_layouts
kk: Fix some missed NIR debug asserts
kk: Allocate temporary command memory from pool
kk: Fix pre-compiled compute grid size
kk: Enable code formatting enforcement
poly: Refactor poly_unroll_restart for general purpose unrolling
poly: Fix range used for index unroll bounds checks
kk: Fix compute system value and algebric lowering in pre-compiles
kk: De-duplicate geometry unroll logic
kk: Support VK_KHR_shader_untyped_pointers
kk: Refactor multi-draws and predicates into kk_draw_data
kk: Enable VK_EXT_nested_command_buffer
kk: Support VK_EXT_conditional_rendering
kk: Fix precomp data buffer alignment
kk: Sanitize image copy through buffer extents
kk: Do not use image-to-image copies for 1D compressed textures
kk: Handle index robustness for fully bound buffers manually
kk: Accurately declare supported samples in image format properties
kk: Support VK_EXT_vertex_attribute_robustness
kk: Support VK_EXT_blend_operation_advanced
kk: Support VK_EXT_custom_resolve
kk: Support VK_EXT_primitive_restart_index
kk: Support VK_EXT_primitive_topology_list_restart
kk: Fence read-write images after write
kk: Support VK_IMAGE_CREATE_BLOCK_TEXEL_VIEW_COMPATIBLE_BIT
kk: Support VK_EXT_external_memory_host
kk: Advertise additional tessellation dynamic state
kk: Perform sink-and-move of instructions
kk: Do not force render image view to all subresources
kk: Fix divide by 0 in non-indexed draw unroll
kk: Support VK_EXT_sample_locations
kk: Remove unused deprecated Metal APIs
kk: Ensure some vertex lowerings happen on hardware stage
kk: Enable shaderTessellationAndGeometryPointSize
kk: Migrate to Metal 4 pipelines
kk: Work around crash with multiple concurrent MTL4Compiler
wsi/metal: Support HDR10 color spaces
kk,wsi/metal: Support VK_EXT_hdr_metadata
kk,wsi/metal: Support VK_(KHR/EXT)_swapchain_maintenance1
kk: Respect precomp-compiler options when setting up kk_clc
kk: Implement draw-related commands using device addresses
kk: Work around Metal index robustness gaps
kk: Pre-declare texture SSA variables
kk: Use safe math for nir_fp_no_reassoc
kk: Enable shaderRoundingModeRTEFloat16/32
kk: Only support 1 sample for storage images
kk: Remove deprecated MTL4CommandQueueErrorDeviceRemoved
utzcoz (4):
gfxstream: Validate guest mapped-memory ranges in flush/invalidate
ci/amd: enable ACO validation on radeonsi jobs
radeonsi: convert gather_instruction to nir_function_instructions_pass
virtio: magma-gpu-rs: accept a null device in virtgpu_kumquat_finish
xueyuli2 (1):
amd/virtio: fix bo use-after-free race condition in amdvgpu_bo_free
yserrr (4):
llvmpipe: fix UB and incorrect value in compute caps shift
v3d: fix stencil blit layer selection
v3d: lower more 64-bit integer operations
v3d: remove duplicate util_blitter_save_so_targets() call