Mesa 26.2.0 Release Notes / 2026-08-05

Mesa 26.2.0 is a new development release. People who are concerned with stability and reliability should stick with a previous release or wait for Mesa 26.2.1.

Mesa 26.2.0 implements the OpenGL 4.6 API, but the version reported by glGetString(GL_VERSION) or glGetIntegerv(GL_MAJOR_VERSION) / glGetIntegerv(GL_MINOR_VERSION) depends on the particular driver being used. Some drivers don’t support all the features required in OpenGL 4.6. OpenGL 4.6 is only available if requested at context creation. Compatibility contexts may report a lower version depending on each driver.

Mesa 26.2.0 implements the OpenCL 3.1 API, but the version reported by the CL_DEVICE_VERSION, CL_DEVICE_NUMERIC_VERSION and CL_DEVICE_OPENCL_C_ALL_VERSIONS clGetDeviceInfo queries depends on the particular driver being used.

Mesa 26.2.0 implements the Vulkan 1.4 API, but the version reported by the apiVersion property of the VkPhysicalDeviceProperties struct depends on the particular driver being used.

SHA checksums

SHA256: efd4bb08cdb7c365a812cd4e6c9202ab55b2f22cdcd13c7d6c4f9647b799a4ef  mesa-26.2.0.tar.xz
SHA512: c7c810c6958fa18eb75eea9968d84d0edd29b579d351a7f22d8c2b2b13b6e04e5b0df31dae3aa943672c792d124325bab03ccbd475a16d27267418bddbec4936  mesa-26.2.0.tar.xz

New features

  • cl_khr_subgroup_rotate on radeonsi

  • cl_khr_subgroup_rotate on iris

  • VK_EXT_shader_uniform_buffer_unsized_array on panvk

  • VK_KHR_shader_constant_data on RADV

  • VK_EXT_dynamic_rendering_unused_attachments on panvk

  • protectedMemory support on RADV/GFX10+ and VEGA10

  • VK_KHR_performance_query on RADV/GFX11

  • VK_EXT_conservative_rasterization on panvk

  • shaderImageGatherExtended on pvr

  • static C++ stdlib required on rusticl to workaround applications using their own C++ stdlib

  • VK_EXT_pipeline_protected_access on RADV

  • VK_EXT_extended_dynamic_state3 on panvk

  • GL_ARB_texture_query_lod on panfrost/v9+

  • VK_KHR_maintenance11 on RADV

  • OpenCL 3.1 support for rusticl on asahi, iris, radeonsi, llvmpipe and zink

  • VK_KHR_workgroup_memory_explicit_layout on pvr

  • VK_KHR_maintenance5 on pvr

  • VK_KHR_calibrated_timestamps on hasvk

  • VK_KHR_present_id on pvr

  • VK_KHR_present_wait on pvr

  • VK_KHR_present_id2 on hasvk

  • VK_KHR_present_wait2 on hasvk

  • VK_KHR_shader_fma on RADV

  • VK_EXT_shader_split_barrier on RADV/GFX12

  • VK_KHR_shader_fma on nvk

  • Support for G1-Ultra, G1-Premium and G1-Pro GPUs on Panfrost and PanVK

  • VK_EXT_shader_atomic_float on nvk

  • VK_{KHR,EXT}_index_type_uint8 on pvr

  • VK_EXT_debug_marker in vulkan runtime

  • VK_EXT_mesh_shader on NVK

  • VK_KHR_device_fault on RADV

  • VK_EXT_present_timing now also on wsi/x11

  • VK_EXT_present_timing on hasvk

  • VK_GOOGLE_display_timing for KHR_display (and opt-in for x11, wayland) on anv, hasvk, hk, nvk, panvk, pvr, radv, tu, v3dv

  • VK_KHR_shader_abort on RADV

  • VK_KHR_shader_fma on panvk

  • VK_NV_shader_atomic_float16_vector on NVK

  • VK_KHR_compute_shader_derivatives on panvk

  • VK_EXT_shader_subgroup_ballot on pvr

  • VK_EXT_shader_subgroup_vote on pvr

  • VK_KHR_shader_subgroup_rotate on pvr

  • VK_KHR_shader_subgroup_uniform_control_flow on pvr

  • VK_EXT_subgroup_size_control on pvr

  • VK_EXT_rasterization_order_attachment_access on panvk

  • VK_ARM_rasterization_order_attachment_access on panvk

  • GL_OES_texture_float on etnaviv/HALF_FLOAT

  • GL_ARB_texture_float on etnaviv/HALF_FLOAT

  • VK_NVX_binary_import on NVK

  • VK_EXT_image_sliced_view_of_3d on panvk

  • VK_EXT_shader_tile_image on panvk

  • VK_KHR_internally_synchronized_queues on panvk

  • VK_KHR_workgroup_memory_explicit_layout on panvk

  • VK_EXT_device_memory_report on pvr

  • VK_EXT_descriptor_heap enabled by default on anv, RADV

  • VK_EXT_device_address_binding_report on panvk

  • VK_EXT_multisampled_render_to_single_sampled on RADV/Android

  • VK_KHR_shader_fma on anv

  • VK_KHR_extended_flags on RADV

  • VK_EXT_display_surface_counter on pvr

  • VK_EXT_display_control on pvr

  • VK_EXT_direct_mode_display on pvr

  • VK_KHR_surface_maintenance1 on pvr

  • VK_EXT_surface_maintenance1 on pvr

  • VK_KHR_swapchain_maintenance1 on pvr

  • VK_EXT_swapchain_maintenance1 on pvr

  • VK_EXT_swapchain_colorspace on pvr

  • VK_EXT_acquire_drm_display on pvr

  • VK_KHR_unified_image_layouts on pvr

  • VK_KHR_shader_quad_control on v3dv

  • VK_KHR_shader_subgroup_rotate on v3dv

  • VK_KHR_shader_maximal_reconvergence on v3dv

  • VK_EXT_shader_image_atomic_int64 on panvk

  • VK_EXT_host_image_copy on RADV/GFX10.3+

Bug fixes

  • “bad target in _mesa_select_tex_object()” when using glXBindTexImageEXT

  • 7900XT, hardware acceleration crashes chromium-based apps unless LIBVA_DRIVER_NAME=null

  • A702: assertions at A6XX_TEX_MEMOBJ_4_BASE_LO

  • ANV: dEQP ASTC tests crash w/ FPE_INTDIV on Xe3

  • AV1 videos dropping frames with AMD card

  • After updating GStreamer, all videos in Showtime are green/purple

  • Ambient occlusion is broken in The Chronicles of Riddick - Assault on Dark Athena

  • Blockland crashes when shader quality turned up

  • Broken rendering in Isonzo (Unity) since vulkan-radeon 26.1.0

  • Celestia hit fallback on r300g from git?

  • Confidential issue #15670

  • Copy paste bug in `gallium/drivers/radeonsi/si_texture.c`

  • Copy paste bug in `vulkan/runtime/vk_debug_utils.c`

  • Crash linking cached GLSL shader with optimized array with explicit uniform location

  • D3D12: last vertex’s non-zero-offset attribute fetched as 0 with a tightly-packed interleaved vertex buffer (AMD D3D12 only, regression since 24.1)

  • DaVinci Resolve 20.2 opencl-mesa crash

  • Dishonored 2 flickering menus on BMG and DG2

  • Doom Eternal - Page Fault - somewhat reproducable. (RX9070)

  • Gunfire Reborn crashes with ring gfx_0.0.0 timeout since Mesa 26.0.0

  • Horizon Forbidden West misrendred lighting effects on BMG

  • Intel binaries using GLX crash on ARM Macs using llvmpipe

  • Is maxFragmentCombinedOutputResources=16 in Honeykrisp reflects an actual HW limit?

  • Issues with shadows on DOOM: The Dark Ages - Revelations DLC

  • Mesa fails to build due to rust bindings error

  • Minecraft with Complementary Shaders and Voxy LOD rendering issue on Radeon 680M on latest mesa main branch

  • NIR: loop unrolling should not create phis for constants defined in the loop

  • NVK The Surge 2 misrendering near text

  • NVK: Rendering corruption in Shadow of the Tomb Raider

  • OpenGL app deadlocks in loader_dri3_swap_buffers_msc → xcb_wait_for_special_event on radeonsi / Mesa 26.1 / X11 (Telegram Desktop)

  • Qt application TrenchBroom hangs in glXSwapBuffers in recent version of Mesa

  • RADV NIR optimization bug

  • RADV/RX 9070 XT: WUCHANG: Fallen Feathers gpu ring reset

  • RADV: RDNA2 Raytracing regression after “aco/lower_branches: Add try_rotate_latch_block() optimization”

  • RADV_DEBUG=bo_history crashes any vkd3d/dxvk game at startup with heap corruption (double free)

  • RADV_PERFTEST=transfer_queue causes assertion failure in prime scenario

  • Regression. VTK polydata. gl_PrimitiveID requires explicit geometry shader with versions 25.3.6 and later.

  • Rusticl causes a crash after compiling an OpenCL program

  • SIGABRT: Invalid free in zink_destroy_resource_surface_cache()

  • Spike in mmap count for vk_cmd_queue regressed dEQP-VK.api.object_management.max_concurrent#command_buffer_*

  • System freeze when seeking in h264 files with gst-play-1.0 using the VA plugin

  • The End is Nigh (Wine): No lighting in The Hollows

  • Uplink text rendering still bugged out

  • [AC/NIR] XPlane 12 failure to start

  • [ANV] Intel arc b580 | Halo Infinite misrenders and GPU hangs

  • [ANV][ARC] Space engineers 2 Artifacts

  • [ANV][LNL] - Split Fiction (2001120) - Vertex explosion on white cloth during tutorial.

  • [ANV][PTL] - Elden Ring (1245620) - Blue artifacts and flickering lights with raytracing enabled

  • [ANV][PTL] - Horizon: Forbidden West regression in shader quirk

  • [ANV][PTL] - Persona 3 Reload (2161700) - Blue color on some objects + reflections do not reflect

  • [Arc B580][Baldurs Gate 3] Hang on opening inventory

  • [BMG] regression: Artifacts in GTK applications in 26.0.x

  • [GR Breakpoint] Inverted frustum culling for grass meshes on Linux

  • [PTL] dEQP-VK.synchronization2.op.single_queue.event.write_fill_buffer_read_ubo_tess_control.buffer_16384_maintenance9 fails

  • [RADV/ACO]: (Bisected) Regression causes flashing reflections in CyperPunk 2077

  • [RADV] REGRESSION: Lighting Bleed-through in Need for Speed games on mesa 1:26.1.1-2 ArchLinux

  • [RADV] REGRESSION: Weird graphics glitches and Lighting Bleed-through in Saints Row 2

  • [RADV] Regression: Infinite loop in NIR compiler during nir_opt_dead_write_vars loading Gemma 4 via llama.cpp on gfx1100

  • [RADV] Regression: Sackboy - A big Adventure, crash if RT is enabled

  • [RADV] Video color artifacts in mpv

  • [RADV][RDNA3][regression] GPU context lost in Sushi Ben (Steam 2419240) when using in-game camera

  • [Security] gallium VA-API AV1 tile_info(): unbounded loop writes past fixed tile_col_start_sb/width_in_sbs arrays on heap

  • [VAAPI] VRAM leak on Polaris, 26.1 regression

  • [VAAPI][Feature Request] Support of VAProcPipelineCaps.blend_flags in radeonsi (vf_overlay_vaapi)

  • [VAOn12] Thread safety issues

  • [VC4/V3D] GL_EXT_shadow_samplers exposed but GL_TEXTURE_COMPARE_MODE/FUNC returns GL_INVALID_ENUM in GLES2

  • [Windows][arm64] Building windows ARM64 with MSVC fails on src\util\u_math.c

  • [Xe][ARC770][Baldurs Gate 3] Hang during shader compilation

  • [amdgpu] Little Inferno rendering issues starting with Mesa 21.1

  • [anv] Intel ARC B390 | Hitman 2 | DX11 | Blinking corruptions seen

  • [anv] Intel ARC B390 | Starfield | DX12 | Crash after starting the game

  • [anv] [bmg] Halo Infinite will not start rendering on an Arc B580

  • [anv] dpas/Vulkan coopmat uses ~double registers due to scalar lowering

  • [anv] negate of INT_MAX calculated as INT_MIN instead of INT_MIN + 1

  • [anv] wrong value for local variable after nested switch merge

  • [meson][etnaviv] Build target etnaviv_isa_rs has no sources

  • [r300] [big-endian] bad colors under memory pressure

  • [radeonsi/VCN] RX 9070 XT (Navi 48, gfx1201): VCN unified ring timeout during VAAPI HEVC encode from Steam Game Recording

  • [radeonsi] SIGSEGV in gallivm during software fallback for GL_FEEDBACK with lighting

  • [radeonsi] eglinfo ‘libEGL warning: failed to get driver name for fd -1’

  • [radv] Regression causes GPU page faults in Crimson Desert

  • [radv] Regression causes black patches on the ground in DOOM The Dark Ages Revelations DLC

  • android14 gpu virgl: Shader memory leak

  • anv: Assert in alias vkCreateImage

  • anv: Missing null check in vkCmdEndTransformFeedback

  • anv: NULL deref in anv_AllocateMemory when importing DMA-BUF without VkMemoryDedicatedAllocateInfo on Xe2 (regression)

  • anv: World of Warcraft dx11 CMAA 2 option buggy on a b580

  • anv: anv_load_fp64_shader consumes 12MB of heap at device creation

  • anv: atan shader compiler bug leads to negative infinity

  • anv: bottom-of-pipe timestamp latched before vkCmdDispatch completes (Arc B390 / Panther Lake, Mesa 26.1.4), breaking wgpu compute-pass timings

  • anv: cooperative matrix loads from shared memory return the wrong tile on Arc B390 (Xe3/PTL), matmul results are wrong

  • asahi: broken image? rendering in firefox with active hw-acceleration since mesa-26.0.5

  • brw/decoder: fragment shader decoding issue

  • brw: Blender performance regression for Barbershop benchmark scene

  • build: intel_eu_stall_viewer fails on 32-bit platform

  • ci: arm64 lava kernel missing zstd support

  • d3d12: stretched + black-bar artifact after fast offscreen FBO resize on Adreno (regression in MR 41322)

  • ethosu: Build error on 32-bit due to %lu on uint64_t

  • gallium/vl: scale_vaapi limited range RGB to limited range YUV broken with red tint

  • glcpp: incorrect macro expansion in token pasting

  • intel/blorp: VK_ANDROID_external_format regression on buildtype=debug

  • intel/blorp: shader_pipeline initialization results in incorrect blorp_key values

  • intel: Investigate HIZ Plane Optimization disable bit for gfx12.5

  • ir3: THREAD64→THREAD128 heuristic change in 25.2.0 causes vertical line artifacts in No Man’s Sky on Adreno 740 (a7xx)

  • ir3: ir3_cf brokenness

  • iris: Blender wireframe rendering broken

  • iris: unhandled -EAGAIN in iris_batch_flush (26.1 regression)

  • kk: Implement timestamps

  • kk: stencil test unexpectedly not working

  • lavapipe: dynamic state for advanced blend is broken

  • llvm23 breaks build of clc_helpers.cpp

  • llvm23 commit d50631f breaks build of ac_llvm_helper.cpp

  • macOS build stuck in infinite loop since 26.1

  • mediafoundation: error C2039: ‘step’: is not a member of ‘_inputQPSettings’

  • mesa 26.1.3 does not compile with ARM64 MSVC

  • mesa: clean up st_context caps flags

  • mesa: glthread crash with MESA_VERBOSE=api

  • nir: possible exactness bug in reassociate

  • nir_opt_copy_prop_vars validation failure with Rusticl OpenCL kernel shaders

  • nvk: various float_controls cts test fails on Turing only

  • nvk: zcull causes SAVE_RESTORE_ADDR_OOB crashes in Horizon Forbidden West

  • panvk: Failures in new (1.4.5.0) texturequerylod*trilinear CTS tests

  • r300 bisected : Broken rendering with applications using MSAA visuals

  • r300: dEQP-GLES2.functional.shaders.struct.uniform.(not_)equal_fragment regression

  • r600, sfn: Lowering to assembly failed on R700 while running ShooterGame demo

  • r600: SFN assersion failed in Transport fever 2

  • radeonsi+ACO jobs should use ACO_DEBUG=validatera

  • radeonsi/VCN: AV1/HEVC encode reports a coded size larger than the coded buffer, client segfaults reading the mapped buffer

  • radeonsi: sqtt missing data for viewperf

  • radv: Use more efficient cache uuid for hardware identifiers

  • radv: acceleration structure update (refit) of AABB geometry loses intersection candidates (NAVI32, Mesa 26.1.4)

  • radv: descriptor heaps do not support non-uniform indexing on buffer pointers

  • radv: incorrect memory accounting for imported bufffers

  • radv: use preload sgprs for 16/8bit push constant loads

  • regression;bisected;radeonsi/video: corrupted VA-API H.264 video encoding (bisected to 4487162a)

  • rusticl + v3d assertion failure

  • rusticl/radeonsi OpenCL behavior regression after 5a298f3560629943e9140c74547183da1635352e

  • rusticl: incorrect float16 constant folding

  • segfault v3d raspberry Pi5 h265 hw decoding / regression mesa 26.1.3/26.1.4

  • some dEQP-VK.image.host_image_copy failures on xe2/Xe3 with 32bit

  • std430 layout not supported on older GLSL profiles with GL_ARB_shader_storage_buffer_object extension

  • the new mesa 26.1.0-1 is causing mutliple graphical issues on intel i5-2400 and higher (integrated graphics)

  • tu: assertions in ir3_ra.c: Assertion `physreg != (physreg_t)~0’ failed

  • turnip: Blender: viewport contents invisible with UBWC

  • v3dv: Compute shader crashes for unknown reason

  • venus: rare flake in dEQP-VK.wsi.android.swapchain.render.basic

  • venus: typo in vn_queue_submit_2_to_1

  • vulkan/runtime: GetPipelineBinaryDataKHR incorrectly assumes `pPipelineBinaryDataSize` must be 0 initialized

  • vulkan/wsi/win32: bgcolor of Vulkan app window becomes lighter under venus

  • vulkan: 5.10 and 5.15 LTS kernels require implicit sync support

  • wsi: Suspicious assertion in wsi_create_buffer_blit_context

  • zink/ci: switch all traces jobs to surfaceless+gbm

Changes

Adam Jackson (7):

  • zink: consolidate resource_create error paths

  • zink: extract format list setup from create_image

  • zink: extract pNext chain construction from create_image

  • zink: extract memory binding from create_image

  • zink: replace image negotiation with candidate-based approach

  • zink: stop find_good_mod from mutating ici in place

  • nvk: use unsigned comparison in UBO bounds checks

Adam Stylinski (1):

  • nv30: fix an issue when this push buffer is NULL

Aditya Swarup (2):

  • anv/pps: Use counter block to stay consistent with Perfetto

  • intel/test: Add support for Perfetto counter groups

Adrián Larumbe (10):

  • pan/kmod: Fix minor version number check for USER_MMIO_OFFSET ioctl

  • pan/kmod: fix double syncop count sum when populating vm_bind syncs

  • drm-uapi: Sync the panthor header

  • pan/kmod: Use kernel-reported page sizes for new VM when available

  • pan/kmod: Pass signal and wait syncs separately

  • pan/kmod: Handle sync object signals in Panthor’s vm_bind

  • pan/kmod: Introduce sparse binding

  • pan/kmod: Introduce vm_op buffering and sparse mapping emulation

  • panvk: Use pankmod instead of panthor drm interfaces in bind queues

  • panvk: Talk directly to pankmod when binding sparse resources

Agate, Jesse (2):

  • amd/vpelib: separating frontend programming

  • amd/vpelib: Fix blending hang issue

Ahmed Hesham (10):

  • pan/bi: Restore b3210 as a valid swizzle

  • pan/nir: Fix 8 and 16 bool reduction lowering

  • pan/bi: Fix function temp lowering with 64-bit pointers

  • pan/bi: Fix MKVEC.v2i8 src2 swizzle lowering

  • pan: report async CSF group faults via context reset status

  • clc: fix fp16 fallback mask for remquo

  • nir: fix vectorising phis with mixed chased sources

  • rusticl: enable panfrost by default

  • ci: Add OpenCL-CTS to GL test infrastructure

  • pan/ci: Add OpenCL-CTS quick job for Mali-G610

Aitor Camacho (37):

  • kk: Add poly dependency to KK

  • kk: Reuse as much poly utilities as possible for unrolling

  • kk: Increase maxFragmentCombinedOutputResources to KK_MAX_DESCRIPTORS

  • kk: Add ds state to fragment key since it’s part of the pipeline we compile

  • kk: Fix subgroup failures on M1/2 due to bcsel

  • kk: Fix global_store writemask

  • kk: Rewrite force position output pass to use lowered io

  • kk: Add residency set to queues

  • kk: Add grid struct for dispatches for convenience

  • kk: Correctly report failures when compiling precompiled shaders

  • kk: Rework draw dispatch

  • kk: Rework shader compilation to handle more than 2 stages

  • kk: Implement tessellation

  • kk: Use index element size instead of Metal enum to avoid asserts

  • kk: Use subgroups for tessellation prefix count since they are now fixed

  • kk: Move poly data out of root buffer

  • kk: Clean up per draw upload for tessellation stage

  • kk: Move to Metal4 command encoding

  • kk: Add GPU hang detection

  • kk: Disable workarounds 1-6 in macOS 27

  • kk: Reduce root buffer pointer by replacing it with the GPU address

  • kk: Expose texture max dimensions based on GPU family

  • poly/lower_tcs: Preserve TCS barriers against lowered outputs

  • kk: refold combined image/sampler packing after vars_to_ssa

  • kk: defer cmd buffer submission and lighten compute barriers

  • kk: Record command buffers live and replay only on resubmit

  • kk: Fix flrp signed-zero preservation with float_controls2

  • nir: compute acos in 32-bit then downgrade to 16-bit

  • kk: Fix metal import assert

  • kk: Handle per alu math controls in MSL

  • kk: Use isnan(x) for x != x and !isnan(x) for x == x

  • kk: Compile all shaders with fast math

  • kk: Expose shaderSignedZeroInfNanPreserveFloat16/32

  • kk: Implement VK_KHR_shader_float_controls2

  • kk: Metal’s precise functions are only fp32

  • kk: expose Vulkan 1.4

  • kk: Fix icd json api version

Alejandro Piñeiro (9):

  • pan/midgard: reorder nir_shader_compiler_options alphabetically

  • panfrost: add explicit casts when assigning ~0 and 0 to enum-typed fields

  • cs_builder: fix trailing comma in cs_builder_init

  • panfrost: track active endpoint scoreboard slot

  • panfrost: define scoreboard slots

  • panfrost: add Perfetto render stage tracing support (v10+)

  • panfrost: add u_trace indirect capture for compute dispatches (v10+)

  • panfrost: add Batch, Barrier and CacheFlush Perfetto render stage (v10+)

  • docs/perfetto: panfrost now supports render stages

Aleksi Sapon (2):

  • llvmpipe: remove unused SSE rasterization code

  • llvmpipe: fix overflow in rasterizer

Alessandro Astone (4):

  • gallivm: Fix armhf build against LLVM 22

  • radv: Support explicit DRM format modifier from android gralloc

  • anv: Support VK_ANDROID_native_buffer older than version 11

  • hasvk: Fix android build with android-strict=false

Alexander Slobodeniuk (1):

  • radeonsi: fix conformance window emission in the SPS

Ali, Nawwar (1):

  • amd/vpelib: update shaper config size

Allen Ballway (2):

  • vulkan/android: Set COLOR_ATTACHMENT_BIT for external format resolve

  • vulkan/android: Map AHARDWAREBUFFER_FORMAT_Y8 to VK_FORMAT_R8_UNORM

Alyssa Rosenzweig (288):

  • nir/opt_generate_bfi: avoid trivial instructions

  • jay: strengthen assert

  • jay: drop dead code

  • jay: generalize last kill code

  • jay: reduce zeroing

  • jay: reduce calloc to malloc when memsetting after

  • jay: fix SEL implied pipe

  • jay: fix the source pinning code

  • jay/register_allocate: use standard builder name

  • jay/opt_dead_code: handle predication

  • jay/lower_pre_ra: skip predication

  • jay/assign_flags: handle predicated CMP

  • jay/register_allocate: tie predicated-defaults

  • jay: allow predication of pure-flag instrs

  • jay: fold logic ops

  • jay/test-optimizer: fuse before/after cases

  • jay: test logic op fusing

  • jay: fix simd32 deswizzle

  • jay: improve spiller debug

  • jay: fix spiller coupling code

  • jay: don’t print internal without the flag

  • jay/register_allocate: don’t depend on indexing

  • jay: call DCE an extra time

  • jay: refuse to propagate ADDRESS copies

  • jay: relax mov type check

  • jay: fix SEL types

  • jay/print: deal with bare r0 copies

  • jay/to_binary: handle packing accumulators

  • jay: validate non-SSA accumulators

  • jay/register_allocate: start using accumulators

  • jay/lower_post_ra: remove SWAP macro

  • jay/lower_post_ra: drop old 2<–>8 lowering

  • jay/ra: don’t reserve registers when not spilling

  • jay/ra: use accumulator for memory copies

  • jay/ra: use accumulator for memory swaps

  • jay/ra: use accumulator for stride=4 swaps

  • jay/ra: drop memory copy reordering

  • jay/ra: only use stride=4 temps

  • gallium: Drop users of post-processing filters

  • gallium: Drop post-processing filters

  • nir/opt_algebraic: add redundant u2u32/unpack_64_2x32_split_x patterns

  • brw/nir_lower_cs_intrinsics: do some math at 16-bit

  • nir/opt_reassociate: fix exactness bug

  • jay/assign_flags: refactor for next commit

  • jay/assign_flags: don’t burn a null flag

  • jay/assign_flags: don’t burn a flag for ballots

  • jay: shrink stack allocation

  • jay: jayize swsb print

  • jay: consolidate file prefixes

  • jay: fix 16-bit predicated compares

  • jay: drop jay_exec_mask

  • jay: inline jay_control()

  • jay/opt_propagate: fold uflag copies

  • jay/opt_propagate: disable f64 opts for now

  • jay: introduce a physical control flow graph

  • jay: drop UGPR->UMEM spilling path

  • jay: convert to LCSSA

  • jay: do not copyprop ballots globally

  • jay: propagate inverse-ballots only locally

  • jay: check for inverse-ballots in jay_uses_flag

  • jay: smarten predication pass

  • jay: adjust flag replication

  • jay: predicate NoMask instructions in uniform IF’s

  • jay: drop a bunch of stale TODO and XXX

  • jay/to_binary: rename grf -> phys_reg

  • jay/to_binary: fix packing of simd-split accumulators

  • jay: model MAC

  • jay: do moves on the float pipe where possible

  • jay: move simd32 deswizzling to float pipe

  • jay: assign accumulators post-RA

  • jay/lower_scoreboard: elide more dependencies

  • jay/lower_scoreboard: refactor wait pipe code

  • jay/lower_scoreboard: fix tracking for A@* and *@7

  • jay/lower_scoreboard: refactor SYNC.nop insertion

  • jay/lower_scoreboard: use .src annotations

  • jay/lower_scoreboard: be the sole emitter of SYNC

  • jay/lower_scoreboard: use SYNC.allrd/allwr

  • jay: swap predication/acc pass order

  • jay: add JAY_DEBUG=noacc option

  • jay: fix bfn with 0xffff constant

  • jay: elide atomic dests

  • jay: optimize pack_32_2x16_split(#0, x)

  • jay: make indirect push data blow up more obviously

  • jay: fix comment

  • jay: have proper UNDEF

  • jay: clarify development model

  • jay/lower_scoreboard: fix trivial scheduling

  • jay/lower_scoreboard: refactor

  • jay/lower_scoreboard: run RegDist globally

  • jay/lower_scoreboard: factor regdist logic out

  • jay/lower_scoreboard: control flow is int pipe

  • jay/lower_scoreboard: compact inst_exec_pipe

  • jay/lower_scoreboard: rename gpr_range -> key

  • jay/lower_scoreboard: use CFG for RegDist scoreboarding

  • jay/lower_scoreboard: use sbid syncs to elide regdist deps

  • jay/register_allocate: set num_regs[MEM] properly

  • jay/register_allocate: tweak roundrobin heuristic

  • jay/opt_propagate: fix NOT propagation

  • jay/opt_propagate: propagate undefs

  • pan/mdg: make clang warning quiet

  • jay: relax fragment payload layout

  • jay: insert simd32 deswizzle in a dedicated pass

  • jay/liveness: remove pointless bitset init

  • jay/liveness: speed up physical CFG merging

  • jay/liveness: drop redundant source filtering

  • jay/lower_scoreboard: handle accumulator hazard

  • jay/lower_scoreboard: add asserts on key bounds

  • jay/opt_propagate: avoid branching on poison

  • jay: annotate pure sends

  • jay: factor jay_op_(starts,ends)_block queries

  • jay: schedule for pressure

  • jay: fix omask on single sample

  • jay: hack for sample position

  • jay: use new fs payload variable more

  • brw/eu_validate: relax EOT requirements on Xe2

  • CODEOWNERS: add Jay

  • brw,jay: add use_src_xy prog data field

  • jay/liveness: use jay_foreach_preload

  • jay: legalize shuffle(ugpr) for now

  • jay/spill: fix reload array size issues

  • jay/register_allocate: inline silly helper

  • jay/register_allocate: remove out of date comment

  • jay/lower_spill: use 1 less temporary

  • jay: fix FS reading too many sysvals

  • jay: add zip_ugpr16 instruction

  • jay: pack jay_stride

  • jay: stop asking for stride=4 ugpr’s

  • jay/validate_ra: use jay_def_stride

  • jay: drop unneeded #include

  • jay: rewrite partition handling

  • jay: merge partition blocks

  • jay: remove send split hack

  • jay: allow simd32 gl_SamplePosition

  • jay: avoid overflow affinities with large UGPR vecs

  • jay: allow SIMD1 imageStore()

  • jay: hide MAD->MAC behind !JAY_DEBUG=strict

  • jay: gate early EOT code behind =strict

  • jay: renumber reg files predictably

  • jay: simplify uniformity checks

  • jay: allow npot operands in RA

  • brw: nir_lower_constant_convert_alu_types only once

  • intel/gen: remove dead #include

  • intel/gen: drop noisy build spam

  • jay/validate_ra: validate against partition

  • jay/register_allocate: do not treat reserved regs as free

  • jay/register_allocate: remove remnant of old partition code

  • jay/register_allocate: split out jay_stride.c

  • jay/register_allocate: drop #include

  • jay/partition: validate we don’t generate g127<2>

  • jay/partition: pick better partitions

  • jay: limit stencil export to simd16

  • jay/to_binary: relax packed float restriction

  • jay: generalize jay_extract_range_post_ra

  • jay: rework lane ID calculations

  • jay: replace GPR_FROM_UGPRs with a simple CVT

  • jay: replace BYTE/WORD_PACK with a simple MOV

  • jay/register_allocate: don’t hang if a block is missing

  • jay/partition: reduce 16-bit partitioning more

  • jay: drop dead if

  • jay: clang-format

  • jay/lower_scoreboard: allow multiple jumps

  • jay: lower JAY_OPCODE_LOOP_ONCE earlier

  • jay/to_binary: big clean up post-gen

  • jay: avoid bogus copyprop with cmods

  • jay: add unit test for bogus copyprop case

  • jay/validate: add validation for bogus uflag cases

  • jay: fix 8/16-bit inline_data loads

  • jay/assign_flags: require ballots to be in the balloted src

  • jay: fix mismatched files with predication

  • jay: workaround the while bug

  • jay: follow source order for mad/bfe

  • jay: fix last-use accounting with ARF sources

  • jay: allow null in jay_collect_vectors

  • jay: uniformize bti indirects

  • jay: optimize out more early eot related copies

  • jay/register_allocate: make phi webs conservative

  • bin: add drm-shim script

  • jay/lower_pre_ra: allow immediate on bfe

  • anv: enable jay ray query

  • nir/opt_sink: sink more Intel block instructions

  • nir/lower_terminate_to_demote: tweak terminate_if lowering

  • jay/spill: spill at definitions

  • jay/spill: do initial find-and-replace for ugpr spilling

  • jay/spill: unstub rematerialization

  • jay/spill: refactor

  • jay/spill: drop sketchy heuristic

  • jay/spill: simplify limit()

  • jay: fix bogus unit tests

  • jay/validate: check for mixed ugpr/gpr problems

  • jay/register_allocate: simplify split copy logic

  • jay/register_allocate: don’t search a 2nd UGPR temp

  • jay: allow more 3-src imms

  • jay: introduce accumulators into the partition

  • jay: pool constants per block under pressure

  • jay: remove #include

  • jay: cache message headers locally

  • jay: forbid 8-bit immediate prop

  • jay: autopep8

  • jay: manually format jay_type_for_glsl_base_type

  • jay: clang-format

  • jay: track skip_helpers

  • jay: rewrite demote/terminate/helper/halt handling

  • nir/opt_dead_cf: delete redundant returns/halts

  • intel, nir: Add {load,store}_global_intel intrinsics

  • jay: distinguish physical & logical loop headers

  • jay/validate: validate backedges

  • jay/lower_scoreboard: simplify trivial swsb

  • jay/lower_scoreboard: fix barriers in trivial SWSB

  • jay/test: drop SSA repair tests

  • jay/to_binary: use dedicated addr reg for shuffle

  • jay/lower_pre_ra: fix oob read

  • jay/register_allocate: fix file prefixes

  • jay/opt_dead_code: drop stale todo

  • jay/opt_dead_code: handle phi properly

  • jay/spill: add an assert()

  • jay/spill: repair as we go

  • jay/spill: implement ugpr spilling

  • jay/spill: don’t try to remat mov_imm64

  • jay/spill: do lazy reloading instead

  • jay: remove unused SSA repair pass

  • jay: drop #include

  • jay/assign_flags: fix ballot handling

  • jay: fix barycentrics

  • jay: improve the stride partition heuristic

  • jay/lower_spill: rename to make easier to follow regs

  • jay/lower_pre_ra: fix f64 negate

  • jay: clang-format

  • Revert “rusticl: fix leak in `util_queue`”

  • jay: fix EOT with indirect message descriptors

  • intel: make more multisampling/coarse state static

  • jay/to_binary: avoid overflowing address reg

  • jay: add jay_bare_regs helper

  • jay: use jay_bare_regs

  • jay: drop jay_extract_range_post_ra

  • jay: rename post-sched lowering

  • jay: zero a0 before divergent shuffles

  • jay: consolidate spilling call into a single file

  • jay: fix JAY_DEBUG=spill

  • jay/lower_spill: set cursor explicitly

  • jay: fix ugpr reloads in divergent control flow

  • jay: fix printing internal shaders

  • people: add Lucas Fryzek

  • util: add __bitset_zero, __bitset_copy helpers

  • treewide: use __bitset_{zero,copy}

  • jay: remove #define

  • jay: unify inline_data and push_data loads

  • jay: remove duplicated bf16 check

  • jay: remove pointless (and wrong) bf16 check

  • jay: drop #include

  • jay: treat init_helpers as a start instruction

  • jay: print the shader after spilling before RA

  • jay/print: make predication syntax more explicit

  • jay/schedule: move schedule into ctx

  • jay/schedule: move function into context

  • jay/dag: inline dag code

  • jay/dag: defer parent_count calculation

  • jay/dag: split into a mutable and immutable part

  • jay/dag: add jay_dag_print helper for debug

  • jay/dag: add jay_dag_iterator_reset helper

  • jay/schedule: add per-block data structure

  • jay/schedule: edit comment about pressure

  • jay/schedule: split block analysis from scheduling

  • jay/schedule: flip the signs on pressure calcs

  • jay/schedule: account for demand per-file

  • jay/schedule: fix accounting for dead flags

  • jay/schedule: add missing whitespace

  • jay/schedule: early-out invalid schedules as we go

  • jay: add mlen to SEND

  • jay: add real cycle model

  • nir: lower boolean shuffle without subgroup size

  • nir/opt_algebraic: optimize ~x != x

  • gen/print: include pipes for jay_print

  • intel/gen: allow dead loops

  • iris: use u_default_get_sample_position

  • crocus: use u_default_get_sample_position

  • intel: remove intel_get_sample_positions getter

  • anv: use vk_standard_sample_locations

  • hasvk: use vk_standard_sample_locations

  • jay/spill: fix variable shadowing

  • jay/lower_post_ra: fix simd16 flag zeroing

  • jay/opt_propagate: fix inverse_ballot(bfn)

  • jay/lower_helpers: fix unconditional discard

  • jay: lower boolean shuffles

  • jay: make uniformity explicit

  • jay: support as_uniform

  • jay: drop #include

  • jay: rewrite flags

  • intel: fuse off Jay in Mesa 26.2

Andrzej Datczuk (2):

  • radv: enable advertising of VK_KHR_pipeline_library under llvm

  • radv/rra,rmv: fix device id written into trace files

Anna Maniscalco (1):

  • ir3: Skip preftech and and warmups for non bindless earlier

Arjob Mukherjee (2):

  • pvr: increase value of maxPerStageDescriptorStorageBuffers

  • pvr: increase maxPerStageDescriptorStorageBuffers to 16

Arkady Shlykov (1):

  • Add offset getter for image intrinsics

Arzaq Naufail Khan (1):

  • spirv: fix resource leak in spirv shader replacement

Ashley Smith (2):

  • panfrost: Fix tiler_desc assignment

  • panfrost: Avoid race between bo import and unref

Assadian, Navid (1):

  • amd/vpelib: FP16 non linear handling

Autumn Ashton (5):

  • nak: Expose max_warps_per_sm

  • nvk: Allow nvk_cmd_upload_qmd to take a custom root descriptor

  • nvk: Add nvk_cmd_dispatch_with_root

  • nouveau/cubin: Add cubin and fatbin parsers

  • nvk: Implement VK_NVX_binary_import

Benjamin Cheng (19):

  • ac/vcn: Rename VCN5 swizzle mode to GFX12

  • radv/video_enc: Use correct swizzle mode for VCN5 with GFX11

  • radv/wsi: Re-use transfer queue if it exists

  • ac/parse_ib: Add parsing for variable slice mode

  • radeonsi/video: Cleanup dpb buffer

  • radeonsi/mm: Disable variable slices when bad input is found

  • util/ycbcr: Fix adjust_to_range

  • util/ycbcr: Add a narrow range RGB coeff helper

  • gallium/vl: Fix RGB narrow range conversions

  • draw: Add lower_opcodes NIR pass

  • mesa/st: run the lower_opcodes pass for draw shaders

  • radv/video: Set accurate minQp/QIndex

  • radv/video: Report MULTIPLE_SLICE_SEGMENTS_PER_TILE_BIT

  • ac/video: Add {min,max}_qp to video enc caps

  • radv/video: Use {min,max}_qp caps from ac

  • mesa/st: use col0 attrib from provoking vertex for feedback

  • ac/surface: Remove GFX10 limitation for FORCE_SWIZZLE_MODE

  • ac/vcn_enc: Disable var slice for preencode with VCN4

  • vl/proc: Guard compositor creation with HAVE_GFX_COMPUTE

Benjamin Gaignard (1):

  • pan/format: Advertise support for AFBC(32x8,sparse)

Benjamin Otte (1):

  • Revert “lavapipe: Don’t advertise support for multiplane drm formats”

Benoît du Garreau (1):

  • docs: Add many missing features

Bo Hu (3):

  • codegen: update scripts/cereal/decoder.py

  • vk-snapshot: do not generate code to save vkQueueFlushCommandsGOOGLE

  • gfxstream: vk-snapshot: update handling of bufferview in vkUpdateDescriptorSets

Boris Brezillon (9):

  • pan/kmod: Don’t pass drmVersionPtr objects around

  • kraid: Fix cross-build

  • pan/kmod: Add a pan_kmod_timestamp_cycles_to_ns() helper

  • pan/props: Make pan_query_core_count() safe with wide shader_present bitmaps

  • pan/props: Split core_count/core_id_range into two helpers

  • pan/props: Add pan_query_perf_counter_per_block()

  • pan/perf: Start relying on canonical mali_perf definitions

  • pan/perf: Replace pan_perf_init() by pan_perf_{create,destroy}()

  • pan/perf: Transition to auto-generated derived counters

Boyuan Zhang (1):

  • radeonsi/mm: use non-tmz buf when session tmz size is 0

Brandon Jones (1):

  • nir/opt_algebraic: fix fabs optimization

Brendan King (1):

  • pvr: move some asserts in pvr_srv_alloc_display_pmr

Caio Oliveira (115):

  • brw: Don’t set saturate for SYNC instruction

  • brw: Use brw prefix to LSC helpers tied to brw

  • brw: Remove various unused fields

  • brw: Fix max_dispatch_width collection for CS with variable size

  • brw: Stop tracking inline parameter usage in prog_key/prog_data

  • intel/dev: Expose list of known platform names

  • brw: Fix some indentation in brw_generator.cpp

  • brw: Move brw_prog_data_init to a different file

  • brw: Remove references to SIMD4x2

  • intel/executor: Map the DPAS check to has_systolic

  • brw/tests: Remove redundant parser test

  • brw/tests: Stop using regions/type for null in assembler tests

  • brw/tests: Stop using regions/type for non-null SEND sources in tests

  • anv: Remove saturating cmat configurations when INTEL_LOWER_DPAS=1

  • anv: When using INTEL_LOWER_DPAS disable BFloat16 cmat configurations

  • intel: Move cmat configurations to anv_physical_device

  • intel/perf: Use intel_perf_context as ralloc parent of sample buffers

  • nir/instr_set: Fix multi-slot intrinsic index equality

  • intel/perf: Add helpers to get names of enums

  • intel/perf: Show type, data type and units in intel_perf_query_layout

  • nir/instr_set: Consider normalization when calculating hash

  • brw/scoreboard: Add disabled tests for RegDist baking on Xe2+

  • brw: Save original regs_written() value in register coalesce

  • brw: Call size_read() once in regs_read()

  • brw: Avoid unnecessary calls to size_read() in flags_read()

  • brw: Pass VGRF numbers to liveness helpers

  • brw: Don’t directly use regs_read/regs_written/size_read as bound for non-trivial loops

  • brw: Bound register coalesce rewrites by live range

  • nir: Add print for other cmat_description slots

  • compiler: Support more than 255 cols/rows in cmat descriptions

  • intel/compiler: Move bison command to shared meson.build

  • intel/executor: Add an overflow check for alloc function

  • intel/executor: Add performance counter support

  • brw: Use a single brw_compile entrypoint

  • brw: Move key and prog_data to base compile params

  • anv: Simplify code that calls brw/jay

  • iris: Simplify code that calls brw/jay

  • spirv: Stop warning about ignored invalid ArrayStride decorations

  • anv, brw: Use previous shader VUE map for FS input layout when available

  • anv: Fill VERTEX_ELEMENT_STATE before further emissions

  • anv: Use empty_vs_input for default VERTEX_ELEMENT_STATE

  • jay: Use TGL_PIPE_NONE when RegDist is zero

  • intel/gen: Add gen encoding module

  • intel/gen: Add various to_string/from_string functions

  • intel/gen: Add validation

  • intel/gen: Add function to finish structured control flow

  • intel/gen: Add gen_print()

  • intel/gen: Add gen_parse()

  • Revert “intel/dev: Remove unused intel_get_device_info_for_build() function”

  • intel/gen: Add integrated `gentool` CLI for asm/disasm

  • brw: Add temporary workarounds for compatibility with old parser

  • brw/tests: Port assembler tests to gen module

  • intel/compiler: Stop replicating narrow immediate values in brw_asm “compat” mode

  • intel/compiler: Stop forcing null source <0;1,0> region in brw_asm “compat” mode

  • intel/compiler: Stop forcing null destination HSTRIDE=1 on pre-xe SEND in brw_asm “compat” mode

  • brw, jay: Use lsc_* symbols from gen

  • brw, jay: Use gen_swsb and related enums

  • brw, jay: Use gen_sfid instead of brw_sfid

  • brw, jay: Use various enums from gen instead of brw

  • jay: Use gen_condition enum and helper

  • jay: Use gen module

  • intel/decoder: Convert to use gen module

  • intel/executor: Update to use gen module

  • brw: Make a copy of brw_generator for using gen

  • brw: Port the copy of generator to use gen_encoding

  • brw: Use the new gen based generator

  • brw: Add brw_to_binary() as the single codegen entrypoint

  • brw: Make brw_generator an implementation detail of brw_to_binary.cpp

  • anv: Use gen_print instead of brw_disasm

  • intel/tools: Use gen_print instead of brw_disasm

  • iris: Use gen_print instead of brw_disasm

  • intel/compiler: Add and use gen_update_reloc_imm()

  • brw: Remove old encoder, generator and related tools

  • intel/gen: Change validation test code to use the parser

  • jay: Unify macro for NIR passes

  • jay: Add INTEL_DEBUG=mda support

  • util: Add runtime parser for boolean lookup tables

  • intel/gen: Support symbolic print/parse of BFN function

  • jay: Add helpers for managing unordered instructions

  • jay: Handle dpas_intel intrinsic

  • util: Fix float8 denorm rounding to min-normal

  • intel: build compiler before blorp

  • intel/gen: Generate opcodes and their metadata

  • intel/gen: Drop unused format parameter from gen_inst_has_dst

  • intel/gen: Don’t encode zero exec_size

  • intel/gen: Add a gentool ‘check-roundtrip’ subcommand

  • intel/gen: Replace gen_parse_print_test.cpp with text tests

  • intel/gen: Import test cases from the brw assembler tests

  • brw: Remove the brw assembler tests

  • nir: Handle nir_var_mem_push_const in divergence analysis

  • intel: Change dpas_intel source order to follow DPAS

  • jay: Handle convert_cmat_intel intrinsic

  • jay: Add SIMD restriction for Math with HF

  • brw: Drop dead spill_writes in brw_opt_fill_and_spill

  • brw: Add unit tests for brw_opt_predicate_logic

  • brw: Refactor DPAS lowering for HF case

  • brw: Fix INTEL_LOWER_DPAS=1 for Xe2+

  • util: Add util_is_half_subnormal()

  • brw, elk: Fix invalid case using float-negation in combine constants

  • intel/gen: Use lookup tables for Gfx12+ short type encoding/decoding

  • anv: Initialize shader debug archive key size

  • brw: Fix comma placement when printing memory logical sources

  • brw: Use shared LSC opcode names in IR printing

  • brw: Remove dead mixed-size MOV special case from def copy propagation

  • brw: Fix initializing matrix with uniform but non-constant values

  • intel/gen: Fix case where IF is a loop header

  • brw: Report scratch memory size in shader stats

  • brw: Track logical scratch offsets for spill cleanup

  • brw: Reuse scratch slots between non-interfering spilled VGRFs

  • nir: Account for cmat memory accesses in copy_prop_vars

  • anv: Include build and device identity in shader binary UUID

  • brw: Match fill/spill optimization scratch accesses by logical offset

  • intel/gen: Fix Gfx11 3-src accumulator file encoding

  • nir, spirv: Match debug printf argument alignment with u_printf

  • nir, spirv: Pad 3-component debug printf arguments to 4 components

Caius-Moldovan-img (5):

  • pco: Replace nir_shader_lower_instructions with nir_shader_*_pass

  • pvr: set SMP component count for TQ frag load shaders

  • pco: Fix metadata invalidation

  • nir: Fix trailing comment generation for variable naming

  • pco: Remove hardcoded metadata location

Calder Young (40):

  • anv: Fix address bit masking for indirect SBTs

  • anv: Fix support for indirect SBTs on Xe3+

  • anv: Store batch buffers in a null-initialized VMA heap

  • anv: Add padding to the shader heap to manage EU prefetch

  • isl: Add usage flag to force SurfaceArray to false

  • isl: Add additional alignment/padding requirements to prevent overfetch

  • isl: Optimize the sampler cache to overlap as few 64B cachelines as possible

  • isl: Add function to calculate the amount of overfetch for an unpadded surface

  • isl: Add and use isl_tiling_get_intratile_range_el/sa

  • blorp: Work around sampler overfetch for buffer copies

  • anv: Make sure robust UBO access does not fault

  • brw: Avoid rounding every convergent block load up to a full register

  • brw: Avoid vectorizing loads in NIR if it could extend into a different page

  • anv: Disable scratch page by default on Xe KMD

  • intel_hang_replay: Don’t force scratch page on Xe KMD unless explicitly requested

  • isl: Make sure isl_device::requires_padding is always initialized

  • anv: Fix some usage flags not propagated to ISL for explicit layouts

  • brw: Allow instruction reordering around memory writes

  • brw: Add support for ACCESS_CAN_REORDER memory ordering

  • spirv: Fix debugPrintfEXT not working with multiple arguments

  • brw: Add workaround pass for shaders using derivatives in control flow

  • anv: Add workaround for vertex explosions in Split Fiction

  • jay: Use gen_arf enums instead of jay_arf

  • jay: Use gen_names.h to print CMODs and ARFs

  • jay: Do not propagate ARF src unless its src0

  • jay: Add support for saturating f2i16 and f2i8 NIR opcodes

  • jay: Disable avoid_ternary_with_two_constants when using jay

  • brw: Move ray payload bitfield generation to NIR

  • brw: Move topology id helper intrinsics to NIR

  • jay: Disable SIMD32 if ray queries are used

  • jay: Implement ray tracing topology id intrinsics

  • jay: Implement ray tracing trace intrinsics

  • nir: Do not mask helper lanes of writes if ACCESS_INCLUDE_HELPERS is set

  • anv: Track more error codes from certain IOCTLs

  • intel: Add common utils for page fault reporting

  • anv: Add function to get the list of page faults

  • anv: Print page faults whenever a queue gets banned

  • anv: Enable support for VK_EXT_device_fault/VK_KHR_device_fault

  • intel: Add function to get the number of SBIDs from device info

  • jay: Implement the global SBID scoreboarding pass

Caleb Callaway (6):

  • docs: fix Intel tracepoints.py path

  • perfetto: v56.1 update

  • pps: generate forward-looking counters

  • perfetto: suppress array-bounds warning

  • perfetto: suppress stringop-overflow warning

  • anv: hide Intel vendor ID for Cyberpunk

Casey Bowman (2):

  • anv: Add option to disable HiZ via drirc

  • anv: Set anv_disable_hiz for Sons of the Forest

Charmaine Lee (1):

  • nir_to_tgsi: fix shared memory index

Christian Gmeiner (149):

  • mesa/st: Extend st_context_invalidate_state with meta-op flags

  • mesa/st: Convert st_cb_drawtex to use st_context_invalidate_state

  • mesa/st: Convert st_cb_bitmap to use st_context_invalidate_state

  • mesa/st: Convert st_cb_clear to use st_context_invalidate_state

  • mesa/st: Convert st_cb_drawpixels to use st_context_invalidate_state

  • mesa/st: Convert st_cb_readpixels to use st_context_invalidate_state

  • mesa/st: Convert st_cb_texture to use st_context_invalidate_state

  • panvk: Advertise VK_EXT_shader_uniform_buffer_unsized_array

  • panvk: Advertise VK_EXT_dynamic_rendering_unused_attachments

  • panvk: Implement vkCmdFillBuffer with panlib kernels

  • panvk: Wire up VK_EXT_conservative_rasterization on v11+

  • egl: Switch to mesa_log(..)

  • etnaviv: Bypass BGRA-internal optimization for shared resources

  • etnaviv: Convert PE-internal BGRA to RGBA when flushing shared resources

  • etnaviv: Select texture format dynamically for shared RB_SWAP resources

  • etnaviv: Add per-RT frag_rb_swap shader key and NIR lowering

  • etnaviv: Use shader R/B swap for LINEAR_PE shared resources

  • panvk: Apply sample mask in single-sample mode

  • panvk: Advertise VK_EXT_extended_dynamic_state3

  • compiler/rust: Move VecPair from NAK to shared compiler crate

  • st/mesa: Zero MaxTextureImageUnits for unsupported stages

  • etnaviv: blt: Add sRGB support to blt_imginfo

  • etnaviv: Map R8G8B8A8_SRGB to BLT_FORMAT_A8R8G8B8

  • etnaviv: blt: Add BLT format conversion support

  • lavapipe: Skip advanced blend lowering when blending is disabled

  • lavapipe: Lower advanced blend at draw time when its state is dynamic

  • lavapipe: Enable extendedDynamicState3ColorBlendAdvanced

  • mesa: Allow GL_TEXTURE_IMMUTABLE_LEVELS query on GLES3

  • mesa/main: Add trace dispatch plumbing for MESA_VERBOSE=api

  • mesa/main: Auto-generate MESA_VERBOSE=api trace dispatch

  • etnaviv: Update headers from rnndb

  • etnaviv: Use the full 64-bit clear value for 64bpp render targets

  • etnaviv: Use integer texture formats for R32/RG32 integer textures

  • etnaviv: blt: Don’t sRGB-roundtrip same-encoding copies

  • etnaviv: Drop unused num_loops shader stat

  • vulkan/wsi: Constify wsi_instance_supports_google_display_timing(..)

  • panvk: Advertise VK_GOOGLE_display_timing

  • panvk: Advertise VK_KHR_shader_fma

  • panvk: Shuffle local ids for quad derivatives

  • panvk: Advertise VK_KHR_compute_shader_derivatives

  • compiler/rust: move ACORN PRNG to shared location

  • panvk: Move maxFramebuffer limits to defines

  • panvk: Derive viewport limits from the framebuffer dimension

  • vulkan/runtime: Track rasterization_order_access in pipeline state

  • vulkan/runtime: Add rasterization_order_access to dynamic graphics state

  • panvk: Disable FPK and force late ZS for rasterization order access

  • panvk: Advertise VK_EXT_rasterization_order_attachment_access

  • etnaviv: Flush texture caches after clears

  • etnaviv: Add a helper for the 128-bit second-plane offset

  • etnaviv: Wrap pipe_framebuffer_state in etna_framebuffer_state

  • etnaviv: Lay out 128-bit color FBO as paired G32R32F render targets

  • etnaviv: Emit paired 128-bit sampler descriptors

  • etnaviv: NIR pass to lower 128-bit color RT and texture access

  • etnaviv: blt: Use block-layout offset for 128-bit second-plane blit

  • etnaviv: Save the framebuffer without 128-bit companion slots

  • etnaviv: rs: Support 128-bit color clears

  • etnaviv: Limit nir_lower_fragcolor(..) to advertised render targets

  • etnaviv: Advertise 128-bit color formats as renderable and samplable

  • etnaviv: Disable TS per render target on mixed TS modes

  • etnaviv: Update headers from rnndb

  • etnaviv: Set per-RT sRGB bit on non-zero render target slots

  • etnaviv: Support split sampler for 128-bit formats on the state path

  • etnaviv: Gate 128-bit render targets on HALF_FLOAT

  • pan/texture: Reuse the layer range as a Z-slice range on 3D views

  • panvk: Slice 3D storage image views on Valhall

  • panvk: Slice 3D storage image views on Bifrost

  • panvk: Advertise VK_EXT_image_sliced_view_of_3d

  • nir: Add load_tile_image intrinsic

  • spirv: Implement SPV_EXT_shader_tile_image

  • panvk: Lower tile image reads

  • panvk: Advertise VK_EXT_shader_tile_image

  • mesa/main: Keep RealPublished in sync when glthread toggles with api trace

  • panvk: Iterate the common queue list instead of a per-family array

  • panvk: Advertise VK_KHR_internally_synchronized_queues

  • pan/va: Only widen constants on 32-bit instructions

  • panvk: Advertise VK_KHR_workgroup_memory_explicit_layout

  • etnaviv: blt: Zero-initialize conv_swizzle

  • mr-label-maker: Add rule for rust files

  • nir: Fix lower_fround_even to round half to even

  • panvk: Add panvk_address_binding_report() helper

  • panvk: Report address binding for device memory

  • panvk: Report address binding for buffers

  • panvk: Report address binding for images

  • panvk: Report address binding for internal allocations

  • panvk: Report address binding for descriptor sets

  • panvk: Advertise VK_EXT_device_address_binding_report

  • u_transfer_helper: Add U_TRANSFER_HELPER_Z32F_S8_IN_Z24S8

  • u_transfer_helper: Convert Z32_FLOAT in the Z32F_S8_IN_Z24S8 path

  • etnaviv: Route resources through u_transfer_helper

  • gallium: Add native_fp32_depth cap to gate ARB_depth_buffer_float

  • etnaviv: Emulate Z32_FLOAT (DEPTH_COMPONENT32F) as D24S8

  • etnaviv: Lower depth32f shadow compare in the shader

  • etnaviv: Emulate Z32_FLOAT_S8X24_UINT (DEPTH32F_STENCIL8) as D24S8

  • etnaviv: blt: Address emulated depth32f as D24S8

  • etnaviv: Size depth32f transfer staging by the internal format

  • etnaviv: Support stencil blit of depth32f_stencil8

  • etnaviv: rs: Resolve depth surfaces

  • etnaviv: blt: Resolve depth_component24

  • etnaviv: Address emulated depth32f as D24S8 in resource copies

  • etnaviv: Decide transfer tileability on the physical format

  • etnaviv: Support MSAA resolve of emulated depth32f

  • etnaviv/isa: Add bit_insert instruction

  • nir: Add bitfield_insert_etna opcode

  • etnaviv: Support native bitfield_insert

  • etnaviv: hwdb: Add UNIFIED_SAMPLERS feature

  • etnaviv: Detect unified sampler support

  • etnaviv: Extract sampler descriptor emit helpers

  • etnaviv: Rework descriptor emit into a single loop

  • etnaviv: Implement unified sampler allocation

  • etnaviv: Allow depth only or stencil only MSAA resolves

  • etnaviv: Allow MSAA resolve of stencil only buffers

  • etnaviv: Fix sampler view leak on unsupported texture target

  • etnaviv: Extract texture descriptor fill helper

  • etnaviv: Keep texture descriptor template CPU-side

  • etnaviv: Compose texture descriptors at emit time

  • etnaviv: Rename etna_sampler_view_update_descriptor()

  • etnaviv: Add perf debug for texture descriptor recompose

  • util: Add u_shader_variant_cache

  • util/u_shader_variant_cache: Add hash + equal callback pair

  • etnaviv: Adopt u_shader_variant_cache

  • draw/llvm: Adopt u_shader_variant_cache

  • llvmpipe: Move FS variant LLVM types to temporary JIT struct

  • util/u_shader_variant_cache: Add cap + refcount + per-list eviction

  • llvmpipe: Migrate FS variants to u_shader_variant_cache

  • llvmpipe: Move setup variant LLVM function pointer to a local

  • llvmpipe: Migrate setup variants to u_shader_variant_cache

  • llvmpipe: Move CS variant LLVM types to temporary JIT struct

  • llvmpipe: Migrate CS/task/mesh variants to u_shader_variant_cache

  • llvmpipe: Remove dead USE_GLOBAL_LLVM_CONTEXT path

  • llvmpipe: Move shader caches and LLVMContext to screen

  • draw: Decouple GS/TCS/TES shader CSOs from draw_context

  • draw: Decouple VS shader CSO from draw_context

  • draw: Move current_variant slot from CSO to draw_context

  • draw/llvm: Bound per-stage variant caches with cap + pin slots

  • draw: Move per-draw state off the GS/TCS/TES shader CSOs

  • draw: Decouple mesh shader CSO from draw_context

  • llvmpipe: Enable PIPE_CAP_SHAREABLE_SHADERS

  • panvk: Restore push descriptor dirty bit after meta operations

  • panfrost: Add pan_nir_lower_image_64bit NIR pass

  • panvk: Alias R64 storage images to R32G32_UINT for color clear

  • panvk: Bounds-check SSBO accesses for robust storage buffer access

  • pan/format: Add R64_UINT/R64_SINT format entries for v9+

  • nir/lower_robust_access: drop out-of-bounds buffer stores

  • panvk: Wire up VK_EXT_shader_image_atomic_int64 on v9+

  • etnaviv: Only emit index buffer state when it changed

  • etnaviv: Skip state emission when nothing is dirty

  • mesa/st: Invalidate FS sampler views after PBO texture transfers

  • nir/opt_vectorize: Remove the phis that were combined

  • etnaviv: Drop stale stream output relocs in etna_update_hwxfb(..)

Christian Meissl (1):

  • nir/lower_tex: skip external texture YUV lowering for query instructions

Christoph Neuhauser (1):

  • anv: Add compute only divergent atomics fusion optimization for Blender Blender uses atomic operations as part of its virtual shadow mapping implementation. Virtual shadow mapping page tagging in compute shaders benefits from divergent atomics fusion, while fragment shaders doing the atomic raster step in general have worse performance with this optimization turned on. Thus, an option is added to only apply divergent atomics fusion to compute shaders in ANV, and this option is enabled for Blender.

Christoph Pillmayer (19):

  • pan/bi: Fix source swizzle in bi_repair_ssa

  • pan/bi: Fix format in bi_repair_ssa

  • pan/kmod: Fix uninitialized timestamp info

  • pan: Set nir_shader::source_blake3 for internal shaders

  • pan/clc: Set source_blake3 for each precompiled variant

  • pan: Add BIFROST_MESA_DUMP_DIR option to dump shader binaries

  • pan/kmod: Add l2 features to pan_kmod_dev_props

  • pan/props: Add BUS_WIDTH query to pan_props

  • pan/perf: Generate derived counter definition from Arm’s XMLs

  • pan/kmod: Add perf counter api

  • pan/kmod: Add a perf implementation to the panfrost backend

  • pan/pps: Delegate more tasks to PanfrostPerf

  • pan/perf: Make counter sum across block instances optional

  • pan/pps: Output counters per block

  • pan/perf: Add timing related getters/setters and use them

  • pan/perf: Use new kmod api

  • pan: Fix BIFROST_MESA_DUMP_DIR

  • pan/kraid: Fix 32bit hw_runner builds

  • pan/perf: Fix 32bit build of panquick

Collabora’s Gfx CI Team (21):

  • Uprev VVL to 8474616c3095756c52c1b810b21bd1366b3fc909

  • Uprev VVL to 4acd00c7a0665c9b1d01604e5fe1454837f87134

  • Uprev VVL to 6a6182c0edb35cba7bab0abc61eaff82d11022fb

  • Uprev VVL to d55be6264a17cd28f436805973b12f12a5d22f2f

  • Uprev ANGLE to 7772c5602d59140204494967ba8ebdf801180054

  • Uprev Piglit to 6fd29fe44f8857b876a67bee962919635f22ecc8

  • Uprev VVL to 36187ee9f2074609d3ec56fa1a315c191366b688

  • Uprev ANGLE to a793c75398c746f3f8a08fd2e74dfc4dff07a0c9

  • Uprev VVL to 315d28985ebd1ff9a2e4380e34ed2d8ebe487531

  • Uprev ANGLE to 196d1b79eadbd8fdbf0590092266bb2d87264988

  • Uprev VVL to d2b091858d12802cc1c3722c81a9f3d865d833d4

  • Uprev VVL to 2ab77a01659e3e46d6ca8a25425b19b3425adb11

  • Uprev ANGLE to 8e09325ebad45c7e11630a79754361e965e5fab0

  • Uprev VVL to e17d63f8fcd967b2ff91efcb8607d2c9ab962e23

  • Uprev ANGLE to 836636df1b06034b39a4a2a68b811ecf6f9b674f

  • Uprev VVL to 5e72d40395d07930c935d0624cd7db5f1a144c5e

  • Uprev ANGLE to a4eea1fbedace7a03bae52cd1bb9d6ebdfffb0f7

  • Uprev VVL to 6906fd94f2f422beb682d43d8b1af872aaff97b0

  • Uprev Piglit to 9a0eab5e1f7f009b4f72c25d23faf937b38354a6

  • Uprev VVL to e875181c4f0fab1bd0c4926c21a21de0bd967bb3

  • Uprev ANGLE to 7a960f346c57daae4b1615ead302b4c58af12258

Connor Abbott (35):

  • ir3: Don’t reset immediate count to 0 after lowering

  • ir3: Use correct immediate size for constlen calculation

  • tu: Optimize sync2 event handling in the non-asymmetric case

  • tu: Don’t zero-initialize query pool

  • turnip, ir3: Use shader for vertex input count

  • tu: Support VK_KHR_maintenance9

  • tu: Fix LRZ+FDM offset+secondaries

  • tu: Disable LRZ when resuming if the GPU doesn’t support tracking

  • tu: Zero out unused parts of descriptors

  • compiler/shader_info: Introduce occupancy_bounded_workgroup_fairness

  • spirv: Remove redundant OpExtInst handling

  • spirv: Use correct opcode in non-semantic OpExtInst handling

  • spirv: Implement concurrent workgroup hint from vkd3d-proton

  • freedreno: Document SP_CS_CNTL_0::COMPUTERRMODEEN

  • ir3, tu, freedreno: Plumb through round-robin mode

  • tu: Add TU_DEBUG=computeroundrobin

  • freedreno: Add round_robin_errata to device info

  • ir3: Implement round-robin workaround

  • nir/lower_amul: fix infinity recursion with different phi source order

  • freedreno/qrisc: Refactor section disassembly

  • freedreno/qrisc: Emulate multiple firmwares more accurately

  • freedreno/qrisc: ISA changes for gen8

  • freedreno/qrisc: Support DDE on gen8

  • freedreno/qrisc: Don’t write CP_LPAC_SQE_CNTL on a7xx+

  • freedreno/qrisc: Add support for gen8 firmwares

  • freedreno/qrisc: Add missing <decode/> to #sqe-base

  • freedreno/qrisc: Add extra preempt entrypoint

  • freedreno: Add some gen8 control registers

  • freedreno/qrisc: Allow limited relocation expressions

  • freedreno/qrisc: Allow labels on two-src ALU instructions

  • freedreno/qrisc: Bump label/instruction count

  • freedreno/qrisc: Add support for “absolute” label references

  • freedreno/qrisc: Test new label features

  • tu: Fix resetting command streams with writeable BOs

  • tu: Fix condition for skipping emitting aprons

Daivik Bhatia (6):

  • broadcom/compiler: Add explicit NOP instruction at page boundaries

  • pan/nir: fix GNU compilation error with clang

  • broadcom/compiler: add support for null descriptors

  • v3dv: Implement and enable nullDescriptor support

  • nir/opt_copy_prop_vars: kill stale entries when source deref is written

  • rocket: simplify input/output tensor creation

Daniel Lang (3):

  • radv/meta: fix samples datatype in radv_meta_nir

  • nir: change type_size return type to unsigned in nir_lower_{amul,io}

  • docs: Add GL_ARB_map_buffer_range and GL_ARB_vertex_array_object to etnaviv

Daniel Schürmann (39):

  • nir: add nir_loop::do_while to indicate do-while loops

  • vtn: set nir_loop::do_while during spirv_to_nir()

  • glsl_to_nir: set nir_loop::do_while

  • nir/opt_loop: Don’t peel initial break from do-while loops

  • nir/opt_loop: stop recursion at loop header phi in can_constant_fold()

  • nir/opt_loop: always try to peel initial break from loops with unrolling hint

  • nir/opt_algebraic: use imul24_relaxed for lowered dot4x8_add

  • nir/opt_algebraic: add some imul24_relaxed pattern

  • nir/opt_constant_folding: create const_value_for_alu() helper

  • nir/opt_constant_folding: constant-fold op(bcsel(), #c) -> bcsel(.., #c1, #c2)

  • nir/opt_algebraic: optimize downcast followed by upcast to extract

  • nir/opt_algebraic: extend some extract_u8 pattern to extract_i8

  • nir/builder: constant-fold nir_mov_alu() if requested

  • nir/lower_bit_size: use nir_builder::constant_fold_alu

  • nir/lower_bit_size: use nir_def_replace() instead of nir_def_rewrite_uses()

  • nir/lower_bit_size: skip conversion for more opcodes

  • anti-lag: rework wait time calculation

  • aco/assembler: Fix s_inst_prefetch insertion after loop latch rotation

  • aco/assembler: pass std::vector to insert_code

  • nir: remove fixed-sized nir_op_bitz / nir_op_bitnz

  • nir: disallow converting to sized booleans from nir_type_convert()

  • nouveau: don’t handle 8- and 16-bit comparisons

  • panfrost: remove fixed-sized binop reductions

  • panfrost: replace fixed-sized with unsized comparison opcodes

  • panfrost: replace fixed-sized bcsel with unsized bcsel_pan

  • nir,panfrost: remove 8-bit and 16-bit booleans

  • aco/ra: Fix get_reg_impl() for operand registers

  • aco: encode unused VOP3 operands as inline constant 0 on RDNA

  • aco/isel: move add64_32() to aco_isel_helpers.cpp

  • aco/isel: use add64_32() for nir_iadd(nir_u2u64(), ..)

  • nir/range_analysis: handle read_first_invocation and friends in nir_def_num_lsb_zero()

  • nir/range_analysis: handle phis in nir_def_num_lsb_zero()

  • amd/lower_global_access: refactor using state struct

  • amd/lower_global_access: only consider constant offsets that are aligned

  • amd/lower_global_access: Only consider 32-bit offsets which are aligned

  • amd/lower_global_access: lower SMEM offsets according to hw capabilities

  • aco: remove alignment handling for global SMEM loads

  • aco: emit global SMEM loads directly

  • aco: Remove SMEM offset optimization for non-buffer loads

Daniel Stone (9):

  • pan/afbc: Code motion for split modifier queries

  • pan/mod: Protect against no usage flags for 64k

  • pan/mod: Reorder linear modifier checks

  • pan/afbc: Properly validate format/parameter combinations

  • ci/panfrost: Switch T860 jobs to another RK3399 device type

  • ci/panfrost: Add two T860 OpenCL fails

  • symbols-check: Ignore more pthread symbols

  • doc/ci: Add custom-kernel testing workflow

  • draw: Avoid warnings for maybe-unused variable

Danylo Piliaiev (54):

  • tu: Fix draw call offset for LRZ warnings in secondaries

  • tu/perfetto: Move away from single timeline for all apps

  • tu: Fix CP_CCHE_INVALIDATE not being applied at the right point

  • freedreno: Fix CP_CCHE_INVALIDATE not being applied at the right point

  • tu/u_trace: Use correct u_trace destination in tu_clone_trace_range

  • tu/u_trace: Prevent cloning stale RB_DONE_TS results

  • tu/u_trace: Correct the order of tracepoints clonning for binning

  • tu/u_trace: Fix explicit toggle_name not being used

  • tu/perfetto: Add a performance warning track to perfetto

  • tu/perfetto: Add performance warning tracepoints

  • tu: Fix tu_bo_make_zombie without queues

  • tu: Fix double free of timestamp_copy_data->trace

  • tu: Don’t leak pre_chain.rp_trace, and correct u_trace_move

  • tu: Fix BV/BR race in tu_clone_trace_range when waiting on barrier

  • tu/a8xx: Fix reading border_color from sampler memory

  • tu: Don’t disable UBWC for D24S8+USAGE_SAMPLED+customBorderColorWithoutFormat

  • tu: Disable concurrent binning by default due to perf regressions

  • tu: Don’t enable FDM when there is FDM attachment is UNUSED

  • tu: Always lazy_init_vsc for tiler rendering

  • u_trace: Lazy init ut->linear_alloc

  • tu: Fix TU_CMD_DIRTY_DRAW_STATE value collision

  • tu: Start/End occlusion query should force depth state recalculation

  • tu: Change of disable_fs state should force depth state recalculation

  • tu/a7xx: Don’t force enable IJ_LINEAR_PIXEL for FragFace/FragCoord

  • freedreno/a7xx: Don’t force enable IJ_LINEAR_PIXEL for FragFace/FragCoord

  • tu: Disable FS in some cases even when FS explicitly writes D/S

  • ir3: Add resbase_ir3 intrinsic

  • tu: Add allow_oob_indirect_ubo_loads to device cache uuid

  • tu: Specify max texel buffer and storage buffer limits via GPU props

  • tu/a8xx: Set real storage/texel buffer size limits

  • tu: Add option to raise the maximum texel buffer size

  • tu: Enable texel buffer / SSBO emulation for known problematic games

  • tu: Match SW depth clear value packing with HW

  • tu: Match SW color clear value packing with HW

  • tu: Don’t process A2R10G10B10 clear values via new pack function

  • tu/lrz: Pick correct depth attachments in msrtss case

  • tu: Force GMEM mode when renderpass has MSRTSS attachments

  • tu: Refactor separate D32S8 and ignore D/S aspect mask for RP attachments

  • tu: Remove depth/stencil-specific blit src/dst helpers

  • tu: Remove event_blit_dst_view

  • tu: Enable tu_dont_care_as_load for all Kex Engine games

  • tu: Use application_name_match instead of exe match for workarounds

  • tu/a6xx: Work around D32S8 EARLY_Z_LATE_Z hang

  • tu: Enable tu_allow_oob_indirect_ubo_loads for Clausewitz engine

  • tu: Custom resolve should always use AVOID_CCU layout

  • tu: Fix subsampled metadata and blit emission for separate stencil

  • tu: Fix gfx_write_access checking for TRANSFORM_FEEDBACK_COUNTER_READ_BIT

  • tu: Fix blit_cache_cleaned never being set to true

  • tu: Fix LRZ handling for VK_EXT_custom_resolve

  • tu: Dirty LRZ after changing attachment locations disable LRZ writes

  • nir: Include scalarized component offsets in UBO ranges

  • tu: Fix tu_event not being reset on creation

  • tu: Merge disable_write_for_rp from secondary to primary

  • tu: Fix memory leak of FDM patch-points

Dave Airlie (15):

  • nouveau: drop sector promotion.

  • gallivm: handle llvm 22 coroutine end change

  • gallivm: handle llvm 22 scatter/gather intrinsic changes.

  • lavapipe: treat NULL pColorAttachmentLocations as no handles

  • nak: fix image size for multisample arrays

  • nak: add more sizes to assert in bindless_image_sparse_load

  • nvk: enable subgroupQuadOperationsInAllStages

  • ci: vmware farm is offline, stop using it

  • st: fix get tex subimage fallback for 1D ARRAY

  • st: drop ununsed arguments to copy_to_staging_dest.

  • u_blitter: only set texcoord.w to sample for multisample sources

  • nak: block pipe_format from nak bindings.

  • ir3: use the correct builder for adding preamble to main.

  • nir: add impl pointer to block to avoid recursive linked list

  • nir: remove a lot of nir_cf_node_get_function calls.

David Airlie (3):

  • nir/coopmat: refactor the split vars to clean it up

  • nir/coopmat: move the row/col into a box and add some helpers.

  • nir/coopmat: rename the box split variables.

David Rosca (70):

  • d3d12: Use HEVC RefPicSet order from frontend

  • ac/parse_ib: Fix printing enc recon VAs on VCN5

  • radv: Fix uint32 overflow in slice offset calculation

  • radv/video: Fix initializing rc structs with default rate control

  • radeonsi: Always use 2D tiling for video dpb

  • frontends/va: Fix finding LTRs from POCs in HEVC decode

  • frontends/va: Fix out of bounds write in AV1 decode tile info

  • frontends/va: Fix setting output color properties from color standard

  • frontends/va: Fix dereference before NULL check in postproc

  • frontends/va: Add missing NULL check for additional output surface

  • vl: Use NV12 as deint format instead of preferred format

  • vl: Don’t check npot textures support when creating buffers

  • pipe/video: Remove unused PIPE_VIDEO_CAP_PREFERRED_FORMAT

  • pipe/video: Remove unused PIPE_VIDEO_CAP_NPOT_TEXTURES

  • pipe/video: Remove unused PIPE_VIDEO_CAP_MAX_LEVEL

  • pipe/video: Remove unused PIPE_VIDEO_CAP_STACKED_FRAMES

  • radeonsi/uvd_enc: Skip extra padding bytes in output bitstream

  • radeonsi: Move si_vpe.* to mm subfolder

  • ac/info: Add video codec caps

  • radeonsi/video: Use new video codec caps

  • radv/video: Use new video codec caps

  • ac/info: Remove old video codec caps

  • ac/info: Print number of VPE instances

  • radeonsi: Add RADEON_FLUSH_FORCE and use it to force flush

  • ac/vcn_dec: Add ac_vcn_dec_init_regs to get register offsets

  • ac/vcn: Add ac_vcn_sq_header/tail and use it for decode

  • ac/cmdbuf: Add ac_emit_video_write_memory

  • radv: Use ac_emit_video_write_memory

  • ac/vcn_dec: Move register defines to ac_vcn_dec.c

  • ac/cmdbuf: Add ac_emit_video_write_timestamp

  • radv: Add support for timestamps on video queue

  • radeonsi/mm: Add support for 2-ref H264 encode

  • radeonsi/mm: Remove comment about kernel AV1 instance scheduling bug

  • ac/parse_ib: Add VCN decode queue parsing

  • ac/parse_ib: Add VCN timestamp command

  • radeonsi/mm: Add si_vid_create_buffer and use it

  • radeonsi/mm: Set PIPE_RESOURCE_FLAG_UNMAPPABLE for buffers

  • va: Set contiguous_planes for DMA-BUF imported surfaces

  • ac/vcn_dec: Add 10 to 8 bit dithering support

  • radeonsi/mm: Select DPB format independently from decode surface format

  • vl: Skip transfer function and primaries conversion when not needed

  • va: Always reset compositor chroma location

  • va: Use RGB format with matching bit depth for YUV->YUV matrices

  • radeonsi/mm: Return error when decoding H264 P/B frame with no refs

  • radeonsi/mm: Only setup ref surfaces with tier3

  • radeonsi/mm: Set correct usage in si_dec_fill_surface

  • radeonsi/mm: Fix setting VPE rotation when horizontal flip is enabled

  • va: Implement vaPutImage for derived images

  • pipe/video: Add out_pipe_fence to pipe_picture_desc

  • vl: Support blending with gfx compositor

  • vl: Add pipe_video_codec proc using vl_compositor

  • vl: Add vl_proc pipe_video_codec using vl_compositor as fallback

  • va: Use vl_proc for processing context

  • va: Add vlVaDestroySurface

  • va: Add vlVaPostProc and use it instead of compositor and vid engine blit

  • va: Use vlVaPostProc in vlVaPutSurface and for subpictures

  • va: Stop using vl_compositor

  • pipe: Remove pipe_video_codec::expect_chunked_decode

  • r600/uvd: Set correct h264 chroma format

  • pipe: Remove pipe_video_codec::chroma_format

  • vulkan/video: Fix coding AV1 decoder/encoder_buffer_delay

  • vulkan/video: Don’t code AV1 decoder model info when not present

  • vulkan/video: Fix coding AV1 operating points

  • d3d12/video: Don’t reset batches in fence_wait

  • vulkan/video: Fix coding H265 ref pic list modification lists

  • vulkan/video: Fix coding H265 SPS pcm block sizes, inter ref pic set and lt refs

  • va: Fix leak when vlVaUploadImage fails

  • va: Ensure templat is valid for temporary surfaces

  • radeonsi: Stop forcing GTT with no gfx/compute

  • ac/video: Fix number of AV1 single refs

Derek Lesho (2):

  • zink: Guard bo map/unmap on map_count.

  • zink: Fix zink_bo_unmap synchronization for client pointer support.

Dhruv Mark Collins (7):

  • tu/autotune: Fail gracefully when CP counters are unavailable

  • fd/pps: Allocate performance counters from high-to-low

  • tu/autotune: Allocate performance counters from low-to-high

  • tu/query_pool: Avoid CP counter conflict with autotune

  • freedreno: Update A6XX_PC_MODE_CNTL definition and values

  • tu/util: Fix tile division algorithm

  • tu: Propagate allocation failures for tu_cs_* functions

Dmitry Baryshkov (3):

  • rusticl: enable freedreno by default

  • tu: limit KHR_internally_synchronized_queues to Vulkan 1.1+

  • tu: limit VALVE_fragment_density_map_layered to Vulkan 1.1 devices

Dmitry Osipenko (3):

  • intel/virtio: Preserve errno properly when handling ioctl

  • drm-uapi: Update virtio-gpu with new hinting field

  • intel/virtio: Support DRM_VIRTGPU_BLOB_FLAG_HINT_DEFER_MAPPING

Dorinda Bassey (1):

  • util/rust: Add atomic memory synchronization support

Duncan Brawley (7):

  • pco: Fix pco_last_igrp returning the first element instead of the last

  • pco: Refactor internal shader pass skipping

  • pco: Add propagating coherent/volatile access qualifiers

  • pco: Add DMA ld/st caching support for ssbo/ubo operations

  • pco: Add DMA sampling caching support

  • pco: Fix smp instruction encoding map order

  • pco: Add DMA ld/st caching support for all ld/st instructions

Dylan Baker (3):

  • intel/brw: Add assert for error case

  • meson: ensure that libdrm auto-features match requirements

  • intel/gen: decode type of src1 in basic 2 source after setting IMM

Emma Anholt (87):

  • spirv: Demote the SPIRV 1.6 OpTypeSampledImage on Buffer failure to a warning.

  • ci: Bump apitrace version to 14.0.

  • ci: Don’t set wine vars in deqp-runner.sh/vkd3d-runner.sh.

  • ci: Build a working wine installation in build-wine.sh.

  • ci/test-vk: Install win64 apitrace 14.0 along with setting up wine.

  • ci/test-vk: Install DXVK 2.7.1 to our wine installation.

  • ci/lava: Fix the name of the fluster overlay.

  • ci/lava: Add a note about an otherwise-mysterious error you can encounter.

  • ci/gfxreconstruct: Disable OpenXR support.

  • ci: Build gpu-trace-perf and include a script to use it.

  • ci: Bump the image tags for the previous build script changes.

  • bin/update-traces_checksum.py: Pull out per-job work to a helper function.

  • ci/update_traces_checksum: Parse gpu-trace-perf’s format for hash changes.

  • ci/update_traces_checksum: Default to updating for the current HEAD.

  • ci/update_traces_checksum: Make it work on restricted traces jobs, too.

  • ci/llvmpipe: Use anholt’s new GPU trace snapshot comparison tool.

  • ci/lavapipe: Use anholt’s new GPU trace snapshot comparison tool.

  • ci/turnip: Drop two 660 vk jobs and tune down the vk coverage fraction.

  • ci/turnip: add an a660 VK restricted traces job.

  • ci: Delete references to various broken traces.

  • ci/amd: Switch radv-raven-traces-restricted over to gpu-trace-replay.sh

  • ci/intel: Switch over to the new tool for restricted traces.

  • ci/piglit-traces: Remove ANGLE trace support.

  • ci/llvmpipe: Disable some traces too close to the timeout.

  • tu: Set HALF_PRECISION on blits to R11G11B10.

  • ir3: Fix shared IMAD24 lowering.

  • tu: Add capture/replay for sparse buffers and descriptor buffer.

  • screenshot-layer: Fix leftover VK queues in the map at DeviceDestroy.

  • screenshot-layer: Fix a bunch of unused variable warnings.

  • screenshot-layer: Fix rename() to final png before the file is flushed.

  • screenshot-layer: Clean up the lifetime management of the copyDone fence.

  • screenshot-layer: Fix race on writing the .pngs vs device destroy.

  • screenshot-layer: Wait on the fence before fallible operations.

  • tu/ci: Drop some old xfails that don’t trigger any more.

  • tu: Report missing layout support for host_image_copy with unifiedLayouts.

  • zink/ci/tu: Fix up skips/xfails for GLCTS testcases that got divided up.

  • tu: Disable storage image support for depth/stencil.

  • ci/panfrost: Drop a set of flakes whose fix had landed.

  • panfrost/ci: Skip dEQP-VK.wsi.wayland.swapchain.render.10swapchains on g52.

  • lvp/ci: Drop an old skip long since fixed in the CTS.

  • tu/ci: Drop a750 VKCTS to 50% coverage.

  • ci: Update VK CTS to 1.4.5.3 with fixes.

  • ir3: Add an env var to prefer single wavesize.

  • ir3: Fix shader bisect crashing out when too many shaders get bisected.

  • ir3: Give some feedback as we shader bisect.

  • ir3/shader_bisect: Allow a ‘r’ response to retry a run mid-bisect.

  • ir3: Drop the “SIMD0” debug print that was apparently added for frameretrace.

  • ir3: Deduplicate shader disassembly generation.

  • ir3: If we’re dumping IR3_SHADER_BISECT=[hash] disasm, include the NIR.

  • tu: Disable 128-wide subgroups on No Man’s Sky.

  • screenshot-layer: Log when we can’t open the output directory.

  • weston: Run at a more reasonable 1920x1080 resolution, not 1024x640.

  • ci: Include Windows renderdoc in with wine and update gpu-trace-perf.

  • ci: Build the Vulkan screenshot layer as part of VK test builds.

  • freedreno/ci: Add restricted traces testing of D3D11 traces on a660.

  • freedreno/ci: Add more explanation of a trace failure that’s not our fault.

  • tu: Always set the kernel’s name for BOs.

  • util/drirc_gen: Add a little documentation of what this does.

  • radv/drirc_gen: Clean up the dependency handling.

  • util/drirc_gen: Reduce manual importing of functions.

  • util/drirc_gen: Move the common VK WSI options to a core helper function.

  • util/drirc_gen: Make the header usable from C++.

  • tu: Move to using drirc_gen.

  • drm-shim: Include the hex of the driver ioctl for unimplemented ioctls.

  • drm-shim/freedreno: Provide a dummy set of UBWC config params.

  • drm-shim/freedreno: report a 48-bit address space.

  • drm-shim/freedreno: Report VM_BIND support.

  • vulkan: Enable GOOGLE_display_timing on KHR_display across multiple drivers.

  • drm-shim/freedreno: Fix VM_BIND support.

  • docs: Link in particular to the difficulty: * issue tags in Help Wanted.

  • zink: Use the new common code for nearest consistency in blits.

  • zink: Also enable the nearest consistency workaround on turnip.

  • zink: Also enable the nearest consistency workaround on anv.

  • .mailmap: Switch to anholt’s current work address.

  • intel/device_info_override_test: Make sure we actually find our device.

  • drm-shim: Share common code for PCI and platform device setup.

  • drm-shim: Generalize overriding of links.

  • drm-shim: Give the device/subsystem links real link values.

  • drm-shim: Lock access to shim_device.fd_map.

  • drm-shim: Remove drm_shim_driver_prefers_first_render_node.

  • drm-shim: Remove unnecessary runtime setup of drm_device_path_prefix.

  • drm-shim: Remove unnecessary runtime setup of various device strings.

  • drm-shim: Fix racy initialization.

  • freedreno/ci: Clear the xfail for texture-immutable-levels.

  • etnaviv/ci: Fix flakes lists that are breaking gc2000 CI.

  • intel/ci: Fix xfails for nightlies.

  • freedreno: Don’t force image component A=1 substitution on R/RG textures.

Emre Cecanpunar (1):

  • jay: allocate shader under memctx

Eric Engestrom (131):

  • VERSION: bump to 26.2

  • docs: reset new_features.txt

  • docs: update calendar for 26.1.0-rc1

  • docs: update calendar for 26.0.5

  • docs: add release notes for 26.0.5

  • docs: add sha sum for 26.0.5

  • docs: add stub of vk_struct_type_cast.h for vk_util.h

  • ci/bare-metal: drop duplicate timestamps now that gitlab-runner has per-line timestamps

  • docs: update calendar for 26.1.0-rc2

  • docs: update calendar for 26.1.0-rc3

  • docs: update calendar for 26.0.6

  • docs: add release notes for 26.0.6

  • docs: add sha sum for 26.0.6

  • docs: update calendar for 26.1.0

  • docs: add release notes for 26.1.0

  • docs: add sha sum for 26.1.0

  • docs: add calendar for the 26.1 cycle, and 26.2 branchpoint and release candidates

  • docs: fix unescaped `*`

  • docs/submittingpatches: fix section nesting

  • docs/ci: explain what Marge saying “Manual Step encountered” means

  • zink+nvk/ci: update expected fails

  • docs: update calendar for 26.0.7

  • docs: add release notes for 26.0.7

  • docs: add sha sum for 26.0.7

  • docs/ci: ignore docs.redhat.com & registry.khronos.org links

  • etnaviv: initialize value before calling etna_gpu_get_param(), in case it fails

  • meson/libmesa: ensure shader_replacement.h is generated before using it

  • meson/amd: only build libaco when requested

  • meson/asahi: only build libagx2_disasm when requested

  • meson/freedreno: only build libfreedreno_common when requested

  • meson/intel: only build libblorp_elk when requested

  • ci/build: restore riscv64 build as it works again

  • Revert “ci/build: restore riscv64 build as it works again”

  • docs: update calendar for 26.1.1

  • docs: add release notes for 26.1.1

  • docs: add sha sum for 26.1.1

  • docs: update calendar for 26.0.8

  • docs: add release notes for 26.0.8

  • docs: add sha sum for 26.0.8

  • util/meson: simplify list of per-driver drirc files

  • drirc: move 00-$drv-defaults.conf to each driver’s folder

  • Revert “drirc: move 00-$drv-defaults.conf to each driver’s folder”

  • docs: update calendar for 26.1.2

  • docs: add release notes for 26.1.2

  • docs: add sha sum for 26.1.2

  • rusticl: skip bindgen for pipe_shader_state_from_tgsi

  • meson: exclude known buggy versions of bindgen

  • ci: bump rust version from 1.90 to 1.96

  • ci: bump bindgen version from 0.71.1 to 0.72.1

  • ci: bump fedora from 42 to 44

  • meson: drop non-existent platforms=xcb check

  • Revert “egl: fix _EGL_NATIVE_PLATFORM fallback for unrecognized native displays”

  • docs: update calendar for 26.1.3

  • docs: add release notes for 26.1.3

  • docs: add sha sum for 26.1.3

  • gen_release_notes_test: don’t evaluate backslash

  • gen_release_notes: add support for “work_items” links

  • docs: fix release notes for 26.1.0

  • docs: fix release notes for 26.1.1

  • docs: fix release notes for 26.1.2

  • docs: fix release notes for 26.1.3

  • ci: fix perfetto download in `make-git-archive` nightly job

  • ci: fix perfetto download in build-perfetto.sh

  • ci: fix the fix for perfetto download in `make-git-archive` nightly job

  • etnaviv/ci: document two fixed tests

  • nvk/ci: document fixed tests, new failures, and recent flakes

  • zink+nvk/ci: fix duplicate fails

  • docs: drop x.org -> x.org/wiki/ redirect and expected url

  • docs: s/issues/work_items/

  • docs/ci: mark yet another domain as blocking linkcheck

  • docs/ci: disable auto-retry on nightly linkcheck

  • docs/ci: use full/explicit option names in linkcheck job

  • docs/ci: only print the linkcheck issues, not the thousands of non-issues

  • mr-label-maker: add ~drirc label on all drirc files

  • util: add support for multiple colon-separated DRIRC_CONFIGDIR entries

  • drirc: move 00-$drv-defaults.conf to each driver’s folder

  • meson: merge two consecutive `if with_egl`

  • meson: add native platform to the summary

  • meson: ensure native platform is one of the undetectable ones

  • meson: drop misleading `-D egl-native-platform` values

  • zink/ci: drop leftover anv-cml deqp suite

  • docs: update calendar for 26.1.4

  • docs: add release notes for 26.1.4

  • docs: add sha sum for 26.1.4

  • docs: fix x.org url

  • etnaviv/ci: update nightly job expectations

  • zink+nvk/ci: update nightly job expectations

  • lvp/ci: update nightly job expectations

  • llvmpipe/ci: update nightly job expectations

  • nvk/ci: update nightly job expectations

  • drm-shim: name the driver name `driver_name` consistently

  • drm-shim: set `driver_name` in `drm_shim_*_device_setup()`

  • docs: gitignore the contents of the `_generated` folder

  • ci: disable auto-retry on rustfmt job

  • docs/helpwanted: url-encode `[]` to avoid a pointless redirection

  • rusticl: document api@clgetmemobjectinfo as fixed for all drivers

  • lavapipe/ci: document fixed dEQP-VK.mesh_shader.ext.misc.emit_in_control_flow_bad_emit_last

  • freedreno/ci: document fixed KHR-GL46.copy_image.smoke_test

  • freedreno/ci: document a recent flake

  • zink+nvk/ci: document a couple of recent flakes

  • ci/piglit: fix nightly expectations after piglit uprev

  • img/ci: add `farm:imagination` tag to all jobs

  • loader: move variable to correct scope

  • broadcom/ci: mark fixed tests as such

  • ci/video: move two single-thread tests to global list

  • ci/video: install the current version of gstreamer

  • ci/video: uprev fluster

  • ci/video: download fluster test suites by codec name

  • ci/video: update comment with the new blocker for AV1 support

  • radv/ci: enable VP9 testing in fluster

  • anv/ci: enable VP9 testing in fluster

  • nvk/ci: document fixed dEQP-VK test

  • nvk/ci: document fixed vkd3d tests

  • nvk/ci: document two vkd3d regressions

  • freedreno/ci: document fixed tests

  • radeonsi: fix truncated cache key

  • mailmap: update my email address

  • zink+nvk/ci: document two fixed tests

  • VERSION: bump for 26.2.0-rc1

  • .pick_status.json: Update to d49a00bdf15fd48b31af93aaf5feed3eebcbda12

  • VERSION: bump for 26.2.0-rc2

  • .pick_status.json: Update to 8b00adbe72f2705985146b057f6fde9256d0dcb0

  • .pick_status.json: Mark 47efd739121d51e2f9049cec715a70c13767a67c as denominated

  • .pick_status.json: Mark 7999060992e9cee91d1962faf65dc4e5c6fe4f69 as denominated

  • .pick_status.json: Mark 2515024a5919ed14fe05471e3f1f89c54a454610 as denominated

  • .pick_status.json: Mark 4fd93a0039a07ec2027f2a6d1d252c73ed033e0a as denominated

  • pick-ui: turn commit.date into a (cached) property

  • pick-ui: show MR number for additional context

  • VERSION: bump for 26.2.0-rc3

  • .pick_status.json: Update to 85c082ddbed727940535911e6bf87f7d274525bf

  • [26.2 only] docs/new_features: mention that VK_EXT_host_image_copy was exposed on RADV/GFX10.3+

Eric Guo (3):

  • compiler: Add missing MESA_SHADER_KERNEL case for SPIR-V dump

  • pan/compiler: Clamp fp16 ldexp exponent range

  • pan/bi: Lower 64-bit hadd on v9/v10

Eric R. Smith (4):

  • glsl, spirv: Improve accuracy of asin() and acos()

  • panfrost: add some sanity checks

  • panfrost: make sure INDEX_OFFSET is cleared

  • panfrost: add helper function for checking for active queries

Erico Nunes (4):

  • ci: lima farm maintenance

  • Revert “ci: lima farm maintenance”

  • CODEOWNERS: add lima maintainers

  • ci: lima farm maintenance

Erik Faye-Lund (88):

  • panvk: drop out-of-date TODO

  • panfrost: use perf-trilinear when doing anisotropic sampling

  • panvk: use perf-trilinear when doing anisotropic sampling

  • pan/lib: fix up afbc and linear layout

  • pan/lib: emit high bits of buffer-size

  • pan/lib: validate data_size_B in drivers

  • panvk: do not artificially limit image dimensions

  • panvk: increase maxResourceSize on v11 and later

  • panvk: increase maxBufferSize on v11 and later

  • nouveau: do not report unsupported feature

  • radeonsi: remove old, unsupported cap

  • d3d12: remove benign but unsupported cap

  • iris,crocus: remove benign but unsupported cap

  • llvmpipe: drop support for tgsi_tex_txf_lz cap

  • ntt: stop emitting TXF_LZ

  • gallium/u_blitter: stop emitting TEX_LZ

  • gallium: remove defunct pipe-cap

  • ttn: do not handle T{EX,XF}_LZ

  • gallium: completely remove T{EX,XF}_LZ opcode

  • panvk: do not enable extension without required feature

  • panvk: do not enable extension without required feature

  • haiku: remove unfinished post-processing support

  • gallium: delete leftovers of post-processing infrastructure

  • pan/ci: add a flake from nightly

  • util/format: make Y8_UNORM an alias of Y8_400_UNORM

  • util/format: make subsampling explicit

  • util/format: mark subsampled RGB formats as actually subsampled

  • util/format: verify subsampling in name

  • pan/va: do not allow force_delta_enable on v9

  • pan/bi: correct computation of lod.x

  • panfrost: enable ARB_texture_query_lod on v9+

  • mesa/main: remove stale prototypes

  • mesa/main: remove incorrect debug-output

  • mesa/main: do not gate performance warning

  • mesa/main: remove low-value debug-output

  • mesa/main: remove unused verbose-flags

  • mesa/main: remove VERBOSE_API

  • mesa/main: remove mesa_print_display_list function

  • mesa/main: remove low-value verbose-switch

  • Revert “mesa: check for ARB_ES3_compatibility in format checks”

  • mesa/main: remove unused array

  • pan/ci: update flakes based on nightly ci

  • pan/ci: remove benign typoed flake

  • meson: update libdrm wrap

  • pan/ci: add missing gitlab rules

  • pan/ci: remove outdated gitlab rule

  • pan/ci: add missing gitlab rule

  • pan/ci: fix gitlab rules after move

  • pan/genxml: correct size of field

  • pan/genxml: add missing modifier

  • pan/genxml: correct size of field

  • pan/genxml: correct size of field

  • pan/genxml: correct size of field

  • pan/genxml: add missing enum value

  • pan/genxml: sort CS structs by enum-value

  • pan/genxml: use consistent name for scissor

  • pan/genxml: use an enum for progress increment

  • pan/genxml: consistently use bool for error reject

  • pan/genxml: consistently use hex for masks

  • pan/genxml: consistently use uint for signal slot

  • pan/genxml: consistently use uint for chunk indexes

  • pan/genxml: remove needless defaults

  • pan/genxml: keep enum ordering from v10

  • pan/genxml: correct casing of names/types

  • pan/genxml: consistently use hex for uint immediates

  • pan/genxml: consistently set default

  • pan/genxml: clean up whitespace

  • pan/genxml: make field consistent

  • pan/genxml: remove some pointless comments

  • pan/genxml: use consistent attribute order

  • pan/ci: add a couple of flakes

  • pan/ci: use slow-skips to only skip slow tests for merge-requests

  • pan/ci: move cts-bug-fails to skips

  • pan/ci: stuff some breadcrumbs in the fails-list

  • pan/ci: reenable passing tests

  • pan/ci: add a few new g925 flakes

  • pan/ci: add back missing skip-list heading

  • pan/ci: drop needless skips

  • pan/ci: move skip to flakes

  • pan/ci: skip slow test

  • pan/ci: move common flake to common flake-file

  • pan/ci: recognize flaking test

  • pan/ci: mark missing xfails

  • pan/ci: move longprim flake into common flake-file

  • pan/ci: add new flake

  • pan/ci: just mark all random-max draw-tests as flakes

  • ci/vulkan: remove long outdated skips

  • panvk: simplify non_polygon calculation

Etaash Mathamsetty (4):

  • vulkan/wsi/wayland: Fix error handling for tearing control.

  • vulkan/wsi/wayland: Move drm syncobj to swapchain.

  • vulkan/wsi/wayland: Move color management surface to swapchain.

  • vulkan/wsi/wayland: Do a roundtrip after retiring the old swapchain.

Faith Ekstrand (399):

  • panvk/csf: Emit INDEX_BUFFER[_SIZE] even for non-indexed draws

  • pan/bi: Improve swizzle propagation

  • zink: Assert if we try to use a dedicated allocation with offset > 0

  • panfrost: Add and use a new pan_nir_res_handle() helper

  • pan,nir: Add cube face intrinsics

  • nir/builder: Allow backend1/2 in nir_build_tex()

  • nir: Add a new nir_texop_gradient_pan

  • panvk: Implement bitfield_select

  • pan/nir: Add a pass for lowering texture ops in NIR on Valhall+

  • pan/nir: Use the NIR lowering on Valhall+

  • nir: Add a new nir_op_f2u32_rtne

  • pan/bi: Implement nir_op_f2[iu]32_rtne

  • pan,nir: Add Bifrost texturing intrinsics

  • pan/nir: Add bifrost support to pan_nir_lower_tex()

  • pan/nir: Lower texturing ops in NIR on Bifrost

  • pan/nir: Load texel buffer conversion descriptors in NIR

  • pan/bi: Allow setting the table on lea_attr_pan

  • pan/nir: Use HW NIR intrinsics for texel buffer addresses

  • pan/bi: Delete the old texel buffer intrinsics

  • pan/nir: Lower texel buffers in nir_lower_tex()

  • pan/nir: Lower texture queries in nir_lower_tex() on Valhall+

  • panfrost: Also remap image handles for image_size/samples

  • pan/nir: Lower image queries in NIR on Valhall+

  • panvk: Let the compiler handle texture queries on v9+

  • pan/nir/tex: Support full index+offset

  • panvk: Add MAX_VS_ATTRIBS to image indices in panvk_nir_lower_descriptors

  • panfrost: Take texture/sampler_index into account in lower_res_indices

  • panfrost: Prefix valhall bits of lower_res_indices

  • panfrost: Handle pre-Valhall images and texel buffers in lower_res_indices

  • pan/bi: Drop lower_index_to_offset from preprocess

  • util/half: Use explicit RTNE rounding for the C++ float16_t

  • util/half: Stop whacking CPU flags to test float_to_half_slow()

  • util/half: Rename the tests

  • util/half: Re-organize the tests a bit

  • util/half: Add float_to_half rounding tests

  • util/half: Add double_to_half tests

  • util/half: Add a simpler double_to_float16()

  • util/half: Add double_to_float16_ru/rd helpers

  • nak: Move Srcs/DstsAsSlice implementations

  • nak: Implement Srcs/DstsAsType directly for Op

  • nak: Implement Srcs/DstsAsSlice directly on ops

  • nak,compiler: Move AttrList into NAK

  • nak: Don’t use the proc macro to implement auto-boxing of ops

  • nak,compiler: Move FromVariants to common code

  • pan/bi: Use LOD_MODE_EXPLICIT for the 2nd half of textureGrad() on Bifrost

  • docs: Move and rename “Development Notes”

  • docs: Add docs with Vulkan/SPIR-V extensions basics

  • docs: Add docs for drafting new MESA extensions

  • panvk/csf: fix VERTEX_SPD dirty tracking when topology changes

  • panvk/csf: Inline the SPD addr helpers

  • nouveau/push: Rename push_method to push_mthd

  • nouveau: Don’t build NAK tests on Android

  • compiler/rust: Add a float16 wrapper

  • etnaviv: Remove f32_to_f16_fallback() in favor of float16::F16

  • meson: Bump the minimum rust version to 1.85.0

  • compiler/rust: Add LowerBoundedU32[Array] types

  • nak: Use LowerBoundedU32 for SSAValue

  • nak: Allow SSA value 0 again

  • nak: Simplify SSARef construction with try_push()

  • meson: Suffix compiler/rust bindings with _compiler_rs_extern

  • compiler/rust/bindings: Add util_dyarray

  • compiler/rust: Add a nir_shader::get_entrypoint() helper

  • compiler/rust: Add a nir_shader::to_string()

  • compiler/rust/nir: Add structured block iterators

  • compiler/rust/nir: Add helpers for getting ALU input/output types

  • compiler/rust/bitset: Add a BitIndex helper struct

  • compiler/rust/bitset: Don’t reserve space in remove()

  • compiler/rust/bitset: Add find_next_[un]set() helpers

  • compiler/rust/bitset: Generalize BitSetIterator

  • compiler/rust/bitset: Implement Into/FromBitIndex for more types

  • compiler/rust/bitset: Add a new ConstBitSet type

  • compiler/rust: Add an EnumAsU8 trait

  • nak: Use EnumAsU8 for RegFile

  • panfrost: Initial rust build system support

  • panfrost: Add the basis for the new Kraid compiler

  • kraid: Add a GPU model abstraction

  • kraid: Add DataType and NumericType enums

  • kraid: Add a swizzle struct

  • kraid: Add SSAValue and SSARef structs

  • kraid: Add Src/Dst data types

  • kraid: Add an Opcode trait and Op enum

  • kraid: Add Instr, BasicBlock, and Shader structs

  • kraid: Add a builder

  • kraid: Start parsing NIR shaders

  • kraid: Parse the NIR CFG

  • kraid: Handle load_const instructions

  • kraid: Handle nir_op_mov/vec/[un]pack

  • kraid: Add some float alu ops

  • kraid: Implement nir_op_iadd

  • kraid: Handle a few NIR intrinsics

  • Kraid: re-indent shaders for prettier printing

  • kraid: Add a validator to check IR invariants

  • kraid: Add a super simple register allocator

  • kraid: Plumb through Model::encode_shader()

  • kraid: Rework swizzles

  • kraid: Print ASM swizzles when we have them

  • kraid: Copy the bitview module from nouveau

  • kraid: Add a FlowCtrl struct

  • kraid: Replace OpEnd with OpNop.end

  • kraid: Move proc/lib.rs to proc/macros.rs

  • subprojects: Pull in the Rust xml crate

  • kraid: Add ISA XML for v9-15

  • kraid: Add the start of encoder code-gen

  • kraid/isa: Add a simple XML parser

  • kraid/isa: Generate enums with [Try]Encode/Decode

  • kraid/isa: Add an encoder for expressiosn

  • kraid/isa: Add an encoder for instructions

  • kraid/isa: Add support for field modifiers

  • kraid: Add the start of a v9 encoder

  • kraid/isa: Specially handle small_constant_t

  • kraid: Add a SmallConstant struct and a Model::small_constants() hook

  • kraid: Add a lower_small_constants() pass

  • kraid: Add a very dumb message slot assignment pass

  • kraid: Implement shifts and logic ops

  • kraid: Implement integer comparisons

  • krai/isa: Expose a new InstructionInfo struct per-instruction

  • kraid: Break v9 instruction encoding out into traits

  • kraid: Use instruction info to implement op_is_message()

  • kraid/isa: Add a special case in to_snake/camel_case() for data types

  • kraid/isa: Emit TryFrom<DataType> for all data-type-like enums

  • kraid: Clean up the data type mess in the encoder

  • kraid: Implement OpCSel and nir_op_[ui]min/max

  • kraid: Claim we use 64 registers

  • kraid: Support signless IAdd

  • kraid: Implement nir_op_u2u/i2i

  • kraid: Add a SrcRef::Zero

  • kraid: Add a 16-bit ALU lowering pass

  • kraid: Implement nir_op_extract_*

  • kraid: Be more lax about immediates

  • kraid: Map H01 and B0123 to None in the encoder

  • kraid: Implement nir_op_f2f*

  • kraid: Make Instruction::get_info() more ergonamic

  • kraid: Add a Model::op_src_supports_imm32() query

  • kraid/isa: Handle field restrictions

  • kraid: Box ops inside Op

  • compiler/rust/smallvec: Implement Clone, Default, and new()

  • compiler/rust/smallvec: Add a push_mut() method

  • compiler/rust/smallvec: Implement Deref[Mut]<Target = [T]>

  • compiler/rust/smallvec: Implement Extend<T> for SmallVec<T>

  • compiler/rust/smallvec: Implement From<Vec<T>>

  • compiler/rust/smallvec: Implement FromIterator and From<[T; N]>

  • compiler/rust/smallvec: Implement IntoIterator

  • compiler/rust/smallvec: Implement From<SmallVec<T>> for Vec<T>

  • nak: Simplify BasicBlock::map_instrs()

  • nak/builder: Use some of the SmallVec improvements

  • nak: Simplify our SmallVec usage

  • compiler/rust/smallvec: Hide the enum

  • compiler/rust/smallvec: Optimize extend()

  • nir: Allow atomic intrinsics to have multiple components

  • spirv,nir: Add support for AtomicFloat16VectorNV

  • nak/nir: Lower f16vec4 atomics to 2xf16v2

  • nak: Rename AtomType::F16x2 to F16v2

  • nak/from_nir: Handle f16v2 atomics

  • nvk: Advertise VK_NV_shader_atomic_float16_vector

  • kraid: Make SrcRef::Imm32 explicitly non-zero

  • kraid: Make SrcRef PartialEq

  • kraid: Add map_instrs() methods to Shader and BasicBlock

  • kraid/builder: Store the model in builders

  • kraid: Split DataType into two enums

  • kraid/v9: Fix encoding of high register numbers

  • kraid/v9: Rework the shift_lop encode macro

  • kraid/v9: Add the rest of the shift/lop ops

  • kraid: Add None logic and shift ops

  • kraid/v9: Allow immediates in logic ops

  • compiler/rust/bitset: Implement Eq and PartialEq for ConstBitSet

  • compiler/rust/enum_as_u8: Add an EnumAsU8::MAX_DISCRIMINANT

  • compiler/rust/enum_as_u8: Add an ConstU8EnumSet struct

  • compiler/rust/as_slice: Document AsSlice

  • compiler/rust/as_slice: Add a new AsArray trait

  • kraid: Add a VirtualOpcode trait

  • kraid: Add a Model::op_is_supported() query

  • kraid: Add a lanes to Dst

  • kraid/isa: Make Enum::meta a weak reference

  • kraid/isa: Rework enum literals

  • kraid/isa: Make Swizzle EnumAsU8

  • kraid/isa: Treat exact= as a field restriction

  • kraid/isa: Expose allowed swizzles through InstructionInfo

  • kraid/isa: Expose allowed lanes through InstructionInfo

  • kraid: Add a Model::op_src_supports_swizzle() helper

  • kraid: Add a Model::op_dst_supports_lanes() helper

  • kraid: Add the hardware MkVec ops

  • kraid/v9: Fix OpShiftLop::src_supports_imm32()

  • kraid/v9: Fold swizzles and modifiers on imm1w sources

  • kraid: Add a virtual OpCopy and the relevant lowering pass

  • kraid/nir: Emit OpCopy instead of OpMov

  • kraid: Allow 8-bit SSA values

  • kraid: RA per-byte

  • kraid/ops: Claim even more variants

  • kraid/nir: Emit 8-bit ops

  • kraid: Widen ALU ops before RA

  • kraid: Expose the guts of Swizzle

  • kraid/builder: Add copy_iN_to() helpers

  • kraid: Add an OpSwz and a lower_mkvec_swz() pas

  • kraid/nir: Use OpSwz for nir_op_u2uN and nir_op_i2iN

  • kraid/nir: Fix 2x16 extract_[iu]8

  • kraid/nir: Use OpSwz op_extract_*

  • kraid/nir: Implement nir_op_unpack_32_*

  • kraid: Add a new legalize_src_swizzles() pass

  • kraid: Add word() helpers to Src/Dst types

  • kraid/nir: Implement nir_op_unpack_64_*

  • kraid: Better document swizzles

  • kraid: Only dump shaders if KRAID_DEBUG=print is set

  • kraid: Fix RA for dead destinations

  • kraid: Add a Model::op_src_is_staging_reg() helper

  • kraid: Add a Model::op_dst_is_staging_reg() helper

  • kraid: Allocate whole registers for staging destinations

  • kraid: Re-materialize constants

  • panfrost: Set the rustfmt edition to 2024

  • kraid/swizzle: Add a Swizzle::is_none() helper

  • kraid/swizzle: Add an is_none() special case in fold_u32()

  • kraid/swizzle: Take a src_bytes param in Swizzle::bytes_read()

  • kraid/validate: Fix 64-bit destination validation

  • kraid/hw_tests: Allow the test to specify swizzles and lanes

  • kraid/swizzle: Return Option<Swizzle> from AsmSwizzleWiden::to_swizzle()

  • kraid: OpShiftLop is unsigned

  • kraid: Add an SSAValue::bytes() helper

  • kraid: Use a tuple struct for SSAValue

  • kraid: Add OpRegIn and OpRegOut

  • panvk/jm: De-duplicate most of cmd_draw[_indirect]

  • panvk/jm: Re-group setting desc tables and SSBOs

  • panvk/jm: Take a desc_info in meta_get_copy_desc_job

  • panvk/jm: Take a desc_info in prepare_desc/dyn_ssbo()

  • panvk/csf: Take a desc_info in fill_dyn_bufs() and prepare_res_table()

  • panvk: Move desc_info to panvk_shader

  • panvk: Call panvk_lower_nir() before lowering multiview

  • panvk: Add a central panvk_cmd_draw() helper

  • panvk/csf: Make various panvk_draw_info pointers const

  • panvk: Plumb index buffers through panvk_draw_info

  • panvk/jm: Plumb IA state through draw_info

  • panvk/csf: Plumb IA state through draw_info

  • panvk: Improve base instance tracking for indirect draws

  • panvk/csf: Add some sanity assertions in prepare_push_uniforms

  • panvk/csf: Break FS descriptor setup into a new helper

  • panvk/csf: Break VS descriptor setup into a new helper

  • panvk/csf: Prepare descriptors first

  • panvk: Patch VS attribute descriptors as a separate step

  • panvk: Improve panvk_shader_foreach_variant()

  • panvk: Add a helper for uploading to cmd mem

  • panvk/csf: Add a helper for dispatching compute shaders with 3D state

  • compiler/rust: Re-add From<Box<T>> to FromVariants

  • compiler/rust: Only allow FromVariants on enums

  • kraid: Make PAN_USE_KRAID per-stage

  • kraid: Use unsafe with no_mangle

  • kraid/v9: Simplify DstLanes logic for staging registers

  • kraid/ir: Rework some RegRange helpers

  • kraid: Automatically swizzle in From<SrcRef> for Src

  • kraid: Add lowering for COPY.i64

  • kraid: Add OpFMul and plumb it through

  • kraid/nir: Implement nir_op_inot

  • kraid/data_types: Add message types

  • kraid/data_types: Add unit tests

  • kraid: Add OpLea/LdTex and plumb them through

  • kraid: Add OpLd/StCvt and plumb them through

  • compiler/rust: Implement Eq/Hash/PartialEq for LowerBoundedU32Array

  • compiler/rust/bitset: Add an iteration test

  • compiler/rust/bitset: Further generalize find_next_set()

  • compiler/rust/bitset: Add a next_set() method

  • compiler/rust/bitset: Further generalize find_aligned_unset_range()

  • compiler/rust/bitset: Add a find_aligned_set_range() method

  • compiler/rust/bitset: Don’t write past the end in insert_range()

  • compiler/rust/bitset: Generalize ConstBitSet::insert_range()

  • compiler/rust/bitset: Add some range methods to BitSet<usize>

  • compiler/rust/bitset: Add an iter_bit_indices() method

  • compiler/rust: Add more methods/traits to U8EnumSet

  • kraid/nir: Implement load_local_invocation_id

  • kraid: Add OpMux and plumb it through

  • kraid/isa: Handle 16-bit replicated destinations

  • kraid: Add OpFrcp/Frsq and plumb them through

  • kraid/ir: Add a Opcode::set_variant() method

  • kraid: Widen more ops

  • kraid/hw_tests: Use a single basic block

  • kraid: Store blocks in a CFG

  • compiler/rust/bitset: Enable From/IntoBitSet for u32

  • kraid: Copy the SimpleLiveness and LiveSet from NAK

  • kraid: Add a parallel copy builder

  • kraid: Use Swizzle::is_none() more

  • kraid: Add new Phi label type and OpPhiSrc/Dst

  • kraid/nir: Handle nir_phi_instr

  • kraid: Implement EnumAsU8 for DstLanes

  • kraid: Rework supported DstLanes queries

  • kraid: Allow RegRef::word() on subregs

  • Revert “compiler/rust/bitset: Add an iter_bit_indices() method”

  • compiler/rust/bitset: Fix a unit test

  • compiler/rust/bitset: Add a count_set_in_range() method

  • compiler/rust/bitset: Implement Eq and PartialEq

  • compiler/rust/bitset: Add a retain() method

  • kraid: Better RA

  • kraid/nir: Use correct zero sizes for unused ALU components

  • kraid/swizzle: Enable Swizzle::swizzle() on word swizzles

  • kraid/ir,v9: Fix swizzles for the accum source of OpMkVecV2I8I16

  • kraid/ra: Re-swizzle 64-bit sources that read 32-bit values

  • kraid/ra: More accurately compute source constraints

  • kraid/nir: Allow i8v3 ops

  • kraid: Use a tuple struct for SSARef

  • kraid: Implement FromIterator for SSARef

  • kraid: Implement load_ubo

  • kraid/lower_copy: Use Src::imm_u8() for shifts

  • kraid: Use a Builder in ParallelCopy

  • kraid/parallel_copy: Emit small constants directly

  • kraid/ra: Delete a left-over debug check

  • kraid/nir: implement nir_op_[ui](add|sub)_sat

  • kraid/data_type: Add more auto types

  • kraid: Add a DataType::SR special case

  • kraid: Add OpLeaBuf and plumb it through

  • kraid: Add OpTex*

  • kraid/nir: Plumb through texture ops

  • kraid: Use flat_map() instead of map().flatten()

  • kraid/data_type: Handle SR in as_data_type()

  • kraid/v9: Actually encode OpTexGradient

  • kraid/v9: Fix src_supports_imm32() for Op[IF]Add

  • kraid: Add a vec src legalization pass

  • kraid/model: Add an op_src_supports_mod() query

  • kraid: Add a word-based copy propagation pass

  • kraid/widen: Don’t widen messages

  • kraid/nir: Enable load_global_constant

  • kraid/v9: Use the right data type for OpShiftLop::src_supports_imm32()

  • kraid: Add OpAtom* and plumb them through

  • kraid/v9: Don’t allow src0 swizzles in OpShiftLop::src_supports_imm32()

  • kraid: Run copy-prop after legalizing_src_swizzles()

  • kraid/copy-prop: Don’t propagate SSA values with mismatched sizes

  • kraid/swizzle: Expose the guts of swizzle composition

  • kraid/copy-prop: Add byte-based copy propagation

  • kraid: Pass the immediate to Model::op_src_supports_imm32()

  • kraid/v9: Support immediate buffer/texture handles

  • kraid/nir: Always use a destination for AtomOp::Xchg

  • kraid/ra: Handle OpPhiSrc with a swizzle

  • kraid: Call pan_shader_update_info()

  • kraid/nir: Respect FLOAT_CONTROLS_ROUNDING_MODE_RTZ

  • pan/nir: Lower read_invocation to 32 bits

  • compiler/rust/cfg: Assert that nodes are in a dominance-respecting order

  • compiler/rust/cfg: Unexpose CFG::from_blocks_edges()

  • compiler/rust/cfg: Make sorting optional in CFGBuilder::as_cfg()

  • kraid: Stop re-sorting blocks with CFGBuilder

  • kraid: Return an Option<RegRef> from Model::preload_reg()

  • kraid/nir: Use FAURef::user_i32()

  • kraid: Add special FAUs

  • kraid: Add OpBarrier and plumb it through

  • kraid/nir: Respect access flags on loads/store ops

  • kraid: Plumb TLS size through to pan_shader_info

  • kraid/nir: Implement load_scratch/shared_base_ptr

  • pan/nir: Lower scratch and shared to global for Kraid

  • kraid: Add OpWMask and plumb it through

  • kraid/nir: Implement load_subgroup_invocation

  • kraid: Add OpClper and plumb it through

  • kraid: Don’t report Src::is_zero() with a BNot modifier

  • kraid/copy-prop: Trivialize zero copies

  • kraid/ir: Rename the raw src/dst type helpers

  • kraid: Add a DataType::total_bytes() helper

  • kraid/validate: Fix source swizzle validation

  • kraid/ra: Fix W1 widens

  • kraid: Fix lower_small_constants() for 64-bit sources

  • kraid: Take a DataType in Opcode::is_valid_variant()

  • kraid/ra: Also handle OpPhi swizzles in the pre-existing live-out case

  • kraid/copy-prop: Try to re-type opcodes for more widening

  • kraid/copy-prop: Fold widen ops into 64-bit sources

  • kraid/copy-prop: Treat F16ToF32 as a widening copy

  • compiler/rust/bitset: Improve test_find_aligned_unset_range()

  • compiler/rust/bitset: Enhance find_aligned_[un]set_range()

  • nvk/image: Style nits

  • nvk/image: Rewrite nvk_image_can_compress() to use early returns

  • nvk/image: Take an nvk_physical_device in can_compress()

  • nvk: Add an NVK_DEBUG=no_compression flag

  • vulkan/meta: Use z_off/scale for 2D array images as well

  • vulkan/meta: Allow resolving a 2D MSAA image to a 3D image

  • kraid/ra: Fix find_unpinned_bytes() for unaligned ranges

  • kraid/ra: Relax alignment requirements for staging registers

  • kraid/nir: Rework mov/vec handling

  • kraid/nir: Implement nir_op_insert_*

  • kraid/nir: Implement as_uniform

  • kraid: Add a Src::fneg_zero() helper

  • kraid: Use FMA instead of FMUL

  • kraid/ir: Don’t compare labels in FAU/RegRef.eq()

  • kraid/nir: Add a special_fau() helper

  • kraid: Add a new FAUModel

  • kraid: Merge legalize_immmmediates and legalize_vec_srcs

  • compiler/rust: Add U8EnumSet::len() and ConstBitSet::len()

  • kraid/legalize: Add a move_src_to_tmp() helepr

  • kraid: Legalize FAU sources

  • kraid: Add OpIDpAdd and plumb it through

  • nvk: Replace nvk_addr_range with VkDeviceAddressRange

  • pan: Take a stage parameter to get_nir_shader_compiler_options()

  • pan: Move PAN_USE_KRAID into pan_compiler.c/h

  • kraid: Expose our own NIR compiler options

  • pan: Use Kraid’s NIR options when it’s enabled

  • kraid/isa,model: Add a op_srs_is_64bit() query

  • kraid: Align registers based on the new ISA query

  • kraid/copy-prop: Handle 64-bit OpShiftLop

  • kraid/nir: Implement 64-bit op_bitfield_select

  • kraid/nir: Enable more 64-bit ops

  • nir: Add combined shift-logic ops for panfrost

  • kraid: Use the new NIR shift+logic ops

  • kraid: Optimize shift+logic ops

  • kraid: Document a couple passes

  • kraid/nir: Implement nir_op_[iu]mul_2x32_64

  • nvk: Advertise minStorageBufferOffsetAlignment=4 for VKD3D

  • docs: Add a note about Vulkan implicit sync in the 25.3.0 release notes

  • compiler/rust/cfg: Remap node edges in remove_unreachable()

  • compiler/rust/nir: Implement Send+Sync for nir_shader_compiler_options

  • kraid: Use nir_shader_compiler_options directly

Feelthepain77 (1):

  • freedreno: add Adreno 613 (Snapdragon 4 Gen 2) to device list

Filip Gawin (5):

  • r300: avoid UB through implicit conversions on 32bit

  • r300: use uint32_t instead of long in vertprog

  • nv30: fix truncated values in line_stipple_pattern

  • nv30: fix 1 << 31 issues

  • nv30: fix another left shift cannot be represented in type ‘int’

Francisco Jerez (4):

  • nir/divergence: Allow local_invocation_id.z to be treated as uniform.

  • intel/brw: Sort scheduling modes by performance after initial RA failure.

  • intel/brw/swsb: Omit redundant read-after-read synchronization for back-to-back DPAS.

  • intel/brw: Add NIR pass to vectorize dot products into DPAS matrix multiplications.

Frank Binns (21):

  • pvr/ci: drop two tests from bxs-4-64-{fails,flakes}

  • pvr: re-enable {EXT,KHR}_index_type_uint8

  • pvr/ci: add AXE-1-16M nightly Vulkan CTS testing

  • pvr/ci: skip timing out VK reconvergence test for AXE-1-16M

  • pvr/ci: add some timing out tests on AXE-1-16M to skips list

  • pvr: drop unused struct member from pvr_render_pass_attachment

  • pvr: drop unused pvr_descriptor struct

  • pvr: enable KHR_external_semaphore{,_fd} unconditionally

  • pvr: define PVR_USE_WSI_PLATFORM for xcb and xlib

  • pvr: move PVR_USE_WSI_PLATFORM_DISPLAY into a header

  • pvr: advertise VK_EXT_display_surface_counter

  • pvr: advertise VK_EXT_display_control

  • pvr: advertise VK_EXT_direct_mode_display

  • pvr: advertise VK_{KHR,EXT}_surface_maintenance1

  • pvr: advertise VK_{KHR,EXT}_swapchain_maintenance1

  • pvr: advertise VK_EXT_swapchain_colorspace

  • pvr: advertise support for VK_EXT_acquire_drm_display

  • pvr: advertise VK_KHR_unified_image_layouts

  • pvr: rearrange some functions in pvr_arch_border.c

  • pvr: setup all format fields for custom border color entries

  • zink: gate some EXT_descriptor_indexing related code

Frank Bouwer (3):

  • pvr: Fix for depth stencil 2d array writes.

  • Revert “pvr: Fix for depth stencil 2d array writes.”

  • pvr: Fix for depth stencil 2d array writes.

Fyodor Kyslov (1):

  • mesa3d: gfxstream: Add P210 format support

GKraats (2):

  • hasvk: unbreak assert format != ISL_FORMAT_UNSUPPORTED

  • crocus: Fix shader precompilation on Gen6 and higher

Ganesh Belgur Ramachandra (8):

  • amd: import gfx11.7 addrlib

  • amd: add initial common code for gfx11.7

  • radeonsi: add gfx11.7

  • radv: add gfx11.7

  • amd: use gfx_level instead of family_id to choose addrlib

  • amd/llvm: fix target feature setting (DumpCode -> dumpcode)

  • amd/llvm: fix LLVM asserts for signed integer constants

  • amd/llvm: truncate const intergers to bitwidth

Georg Lehmann (118):

  • nir: remove nir_link_xfb_varyings

  • radv: allow input attachment to use pixel coord optimization

  • radv: move per-primitive fixup closer to radv_nir_lower_io

  • radv: move fs view_index handling after lowering io

  • radv: remove unused vs/tes num_outputs from shader info

  • radv: never call nir_assign_io_var_locations

  • radv: remove draw_id from mesh shader a bit later

  • radv: export multi view index as layer after lowering io

  • radv: remove radv_graphics_shaders_link

  • nir: disable fp class analysis for 64bit transcendentals

  • intel/nir_opt_peephole_ffma: fix fp_math_ctlr for modifiers

  • nir/instr_set: allow cse with fp_math_ctrl mismatches for intrinsics

  • nir/opt_varyings: back propagate signed zero information to outputs

  • nir/opt_varyings: do no_signed_zero linking even for non removable stores

  • nir/opt_algebraic: add more fmulz pattern

  • ac/nir/lower_tex_coord: fix moving wqm coordinates

  • nir: fix fp_math_ctrl in fisnan

  • nir/opt_peephole_select: do not count fmul towards the limit when only used by fadd

  • nir/loop_analyze: do not count fmul towards the limit when only used by fadd

  • nir,amd: reassociate fadd to create more fma/mad

  • radv/ci: update restricted trace checksums

  • radv: fix amount of sample shading with required sample shaded inputs

  • ac/nir/lower_tex_coords: fix optimizing cube txd to tex

  • aco: add tests for cube txd to tex opt

  • nir/opt_uniform_subgroup: preserve divergence during optimization

  • tgsi: delete unused lowering pass

  • aco/tests: use explicit lod in sparse texture test

  • spirv: always preserve infinities for FMin, FMax and FClamp

  • radv: use radv_get_sampled_image_desc_size instead of open coding it

  • radv: add radv_force_64_byte_sampled_image dri conf option

  • radv: enable radv_force_64_byte_sampled_image for Forza Horizon 6

  • aco/optimizer: only create v_fma_legacy_f32 when denorms are disabled

  • nir: seperate ffmaz from has_fmulz

  • ac/llvm: don’t assert on 32bit ffma before gfx9

  • ac/llvm: never create ffmaz for broken llvm

  • radv: support VK_KHR_shader_fma

  • aco/gfx8: fix 16bit nir_op_ffma

  • nir/deref: consider atomics that store derefs as complex use

  • aco/gfx6: fix fp64 floor lowering

  • radv: don’t lower dfloor in NIR

  • aco/gfx6: fix fceil lowering

  • aco/gfx6: use shorter lowering for ftrunc

  • aco: add rtne pseudo opcodes for fp64 add and fract

  • aco/gfx6: always use rtne for floor/ceil lowering

  • aco/gfx6: fix fround_even(-0.0)

  • aco/gfx6: always use rtne adds for fround_even lowering

  • radv: enable fp64 float controls on gfx6-7

  • aco/isel: never manually flush denorms after 32bit fma

  • nir: preserve infinities and signed zero during atan2

  • amd/gpu_info: precompute instruction prefetch distance

  • amd/common: don’t pass radeon_info to ac_align_shader_binary_for_prefetch

  • amd/common: add helper for INST_PREF_SIZE

  • radv: remove gfx6 code from ngg emission

  • radv/gfx11+: program INST_PREF_SIZE for compute

  • radv/gfx11+: program INST_PREF_SIZE for pixel shaders

  • aco: add exec_size to prolog/epilog callback

  • radv/gfx12: program SPI_SHADER_PGM_RSRC4_GS for seperately compiled gs

  • radv/gfx11+: program INST_PREF_SIZE for NGG and HS

  • radeonsi: use ac_get_instr_prefetch_size

  • radeonsi: use exec_size from the aco prolog/epilog callback

  • aco/ra: fix inline constants with v_dot2c_f32_f16

  • aco/sched_vopd: fix v_dual_dot2acc_f32_f16 created from VOP2 with inline constant

  • radv: fix setting inline push constants when only the last one is used

  • radv: inline 8 and 16bit push constant loads

  • aco/tests: test v_pk_fmac_f16 and v_dotc_f32_f16 with inline constants

  • aco/tests: test creating v_dual_dot2acc_f32_f16 from v_dot2c_f32_f16 with inline constant

  • nir/skip_helpers: fix stores with ACCESS_INCLUDE_HELPERS

  • nir/skip_helpers: handle vendored store_scratch

  • nir/skip_helpers: keep descriptors uniform even for stores that skip helpers

  • nir/skip_helpers: don’t require helpers for non uniform descriptors

  • aco/assembler: chain branches in emit order

  • aco/assembler: do not abort when exec is written after position exports

  • ac/nir/mem_vectorize: never create vec5 stores

  • aco/assembler: don’t reorder branch insertion block index twice

  • radv/gfx11+: do not use s[0:1] for unused scratch VA in compute shaders

  • radv: remove some dead compute scratch code

  • zink/ci: skip unvanquished-ultra trace on van gogh too

  • panfrost/lower_bool_to_bitsize: do not assume loop phi source order

  • nir/phi_builder: do not sort predecessors for phi sources

  • nir/to_lcssa: do not sort predecessors for phi sources

  • nir: generalize loop simplification

  • nir: add pass to optimize shared variables to subgroup operations

  • nir/opt_algebraic: fix vkd3d-proton pack_half_rtz pattern

  • aco/isel: emit v_mul_i32_i24 for imul with negative constant

  • aco/isel: emit v_mul_hi_i32_i24 for imul_high if possible

  • aco: remove isel setup code for no longer implemented intrinsics

  • nir,amd: split SGPR input intrinsic to specify workgroup divergence

  • ac/lower_intrinsics_to_args: use workgroup divergent ttmp intrinsic for subgroup id

  • nir/divergence: always consider load_ttmp_register_amd uniform

  • amd: use load_scalar_arg_wg_div_amd for workgroup divergent sgprs

  • nir/divergence: alyways consider load_scalar_arg_amd uniform

  • spirv: add option to treat FMax/FMin/FClamp like NMax

  • radv: add radv_force_nan_preserve_min_max option

  • radv: enable radv_force_nan_preserve_min_max for DOOM: The Dark Ages

  • nir: remove explict num_components from nir_def_rewrite_uses_with_alu_src

  • nir/opt_vectorize: prefer to swizzle vector phis at the destination, not the source

  • ac/nir: vectorize phis

  • radv: call nir_opt_phi_precision

  • nir: clean up weird qsort_r usage

  • nir/opt_shrink_vectors: restore load_const deduplication

  • nir/opt_sink: don’t sink comparisons that use ballot(true)

  • radv: run nir_opt_reassociate_for_fma for VS/GS too

  • vulkan/nir_lower_heaps: assume no heap addressing can overflow

  • nir: add num_lsb_zero analysis for 64bit pack and u2u

  • ac/nir_lower_global_access: assume both addition operands are aligned if one is

  • nir: add num_lsb intrinsic index for amd arg loads

  • radv: add dword alignment information to descriptor set/heap pointers

  • ac/nir/lower_ngg: use workgroup divergence analysis for culling

  • ac/nir/lower_ngg: allow reuse of workgroup divergent variables even when subgroup ops are used

  • nir/opt_dead_write_vars: handle atomics as reads

  • nir: fix divergence for deref_cast

  • aco/live_var_analysis: make sure shared vgprs are within the encodable vgprs

  • nir/unsigned_upper_bound: fix float to int conversions

  • nir: mark some AMD specific shuffles as subgroup ops

  • aco/optimizer: fix skip_smem_offset_align

  • nir: support phi sources in nir_rematerialize_deref_in_use_blocks

  • nir/to_lcssa: fix progress for derefs

  • nir/to_lcssa: move constants before the loop instead of creating a phi

Gert Wollny (67):

  • r600/sfn: Add lowering of tess inner and outer default intrinsics

  • r600: replace TGSI TCS passthrough with NIR version

  • r600: replace TGSI query shader with nir

  • r600/sfn: run nir_opt_idiv_const

  • r600/sfn: Avoid creating group-tagged registers for ALU dests

  • r600/sfn: signal progress when splitting address loads

  • r600/sfn: run additional optimization only after successful address split

  • r600/sfn: Extract some helpers from schedule_alu

  • r600/sfn: don’t use return parameters in extracted method

  • r600/sfn: Extract schedule alu groups first

  • r600/sfn: extract fill_alu_group

  • r600/sfn: pass reference to group when possible

  • r600/sfn: Extract group fill failure handling

  • r600/sfn: extract t-slot allocation when filling ALU groups

  • r600/sfn: extract idx load state handling in scheduler

  • r600/sfn: make ALU scheduling return values more meaningful

  • r600/sfn: collaps no_schedule and scheduled

  • r600/sfn: simplify ALU scheduling failure handling

  • r600/sfn: split kcache evaluation into try and commit

  • r600/sfn: make try_kcache_reservation const

  • r600/sfn: move tracking of kcache reservation failure to scheduler

  • r600/sfn: Move tracking of kcache reservation to AluScheduleContext

  • r600/sfn: extract kcache check out of schedule_alu_to_group_vec

  • r600/sfn: refactor BlockScheduler::schedule_block

  • r600/sfn: Move exports emission to helper

  • r600/sfn: extract check and report for unscheduled instructions

  • r600/sfn: deduplicate some code in DCE

  • r600/sfn: deduplicate optimizer logging code

  • r600/sfn: deduplicate fixpoint loop for optimizers

  • r600/sfn: refactor CopyPropFwdVisitor::propagate_to

  • r600/sfn: refactor CopyPropFwdVisitor::visit(AluInsr*)

  • r600/sfn: extract logging from CopyPropFwdVisitor::visit(AluInstr*)

  • r600/sfn: refactor CopyPropBackVisitor::visit(AluInstr*)

  • r600/sfn: Fix typo with AssemberVisitor

  • r600/sfn: Make some member variables references

  • r600/sfn: Refactor AssemblerVisitor emit_alu_op

  • r600/sfn: de-duplicate emit_wait_ack

  • r600/sfn: use c++ pattern for zero-init of structs

  • r600/sfn: extract some byte code emission from assembler

  • r600/sfn: Drop index register handler in assembler

  • r600/sfn: simplify fill bytecode

  • r600/sfn: extract emitting the bytecode of Rat Instr too

  • r600/sfn: extract and decouple ALU post-emit state update

  • r600/sfn: return LDS opcode properties as tuple

  • r600/sfn: use opcode switch in ALU post-emit update

  • r600/sfn: use local opcode consistently in emit_alu_op

  • r600/sfn: Move lds_queue_read decrement out of prepare_alu_src to caller

  • r600/sfn: Move copy_src to sfn_fill_bytecode.cpp, rename to fill_alu_src

  • r600/sfn: Move prepare_alu_src to sfn_fill_bytecode.cpp, rename to fill_alu_src_operands

  • r600/sfn: Move prepare_alu_dst/copy_dst to sfn_fill_bytecode.cpp, rename to fill_alu_dst

  • r600/sfn: drop unused literals tracking in assembler

  • r600/sfn: minor reordering of operations in assembler

  • r600/sfn: Move last_addr handling out of fill_alu_dst

  • r600/sfn: Simplify m_last_addr tracking in emit_alu_op

  • r600/sfn: Validate ALU dst writes in emit_alu_op

  • r600/sfn: Extract ALU bytecode emission helper

  • r600/sfn: Handle dst write checks before mova setup split

  • r600/sfn: Extract ALU dst state update into AssemblerVisitor

  • r600/sfn: Move LDS ALU emission to fill_bytecode

  • r600/sfn: Make emit_alu_op return success status

  • r600/sfn: Move opcode_map to fill_bytecode, pass EAluOp to emit_bytecode_alu

  • r600/sfn: Move ds_opcode_map ownership to fill_bytecode

  • r600/sfn: Drop unused AssemblerVisitor members

  • r600/sfn: Add pin_to_chan method to Register and use it

  • r600/sfn: rename pin_dest_to_chan to pin_registers

  • r600/sfn: Pin alu sources as well when registers are pinned

  • r600/sfn: Drop assertions when emitting IF asm instruction

Gleb Mazovetskiy (1):

  • os_misc.c: add missing include for mach_host_self()

Gleb Popov (1):

  • Rename the CACHE_LINE_SIZE define to MESA_CACHE_LINE_SIZE

Grant Nichol (1):

  • ethosu: Fix -Werror=format build error on 32-bit

Gu, Wangfeng (3):

  • radv/sqtt: add instruction timing SE mask controls

  • radv/sqtt: emit pending barrier end before API markers

  • ac/spm: clamp cache miss counts in derived counters

Gurchetan Singh (14):

  • gfxstream: fix string array marshalling

  • gfxstream: emit global state wrapped decoding for vkCmdEvent

  • subprojects: update libc-rs to 0.2.185

  • subprojects: update to rustix 1.1.4 + downstream patches

  • freedreno: fix ignored qualifier

  • tu: fix -Wmissing-prototypes errors

  • tu: fix implicit fallthrough

  • tu: kgsl: fix -Wgnu-alignof-expression warning with Clang

  • freedreno: explicitly declare required depend_files, part 1

  • freedreno: explicitly declare required depend_files, part 2

  • util: rust: sync error handling fixes from downstream

  • util: rust: minor fixups

  • virtio: add magma-gpu-rs subdirectory

  • docs: fix references to moved crates

Han, Mike (3):

  • amd/vpelib: complete 16bpc RGBA format mapping for 10/12bpc msb/lsb support

  • amd/vpelib: add format support check

  • amd/vpelib: Add missing argb variant support

Hans-Kristian Arntzen (21):

  • wsi/common: Report correct time domain in VkPresentTimingInfo.

  • loader: Separate out X11 specific screen queries from dri_helper.h.

  • loader: Clear screen resources struct on init.

  • wsi/x11: Setup screen resources on x11_connection creation.

  • wsi/x11: Add helper to find appropriate screen resources for a window.

  • wsi/x11: Set up screen resources on swapchain creation.

  • wsi/x11: Add helper to compute xrandr rate estimate.

  • wsi/x11: Update xrandr refresh estimate on geometry change.

  • wsi/x11: Update refresh rate estimate based on MSC feedback.

  • wsi/x11: Implement main body of present timing.

  • wsi/x11: Add Xwl support for present timing.

  • wsi/x11: Only accept VRR refresh rates when we’re flipping.

  • wsi/common: Prefer host query resets when available.

  • wsi/common: Pass along requested timing feedback as well.

  • wsi/x11: Avoid non-causal present timings when not flipping.

  • wsi/common: Refactor out the search for a present_timing struct.

  • wsi/common: Ensure that google display timing results propagate.

  • wsi/common: Always ensure that we can get a GPU done timestamp.

  • wsi/x11: Be more adaptive in how much the sleep is pulled back.

  • radv: Consider VkImageView usage rather than VkImage usage in feedback.

  • radv: Only consider default feedback loops for appropriate layouts.

Hsieh, Mike (2):

  • amd/vpelib: add optional __stdcall calling convention via build option

  • amd/vpelib: add indirect shaper config support

Hyunjun Ko (13):

  • anv/video: fix up H.264/H.265 encode session parameters to match advertised caps

  • anv/video: fix to set the upper bound of the bitstream of h265.

  • anv/video: Add to check size mismatch during motion field estimation.

  • anv/video: define ANV_VIDEO_AV1_MAX_DPB_SLOTS

  • anv/video: Change size of the cached array of recently decoded AV1 frames.

  • intel/genxml: update VDENC commands for gen125

  • anv/video: Add h264 vdenc tables from media-driver

  • anv/video: Make H264 encoder work on Gen125

  • anv/video: fix to set valid coded size for the source pictures.

  • anv/video: Add h265 vdenc tables from media-driver

  • anv/video: Make H265 encoder work on Gen125

  • anv/video: Enable video encoding on gen125

  • anv/video: Support H265 10-bit encoding

Iago Toral Quiroga (2):

  • pan/bi: TEX_GRADIENT may need helper invocations

  • CODEOWNERS: update broadcom maintainers

Ian Romanick (21):

  • brw: Lower all phis to scalar

  • brw: Don’t lower phis involved in DPAS instructions to scalar

  • brw: Calcuate divergence before brw_from_nir

  • nir/opt_constant_folding: Don’t fight with nir_lower_bit_size

  • nir: Use nir_instr_remove_v in nir_def_replace

  • nir/opt_if: use nir_def_replace() instead of nir_def_rewrite_uses()

  • nir/opt_if: Merge if-statements with inverted conditions

  • nir/algebraic: Convert bcsel of addition to addition of b2i or b2f

  • nir/opt_shrink_stores: Don’t shrink ivec2 stores to int64 images

  • brw: Use nir_opt_shrink_stores

  • brw: Use nir_opt_shrink_vectors

  • brw: Add functions to calculate flags usage without a brw_inst

  • brw: Replace logical operations with predication

  • brw: Use nir_opt_uub

  • brw: Use nir_opt_fp_math_ctrl

  • elk: Use nir_opt_uub

  • elk: Use nir_opt_fp_math_ctrl

  • nir/divergence: Handle SYSTEM_VALUE_INSTANCE_INDEX

  • brw/predicate: Add missing test with farther_flags

  • brw: Handle empty top block in brw_nir_move_interpolation_to_top

  • brw/validate: Gfx11 can’t have accumulator src0 in 3-src instructions

Icenowy Zheng (36):

  • pvr: follow other drivers’ practice for copying build ID

  • pvr: skip emitting query program when copy result / reset with 0 queries

  • isaspec: decode: manually print the sign when printing NaN float values

  • pvr: wait for graphics jobs in CopyQueryPoolResults

  • pvr: increase maxPerStageResources for new maxPerStageDescriptorStorageBuffers

  • pvr: do not setup deferred RTA clear for active render targets

  • pvr: properly handle deferred RTA clears for 2D array view of 3D image

  • pvr: add deferred RTA clear command to list after checking it’s not NULL

  • pvr: record deferred RTA clears for secondary cmdbuf subcmds

  • pvr: ignore DS attachment’s D or S when it’s unused in dynamic rendering

  • dri: try to enable GL_ARB_compatiblity when supported GL core version is 3.1

  • pvr: setup viewindex if the shader wants it even when multiview disabled

  • pvr: prohibit clang-format from touching the dri options list

  • pvr: add dri options used by common WSI code

  • pvr: fix handling of invalid attachment info in pvr_init_fs_outputs_mrt

  • pvr: copy sub_cmd flags except owned when executing subcmds out of pass

  • pvr: stop to derive rt datasets based on geometry_terminate

  • pvr: add a structure containing data kept for suspended renderpasses

  • pvr: preserve and pass more data for suspending render passes

  • pvr: remove dEQP-VK.pipeline.monolithic.misc.no_rendering from fail list

  • pvr: return FORMAT_NOT_SUPPORTED for unknown image types

  • pvr: prevent direct access to VkImageSubresourceLayers::layerCount

  • pvr: implement CmdBindIndexBuffer2

  • pvr: implement GetRenderingAreaGranularity

  • pvr: implement GetImageSubresourceLayout2

  • pvr: implement GetDeviceImageSubresourceLayout

  • pvr: advertise VK_KHR_maintenance5

  • zink: move maint5 to gl21_baseline capabilities set

  • docs/zink: add maint5 to the list of required extensions

  • llvmpipe: stub other functions inside compute shaders for ORCJIT

  • pvr: bump conformance version to 1.4.3.3

  • Revert “pipe-loader: fallback to zink instead of kmsro for render nodes”

  • pipe-loader: use zink for powervr device nodes

  • zink: check Z/S aspect before creating Z/S image view

  • pvr: apply the culling everything viewport shift for only triangles

  • vulkan: update spec to 1.4.354

Iván Briano (15):

  • anv: silence warning

  • intel/brw: add load_coverage_mask_intel intrinsic

  • intel/brw: add load_msaa_rate_intel intrinsic

  • intel/brw: add load_frag_shading_rate_intel

  • anv/brw: add conservative raster on/off to FS_CONFIG

  • anv/brw: handle FullyCoveredEXT

  • anv: add and use a drirc option to enable FullyCovered for vkd3d

  • anv: fix return of cmd_buffer_set_indirect_stride() function

  • anv, iris: fix MOCS Index setting of EXECUTE_INDIRECT_* commands

  • intel/dev: ARL-H supports EXECUTE_INDIRECT_*

  • anv: don’t try to clear d/s attachments not backed by an image

  • brw/rt: fix max_t selection on intersection report

  • brw/rt: split HitAttribute area in pending/committed

  • brw/rt, anv: reduce maxRayHitAttributeSize

  • anv: fix 2d-array to 3d blits

Jaakko Jokinen (1):

  • nir: Add cases to nir_get_io_offset_src_number()

JaeHoon Lee (25):

  • v3d: release the texture reference if shadow resource creation fails

  • v3d: drop the tiled temporary when bailing on unsupported blits

  • v3d: free the cache buffer when loading a corrupt disk cache entry

  • v3d: create the compute job after the zero-sized dispatch check

  • v3dv: only report 16-bit float formats as blendable at 32/64 bpp

  • vc4: fix last_layer selection in the blit sampler view

  • v3d: fix slot and input indexing in v3d_set_global_binding

  • v3dv: report maxDrawIndirectCount of 1 without multiDrawIndirect

  • v3dv: honor wait dependencies for job-less submissions

  • v3d: clamp transform feedback offset to buffer size

  • v3dv: report the correct dynamic storage buffer UAB limit

  • v3dv: fix blake3 key truncated to 20 bytes in pipeline cache

  • v3d: fix blake3 key truncated to 20 bytes in shader cache

  • vc4: fix incorrect resource unref in vc4_flush_resource

  • vc4: free vertex and constant buffers on context destroy

  • broadcom/compiler: really enable GFXH-1625 TMUWT validation

  • broadcom/compiler: validate magic waddr writes

  • broadcom/qpu: remove empty qpu_validate.c

  • v3dv: make room in the descriptor map for the no-sampler entries

  • v3dv: use the binning VS variant for the binning VPM config

  • v3dv: record the multiview geometry shader with its Vulkan stage bit

  • v3dv: record the no-op fragment shader with its Vulkan stage bit

  • nvk: free copy_memory_indirect_temps on command buffer destroy

  • v3dv: avoid restoring stale descriptor state after a meta op

  • nvk: report fills from memory correctly

Jaishankar Rajendran (2):

  • vulkan/runtime: enable parametrization of ASTC software decode

  • anv: tune parameters of the ASTC software decoding

Jakob Sinclair (13):

  • panvk: Enable scissor_mode for draws

  • panvk: Remove unnecessary functions

  • vulkan/meta: Don’t issue a full drawcall for clears

  • pan: Support lowering D24X8 to D24

  • gallium: fix type size in z24_unorm_packed_pack_z_32unorm

  • pan/va: Decode support for ARSHIFT_OR on Valhall

  • pan: Add G52 skip for xlib wsi failure

  • pan/compiler: fix spilling for 64-bit values

  • pan: Add missing v14 primitive flag

  • panvk/draw: Separate build from prepare functions

  • panvk/csf: Use RUN_FULLSCREEN for cmd_draw_rects

  • panvk/csf: Use RUN_FULLSCREEN for cmd_draw_volume

  • pan/crc: Fix CRC check for sparse AFBC images

Jan Meisel (3):

  • nir/range_analysis: handle msad_4x8 in unsigned upper bound

  • radeonsi/vcn: fail feedback for truncated encodes

  • radv: fix RADV_PERFTEST=nircache enablement

Janne Grunau (7):

  • nir/gather_info: clear interpolation qualifiers only in fragment stage

  • asahi: nir: lower flrp64

  • asahi: ci: Drop no longer failing VK.wsi.xcb.present_timing test

  • panfrost: ci: Drop no longer failing VK.wsi.xcb.present_timing test

  • asahi: ci: Add failing b10g11r11 and e5b9g9r9 copy tests

  • hk: xfb: Avoid assertions in nir_slot_num_components

  • poly: Fix comment after moving passthrough_gs

Jason Macnak (6):

  • gfxstream: Override VkDeviceDeviceMemoryReportCreateInfoEXT vk.xml

  • virtgpu_kumquat_ffi: replace mutex.get_mut() with mutex.lock()

  • gfxstream: support testing d32 s8

  • gfxstream: kumquat: validate device dmabuf support before use

  • gfxstream: route vkGet*ProcAddr to VkDecoderGlobalState

  • gfxstream: Avoid transfering VkAllocationCallbacks between guest and host

Jeremy Gebben (5):

  • kk: Implement VK_KHR_dynamic_rendering_local_read

  • kk: Refactor encoder state

  • kk: Set availability for extra multiview queries in vkCmdEndQuery()

  • kk: Implement VK_QUERY_TYPE_TIMESTAMP

  • kk: Fix Vulkan to Metal stage translation for timestamps

Jeremy Huddleston (38):

  • bin/install_megadrivers: Bail out if libname suffix is never reached

  • gallium/targets: Use libname_suffix for installed driver names

  • glx/apple: Convert K&R-style declarations to ANSI prototypes

  • glx/apple: Switch logging to os_log on macOS 10.12+

  • glx/apple: Replace apple_glx_diagnostic with apple_glx_log_*

  • glx: free visinfo on BadMatch in glXCreateWindow’s AppleGL path

  • glx: Fix stale end-comment on __glXInitialize direct-rendering block

  • glx: drop redundant __glXErrorString forward declaration

  • glx: fix DRI3-not-available diagnostic skip on macOS

  • glx: NULL-check frontend_screen in glXCreateContextAttribsARB

  • glx: simplify FBConfig wire decode

  • glx: free glx_drawable on CreateDRIDrawable failure

  • glx: bail bind_extensions on screens without frontend_screen

  • glx: drop dead AppleGL glXGetProcAddressARB fallback

  • glx/apple: Add create_context_attribs entry to the applegl_screen_vtable

  • zink: fix GLX_USE_APPLE typo (should be GLX_USE_APPLEGL)

  • glx: drop dead GLX_USE_APPLE check inside glXSwapBuffers

  • glx: Guard declaration of glx_accel and kopper to match use

  • glx/apple: free gc in applegl_destroy_context

  • glx: fix per-display drawHash / zombieGLXDrawable / dri2Hash leak on GLX_USE_APPLE builds

  • glx/apple: silence OpenGL deprecation warnings

  • glx/apple: return CGLError from apple_visual_create_pfobj instead of aborting

  • glx: route copy_context through a vtable slot

  • glx: route swap_buffers through a vtable slot

  • glx: extract drawable lifecycle into a vtable

  • glx/apple: allow selection between AppleGL and Gallium at runtime for GLX_USE_APPLE=1 builds

  • glx/apple: skip AppleGL election on macOS 26 and newer

  • glx: remove GLX_USE_APPLE and collapse the guards it gated

  • glx: Fold __glXGetDrawableAttribute and __glXQueryDrawable together

  • glx: Fold CreatePbuffer/DestroyPbuffer into CreateDrawable/DestroyDrawable

  • zink: Add missing link against libxcb-present

  • zink: Address libvulkan.1.dylib dlopen failure on macOS

  • glx/apple: honor the client-requested GLX context version and profile

  • llvmpipe: link all LLVM targets on Apple to fix build failure when using static LLVM libraries

  • dri/st: Fall back to Z32_FLOAT depth configs when Z32_UNORM is unsupported

  • glx/apple: only skip AppleGL election on macOS 26.0 through 26.5

  • llvmpipe: don’t create a screen when the process is not allowed to JIT

  • glx/apple: silence OpenGL deprecation warnings in libglx

Jesse Natalie (23):

  • d3d12: Handle THREAD_SAFE maps and use them for async query results

  • microsoft/compiler: Back-propagate interpolator modes from FS

  • wgl: Use an hwnd xor hdc for framebuffers

  • d3d12: add screen pending-free list plumbing

  • d3d12: clear stale per-context BO state at context destroy

  • d3d12: transfer batch local_bos refs to screen at submit

  • d3d12: transfer batch->bos refs to screen at submit

  • d3d12: reclaim in-flight BO memory on allocation failure

  • d3d12: implement pb_fence vtbl for cache/slab reuse

  • d3d12: drop peer-batch peeking in resource_is_busy / wait_idle

  • d3d12: proactively trim completed pending-free entries

  • nir_lower_non_uniform_access: Add ASSERTED for assert-only var

  • va: Wrap assert-only code in NDEBUG

  • microsoft/compiler: Don’t assume phi ordering

  • util: Fix u_math on MSVC arm64

  • mesa/st: PBO memory barriers imply image barrier if PBO download goes through compute

  • d3d12: Use enhanced barriers for memory_barrier when we can

  • d3d12: Fix transition_array_size for 3D textures

  • d3d12: Fix WARP version detection for broken int64

  • ci/windows: Update WARP to 1.0.20

  • wgl: Move sub-8bpc pixel formats to extended format list to match other Windows drivers

  • d3d12: Disable vao fast path for AMD

  • mesa: Fix shared state bookkeeping for dynamic share list changes (wglShareLists)

Jhanani Thiagarajan (1):

  • intel/mda: Change the default output directory

Jianfeng Liu (1):

  • freedreno/drm: Fix uninitialized read of BO metadata on import

Jianxun Zhang (1):

  • intel/decoder: Print more information in shader’s headline

Jiyu Yang (3):

  • nir/loop_analyze: Use pass_flags for memoization in is_only_uniform_src

  • panfrost: cleanup precomp_cache on screen destroy

  • egl/dri2: exclude >8bpc configs from GLES1 renderable/conformant bits

Job Noorman (60):

  • ir3/ra: fix killed src detection while spilling

  • ir3/shared_ra: fix live-out reload after src reload

  • nir/get_io_offset_src_number: support @load/store_global_ir3

  • ir3/isa: use same src for ldg.a OFF field on a6xx/a7xx

  • ir3: always use byte offset for @load/store_global_ir3

  • nir/opt_offsets: add support for @load/store_global_ir3

  • ir3: move feature check down in ir3_nir_max_imm_offset

  • ir3: enable opt_offsets for load/store_global_offset

  • ir3: mark __alias_n as UNUSED in foreach_src_in_alias_group_n

  • ir3/cf: fix rewriting uses with different dst types

  • ir3/shared_ra: use ir3_cursor instead of instr in reload helpers

  • ir3/shared_ra: insert reloads before tied dst pcopies

  • ir3/cp: support propagating const vecs

  • ir3: allow const src0 for ldg.a/stg.a/ray_intersection

  • ir3: don’t cache driver param instructions

  • ir3: allow (ss) on all cat7 instructions

  • freedreno/computerator: fix UAV view size

  • ir3/spill: extract child intervals for live-in reloads

  • ir3/ra: add ir3_ra_src_is_killed helper

  • ir3/ra: fix killed src detection for spillall min limit

  • ir3: use a1.x addressing for ldg.k with dst 256

  • ir3: don’t use bitfields in ir3_shader_output

  • ir3: don’t store shader_options in the cache

  • freedreno/drm-shim: allow chip selection by chip_id

  • tu: use chip_id instead of gpu_id for the cache UUID

  • tu: add option to override the build ID

  • vulkan: add vk_shader_module_hash helper

  • vulkan: use consistent module hashing for pipeline stages

  • nir/lower_undef_to_zero: add filter argument

  • ir3: lower undef booleans to zero

  • nir/get_io_index_src_number: support @load_ssbo_address

  • nir/lower_ssbo: take offset_shift into account

  • nir/lower_ssbo: add option to only lower large SSBOs

  • nir/lower_ssbo: add option to insert bounds checks

  • tu: Add option to raise the maximum SSBO size

  • ir3: fix possible signed overflow in ir3_link_add

  • ir3/opt_prefetch_descriptors: rematerialize defs at preamble start

  • nir/lower_vars_to_scratch_global: make callback deterministic

  • ir3/lower_vars_to_scratch_global: use stable sort for variables

  • nir: add nir_shader_deref_pass

  • nir: add nir_src_as_{alu,tex,phi}_src helpers

  • nir: add nir_convert_address_format pass

  • nir/lower_explicit_io: add support for 64bit_global_32bit_offset vars

  • nir/lower_explicit_io: support shifting non-const array index

  • nir: add load/store_global_offset intrinsics

  • nir/lower_explicit_io: add support for load/store_global_offset

  • nir/set_io_offset: add support for adjusting BASE

  • nir/lower_explicit_io: support offset_shift for 64bit_global_32bit_offset

  • rusticl/kernel: add support for 64bit_global_32bit_offset

  • ir3: add support load/store_global_offset

  • ir3: enable opt_offsets for load/store_global_offset

  • tu,ir3: use 64bit_global_32bit_offset for global memory

  • ir3/lower_tess: use load/store_global_offset

  • ir3/lower_shader_clock: use load_global_offset

  • ir3: don’t manually lower load/store_global

  • tu/lower_ray_query: use load_global_offset

  • tu,ir3/analyze_ubo_ranges: use load_global_offset

  • tu/lower_ssbo_address_size: use load/store_global_offset

  • nir,ir3: remove load/store_global_ir3

  • ir3/opt_preamble: lower load_global_offset to preamble

Joe Wang (2):

  • ac/spm: bump AC_SPM_MAX_COUNTERS_PER_GROUP to 16

  • radv,ac/spm: add user-defined raw counter collection

Jon Turney (7):

  • ddebug: Fix use of alloca() without #include “c99_alloca.h”

  • glx/windows: Avoid shadowing ‘type’ parameter of driwindowsCreateDrawable()

  • glx/windows: Add stdbool.h include to ‘direct GLX via WGL’ implementation

  • glx/windows: Fix compilation of driwindows_glx after driscreen changed from pointer to member

  • glx/windows: Fix compliation after code motion to put event base in ‘dri’ context

  • glx/windows: Add GLX_USE_WINDOWSGL in new places it’s needed to build libGL

  • glx/windows: Drop static from driwindowsCreateScreen()

Jordan Justen (48):

  • brw: Don’t set header_size at init since it will be re-set in later code

  • brw/compact: Precompact using 2src fields on 3src instructions

  • intel/gen: Add gen 9 through Xe2 instruction formats in JSON

  • intel/gen: Add gen_inst_info.py script to generate C++ headers

  • intel/gen: Create gen_info_util.h

  • intel/gen: Make use of generated instruction info

  • intel/gen/compact: Add compact tables from brw/brw_eu_compact.c

  • intel/gen: Add gen_raw_compact_inst type

  • intel/gen: Add gen_compact_accessor for compact/uncompact

  • intel/gen: Implement compact support

  • intel/gen: Split out type decode functions for use with uncompact

  • intel/gen: Implement uncompact support

  • intel/gen: Account for compact nop pad instruction in gen_scan_raw_layout()

  • intel/gen: Support declaring ISA fields with disconnected bits

  • intel/gen: Support accessing fields & sub-fields with disconnected bits

  • intel/gen: Merge THREE_SRC0_VSTRIDE HI/LO into a gen_split_range

  • intel/gen/xe: Merge THREE_SRC1_VSTRIDE HI/LO into a gen_split_range

  • intel/gen/xe: Merge BFN_FUNC_CONTROL HI/LO into a gen_split_range

  • intel/gen: Merge uncompat control bits into a gen_split_range

  • intel/gen: Merge uncompat datatype bits into a gen_split_range

  • intel/gen: Merge uncompat subreg bits into a gen_split_range

  • intel/gen: Merge uncompat src0 bits into a gen_split_range

  • intel/gen: Merge uncompat src1 bits into a gen_split_range

  • intel/gen: Merge uncompat 3src control bits into a gen_split_range

  • intel/gen: Merge uncompat 3src source bits into a gen_split_range

  • intel/gen/xe: Merge uncompat 3src subreg bits into a gen_split_range

  • intel/gen: Merge SRC_A16_SWIZZLE HI/LO ranges

  • intel/gen/xe: Merge Xe2 DATATYPE_INDEX HI2/LO3 into a gen_split_range

  • intel/gen/xe: Merge Xe2 compact 3src subreg HI2/LO3 into a gen_split_range

  • intel/gen: Assert that the gen opcode is supported by this platform

  • intel/gen: Start Xe3P support

  • intel/gen: Add gen_byte_stride()

  • intel/gen: Add Xe3P validation for src1 byte stride matching dst

  • intel/gen/validation: Start enabling Xe3P tests, but skip for now

  • intel/gen/validation: Update tests for new Xe3P src1 restriction

  • intel/gen/validation: Drop WA 22016140776 on Xe3P

  • intel/gen/validation: Enable running validation tests for Xe3P

  • intel/gen: Remove mac/mach/macl on Xe3P

  • intel/gen: Add mullh instruction for Xe3P

  • intel/gen/xe: Rename decode/encode_type_3src to decode/encode_type_short

  • intel/gen: Disable compact on Xe3P for now

  • intel/gen: Add Xe3P two source encoding changes

  • intel/gen/tests/basic: Add 2src round-trip tests covering nvl src1 changes

  • intel/gen: Add Xe3P three source src1 encoding changes

  • intel/gen/tests/basic: Add 3src round-trip tests covering nvl src1 changes

  • intel/gen/compact: Split datatype into 1src / 2src versions

  • intel/gen: Add Xe3P compact support

  • intel/executor: Enable Xe3P

Jose Maria Casanova Crespo (59):

  • broadcom/compiler: Add V3D 7.1 v8dot dot product QPU instructions

  • broadcom/compiler: hardware-accelerated 4x8-bit dot products on V3D 7.1+

  • broadcom/compiler: Add v8dot and setnnmode scheduler dependencies.

  • broadcom/compiler: Eliminate redundant setnnmode instructions

  • v3dv: Expose hardware-accelerated integer dot products on V3D 7.1+

  • broadcom/compiler: move nir_lower_undef_to_zero out of optimization loop

  • v3dv: bump maxComputeSharedMemorySize to 32 KB

  • v3d/v3dv: Use new V3D_MAX_CSD_WG_SIZE = 256

  • v3dv: lower oversized compute workgroups to 256 invocations

  • v3dv: include mem_offset in vkCmdFillBuffer destination

  • v3dv: Enable KHR_shader_subgroup_extended_types

  • v3dv: expose maxFragmentOutputAttachments as max_rts

  • v3dv: avoid duplicate bo_handles between cpu_job and CSD lists

  • v3dv: assert timestamp pool BO is disjoint from dst buffer BO

  • broadcom/ci: skip SSBO tests close to the 60s threshold on rpi4

  • v3dv: avoid 16F TLB usage for B10G11R11_UFLOAT copies

  • v3dv: advertise VK_EXT_scalar_block_layout on V3D 7.1+

  • v3dv: Enable meta_copy_buffer with TFU for V3D 7.1

  • v3dv: move destroy_update_buffer_cb to a generic helper

  • v3dv: use TFU copy with stride-0 for vkCmdFillBuffer

  • v3dv: extract TFU helpers for format-plane and slice-stride args

  • v3dv: rename copy_buffer_to_image_tfu to copy_buffer_image_tfu

  • v3dv: implement TFU image-to-buffer copy on V3D 7.1

  • v3dv: relax buffer padding in TFU buffer<->image copy

  • v3dv: share zero-fill TFU staging BO at device level

  • v3dv: expose the full simulator memory to applications

  • broadcom/qpu: support output pack on itof/utof

  • v3d: move nir_lower_frexp after nir_lower_bit_size

  • v3dv: lower flrp16 for consistency with flrp32

  • v3d: widen sub-32-bit subgroup arithmetic and vote ops

  • v3d: improve liveness analysis for packed partial writes

  • broadcom/qpu: expose V3D 7.1 packed-f16 instructions

  • v3d: emit packed-f16 ALU ops natively on V3D 7.1

  • v3dv: enable lowered shaderFloat16/Int16/Int8 + VK_KHR_shader_float16_int8

  • broadcom/compiler: fix payload-register liveness condition

  • v3d: Enables GL_ARB_clip_control for v71+

  • v3d: use NO_GUARDBAND clipper for near-zero viewport Z scale

  • v3dv: set non-zero array stride in null texture descriptor state

  • broadcom: add and use max_render_targets to devinfo

  • broadcom: raise framebuffer size to 7680 on V3D 7.1

  • v3dv: gate Dawn-required limits and features behind V3D_WEBGPU_OVERRIDE

  • ci: igalia farm maintenance

  • v3dv: allow TFU readahead padding above maxMemoryAllocationSize

  • v3dv: route blending of UNORM16/SNORM16 RTs through software lowering

  • v3dv: rename format_plane unorm/snorm flags to sw_unorm/sw_snorm

  • v3dv: fix crash on device creation failure before meta initialization

  • v3dv: close the primary node fd on physical device destruction

  • broadcom/compiler: reduce the compile-strategy fallback ladder

  • broadcom/compiler: split v3d_nir_to_vir_finish out of v3d_nir_to_vir

  • broadcom/compiler: support probing a compile’s pre-spill register pressure

  • broadcom/compiler: consolidate the compile-strategy logging

  • broadcom/compiler: move the 2-thread strategies to compile_2t_strategies

  • broadcom/compiler: pick the 2-thread compile strategy by register pressure

  • broadcom/compiler: don’t leak the compile on assembly allocation failure

  • broadcom/compiler: drop V3D_DEBUG=opt_compile_time

  • broadcom/ci: unskip CTS tests that are no longer slow

  • broadcom/compiler: abort 2-thread spill loops over the best result so far

  • vc4: save the fragment constant buffer around the YUV blit

  • vc4: unbind the textures around the blitter clears

José Roberto de Souza (40):

  • anv: Change fill_inline_params() first parameter from struct GENX(COMPUTE_WALKER_BODY) to uint32_t *

  • anv: Move VMA heaps init and finish of vma heaps to anv_va.c

  • anv: Move init and finish of state pools to its own functions

  • anv: Move code to load color border to memory to a function

  • intel/brw: Explicitly upcast UB to UW for SHR with vector immediates

  • intel/tools: Fix parse of ‘[HWCTX].replay_*’ in aubinator_error_decode_xe

  • intel: Sync xe_drm.h

  • intel: Add support for madvise purgeable VMAs in Xe KMD

  • iris: Improve and standardize the behavior of madvice in i915

  • intel/brw: Fix nir_intrinsic_load_inline_data_intel register offset calculation

  • intel/dev: Remove unused intel_get_device_info_for_build() function

  • intel/dev: Add URB max entries values

  • intel/dev: Use URB mesh/task min/max values in intel_device_info

  • intel/dev: Add a Xe2+ table of URB min and max entries

  • anv: Add assert to make sure we don’t push more than max_push_regs to push constants

  • anv: Replace most parameters of fill_inline_param() by a struct

  • anv: Add function to get each anv_state_pool

  • anv: Use anv_device_get_general_state_pool()

  • anv: Use anv_device_get_aux_tt_pool()

  • anv: Use anv_device_get_dynamic_state_pool()

  • anv: Use anv_device_get_binding_table_pool()

  • anv: Use anv_device_get_scratch_surface_state_pool()

  • anv: Use anv_device_get_internal_surface_state_pool()

  • anv: Use anv_device_get_bindless_surface_state_pool()

  • anv: Use anv_device_get_indirect_push_descriptor_pool()

  • anv: Use anv_device_get_push_descriptor_buffer_pool()

  • anv: Replace va.bindless_surface_state_pool access with a function

  • anv: Replace va.indirect_descriptor_pool access with a function

  • anv: Replace va.dynamic_visible_pool access with a function

  • anv: Replace va.internal_surface_state_pool access with a function

  • anv: Replace va.dynamic_state_pool access with a function

  • anv: Replace va.indirect_push_descriptor_pool access with a function

  • anv: Replace va.push_descriptor_buffer_pool access with a function

  • anv: Replace va.aux_tt_pool access with a function

  • anv: Replace va.binding_table_pool access with a function

  • anv: Replace va.scratch_surface_state_pool access with a function

  • anv: Fix memcpy overflows around sampler state

  • anv: Support sampler state of different sizes

  • anv: Replace anv_descriptor_set_binding_layout::descriptor_data_sampler_size by a local variable

  • anv: Fix copy of sampler state for bindless

Juan A. Suarez Romero (36):

  • v3d: mark mapped BO as initialized for valgrind

  • v3d/ci: add OpenCL regressions

  • ci: igalia farm maintenance

  • Revert “ci: igalia farm maintenance”

  • broadcom/ci: update kernel for nightly runs

  • vc4/ci: update expected results

  • loader: check if the kernel driver is amdgpu

  • broadcom/simulator: V3D is always 4.2 or above

  • v3dv: allow device with only render node

  • v3d/drm-shim: add GPU selection

  • v3dv: disable threadeded submissions under drm-shim

  • broadcom/ci: update kernel for nightly jobs

  • Revert “ci: igalia farm maintenance”

  • mesa: allow GL_TEXTURE_COMPARE_{MODE,FUN} with EXT_shadow_samplers

  • v3dv: increase max push constants size

  • Revert “people: update Marek’s email”

  • broadcom/ci: upgrade kernel in DuTs

  • v3d/ci: update expected results and document failures

  • v3dv: fix assertion on push constants

  • st/mesa: release sampler view

  • rusticl: fix leak in `util_queue`

  • v3d: initialize value in query info

  • v3d: free vertex and constant buffers on context destroy

  • v3d: add more blitter ops for saving resources

  • vc4: mark mapped BO as initialized for valgrind

  • vc4: initialize value in query info

  • vc4: move util_copy_constant_buffer to the function beginning

  • vc4: add blitter operations

  • doc/features.txt: enable VK_KHR_shader_float16_int8 for v3dv

  • v3dv: fix buffer creation usage flags validation

  • doc/features.txt: fix VK_KHR_shader_float16_int8 for v3dv

  • v3dv: enable VK_KHR_shader_quad_control / clustered subgroups

  • v3dv: enable VK_KHR_shader_subgroup_rotate / rotate subgroups

  • v3dv: enable VK_KHR_shader_maximal_reconvergence

  • v3d: add support for array of textures blit with TFU

  • v3d: save fragment constants on sand8/sand30 blit

Julia Zhang (15):

  • radv: add new option RADV_DEBUG=notmz

  • radv: allocate encrypted rings BOs

  • radv: create encrypted BOs for protected cmd_buffers

  • radv: enable surface protected capability

  • radv: set TMZ bit in sdma_copy packet

  • radv: save protected queue and non-protected queue seperately

  • radv: advertise VK_EXT_pipeline_protected_access

  • vulkan/wsi: return image compression properties for surface formats

  • vulkan/wsi: pass compression control when filtering DRM modifiers

  • vulkan/wsi: copy swapchain compression fixed-rate flags

  • radv: reject DCC modifiers when image compression is disabled

  • radv: advertise EXT_image_compression_control_swapchain

  • radv/ci: skip compression_control cases

  • radv: implement bo_wait_for_idle

  • radeonsi: avoid unmatched Perfetto events

Julien Schueller (4):

  • glx: avoid crash on glXBindTexImageEXT when no texture target set

  • egl: fix _EGL_NATIVE_PLATFORM fallback for unrecognized native displays

  • drisw_glx: handle XGetGeometry failure in get_drawable_geometry

  • st/drawpixels: tile images larger than max texture size instead of clamping

Kajal Kajal (1):

  • freedreno/blitter: copy full depth of src box in resource_copy_region

Karmjit Mahil (34):

  • freedreno/decode,ir3: Mark decoded dwords as const

  • freedreno/decode: Fix error() in script.c

  • freedreno: Don’t set UCHE_CLIENT_PF

  • gbm: Remove unused ARRAY_SIZE macro

  • gbm: Replace VER_MIN with common MIN2

  • docs: Fix struct redefinition errors

  • util: Add heap_memory_percent driconf option

  • vulkan: Add heap budget helper function

  • hk: Add heap_memory_percent driconf support

  • asahi: Add heap_memory_percent driconf support

  • v3dv: Add heap_memory_percent driconf support

  • v3d: Add heap_memory_percent driconf support

  • tu: Add heap_memory_percent driconf support

  • freedreno: Add heap_memory_percent driconf support

  • panvk: Add heap_memory_percent driconf support

  • panfrost: Add heap_memory_percent driconf support

  • nvk: Add heap_memory_percent driconf support

  • crocus: Add heap_memory_percent driconf support

  • pvr: Add heap_memory_percent driconf support

  • vc4: Use os_get_gpu_heap_size()

  • freedreno/computerator: Remove VLA giving a build warning

  • util/u_trace: Fix copy_func indentation in generated code

  • util/u_trace: Fix indentation of generated _trace function

  • util/u_trace: Evaluate copy_func expression in python

  • util/u_trace: Refactor TracepointArgStruct

  • util/u_trace: Commonize emitted print functions

  • util/u_trace: Move copyright into a Mako def

  • util/u_trace: Commonize trace function header with a Mako def

  • util/u_trace: Avoid sprintf when we already have a string

  • util/u_trace: Add ArgBlob for storing structs in traces

  • util/u_trace: Add u_trace_backend_type

  • tu: Emit bin info in Perfetto render_pass

  • drm-shim/freedreno: Fix fprintf format specifier

  • android_stub: Replace __ANDROID_API_V__ with 35

Karol Herbst (195):

  • nak: the MS location comes last in TLD, same spot as depth compare in TEX

  • mesa/st: do not advertise CL subgroup features on the GL side

  • radeonsi: advertise support for subgroup rotate

  • iris: advertise support for subgroup rotate

  • nak/lower_cf: remove single src phis

  • nak: call nir_opt_fp_math_ctrl

  • nak: call nir_opt_algebraic_distribute_src_mods

  • ci: install libstdc++-static on fedora

  • rusticl: link the C++ runtime statically

  • softfloat: make sign bit an unsigned int

  • nir: add fmul_rtz

  • nir: handle fmul_rtz in a couple of places

  • nak: handle nir_op_fmul_rtz

  • nak: use fmul_rtz for NAK_INTERP_MODE_PERSPECTIVE

  • nir: add fmul_rtz optimizations

  • nir/lower_cl_images: call nir_progress on every function

  • gallivm/nir/soa: use uint for booleans

  • llvmpipe: never pass a NULL function name to LLVMAddFunction

  • nvk: Use nvk_cmd_fill_memory in CmdResetQueryPool when possible

  • include: update CL headers

  • rusticl/program: handle CL_INVALID_CONTEXT for clCompileProgram and clLinkProgram

  • rusticl/kernel: update error code handling for clSetKernelExecInfo

  • rusticl/kernel: return CL_INVALID_WORK_GROUP_SIZE in clEnqueueNDRangeKernel for an explicit 0 workgroup

  • rusticl: start implementing CL 3.1 support

  • rusticl: implement CL 3.1 platform features

  • rusticl: implement CL 3.1 device features

  • docs/features: add OpenCL 3.1 section

  • bin/gen_release_notes: add paragraph on OpenCL support

  • rusticl: update names of types now core in 3.1

  • ci: update OpenCL 3.1 piglit fails

  • clc: do not use std::filesystem

  • Revert “rusticl: link the C++ runtime statically”

  • glsl/softfp: rename ffma to fmad

  • lima: rename ppir_op_ffma to ppir_op_fmad

  • nir: rename ffma to ffma_old

  • nir: rename nir_fmad to nir_fmad_old

  • nir: add new float multiply-add opcodes

  • nir: validate new float_mul_add options

  • nir/tests: handle new multadd opcodes

  • nir/opt_algebraic: add fmad and ffma_weak lowering rules

  • nir: handle new multadd opcodes in lowerings and opts

  • nir: handle new multadd opcodes in helpers

  • nir: duplicate old ffma opts where necessary for new multadd ones

  • nir/tests: use ffma_weak

  • ntt: use ffma_weak

  • llvmpipe: port over to ffma_weak

  • softpipe: keep weak_ffmas around

  • intel/elk: port over to nir_op_ffma

  • intel/jay: support nir_op_ffma

  • intel/brw: port over to nir_op_ffma

  • i915: support nir_op_fmad

  • nv50/ir: port over to new multadd opcodes

  • nv30: advertize new float multadd options

  • nak: port over to nir_op_ffma

  • zink: port over to nir_op_ffma_weak

  • ir3: port to nir_op_fmad

  • freedreno/ir2: use nir_op_fmad

  • tu: use nir_op_ffma_weak in lowering

  • ac: handle new float multadd opcodes

  • ac: use nir_op_ffma_weak

  • ac/llvm: support new multadd opcodes

  • aco: support new multadd opcodes

  • radv: use nir_op_ffma_weak

  • radeonsi: advertize new float multadd options

  • r600,sfn: support new multadd opcodes

  • r300: port over to nir_op_fmad

  • agx: port over to nir_op_ffma

  • kk: support nir_op_ffma

  • bitfrost: support nir_op_ffma

  • microsoft/compiler: support nir_op_ffma

  • d3d12: use nir_op_ffma_weak

  • etnaviv: port over to nir_op_fmad

  • pco: port over to nir_op_ffma

  • pvr: use ffma_weak for lowering

  • lima: support nir_op_fmad

  • svga: use weak_ffma

  • virgl: advertise new muladd options

  • nir: add fmad_or_ffma helpers and use it in lower_double_ops

  • nir: update ffma helpers to use new opcodes

  • nir: make lowering use new ffma opcodes

  • mesa: use ffma_weak

  • vulkan/meta: use nir_op_ffma_weak

  • tgsi_to_nir: translate MAD as ffma_weak

  • glsl: translate fma as fma_weak

  • vtn/glsl: translate fma as ffma_weak

  • vtn: handle OpFmaKHR

  • vtn: use ffma_weak

  • vtn/opencl: map mad to ffma_weak and fma to ffma

  • vtn_bindgen2: keep ffma_weak

  • nir: remove ffma_old

  • ci: update traces due to ffma rework

  • ci/windows: add dEQP-VK.glsl.builtin.precision_double.mix.compute.vec3 fail

  • zink: keep ffma_weak and use GLSLstd450Fma for it

  • zink: support nir_op_ffma

  • nir: add nir_intrinsic_cmat_load_shared_nv to nir_get_io_offset_src_number

  • nak/sm70: add helper for memory load store addresses

  • nak: wire up UGPR Ld/St/Atom encoding

  • nir: add uniform address to nvidia IO intrinsics

  • nak: add UGPR/GPR lowering for load/store/atom instructions

  • nak: optimize iadds with an uniform operand in iadds of address calculations

  • zink: proper advertise keep_weak_ffma for fp16

  • rusticl/kernel: handle nir shader compilation failures gracefully

  • rusticl: more intel compat stuff

  • rusticl/spirv: add SPIRVToNirOptions type

  • rusticl/spirv: silence GenericPointer cap warning

  • rusticl/spirv: properly set float execution mode at spirv_to_nir time

  • gallium: add fp16_no_denorms cap

  • nvk: enable VK_KHR_shader_fma

  • nir/opt_algebraic: add missing fmadz lowering for lower_fmulz_with_abs_min

  • nir/opt_dead_write_vars: cache is_entrypoint of the function

  • nir/lower_alu: fix lower_fminmax_signed_zero for denorms

  • asahi: fix dst range in buffer copy region

  • asahi: fix compute blitter for float16 image copies

  • asahi: move batch flushing into agx_launch_internal

  • asahi: fix fdiv lowering

  • meson: enable more rust 2024 lints

  • rusticl/util: add Traits to help with usage of CString

  • rusticl/kernel: store kernel names as CString

  • rusticl/program: store log as a CString

  • rusticl/program: wrap compiler option parsing

  • rusticl/util: fix rustc-1.95 compilation error

  • vtn/opencl: convert libclc workaround handling to a switch statement

  • vtn/opencl: fix edge case behavior for cospi

  • vtn/opencl: fix edge case behavior for sinpi

  • vtn/opencl: fix edge case behavior for tanpi

  • rusticl/program: print compiler output as Rust string

  • rusticl/util: add CStrExt trait

  • rusticl/util: add CStrExt::from_ptr_or_empty

  • rusticl/program: add CompileOptions::get_clang_args

  • rusticl/program: construct __OPENCL_VERSION__ inside CompileOptions::get_clang_args

  • rusticl/program: turn iter map into loop inside CompileOptions::new

  • rusticl/program: handle -create-library inside CompileOptions::new

  • rusticl/program: move -cl-std handling inside CompileOptions::get_clang_args

  • rusticl/program: set __OPENCL_C_VERSION__ ourselves

  • rusticl/program: store build options as CString

  • rusticl/program: implement CL_PROGRAM_BUILD_OPTIONS without a copy

  • rusticl/program: implement CL_PROGRAM_BUILD_LOG without a copy

  • spirv: set num_components for OpAtomicFlagTestAndSet

  • rusticl/kernel: override libclc shader config helpers

  • nak/instr_sched_prepass: Take predicate spilling into account when scheduling instrucitons

  • nak: normalize lop3 constant sources

  • nak: convert base to iadd for non-uniform ldcx lowering

  • nak: run nir_opt_constant_folding after nak_nir_lower_load_store

  • nir/opt_phi_precision: bail on load_const conversions between float and ints

  • rusticl: move the worker queue into the Platform

  • Reapply “rusticl: fix leak in `util_queue`”

  • rusticl/kernel: updated dim_threads in Kernel::suggest_local_size

  • rusticl/kernel: extract impl of suggest_local_size

  • rusticl/kernel: adjust grid at the end of suggest_local_size_impl

  • rusticl/kernel: add suggest_local_size tests

  • rusticl/kernel: add suggest_local_size_impl_gcd

  • rusticl/kernel: remove code to fill non full subgroups

  • rusticl/kernel: rework block size selection

  • gallium: remove PIPE_BARRIER_GLOBAL_BUFFER

  • rusticl/device: fix long vector_width queries on devices without int64 support

  • nak: implement shfr

  • nir: rework float compare late algebraic opts

  • nir: enable more opts for unordered and neo float compares

  • nak: implement and enable has_fneo_fcmpu

  • nak: implement ford and funord

  • nir/algebraic: pattern-match manual iadd64

  • brw: advertise fp64 fma on hw with fp64 support

  • anv: enable VK_KHR_shader_fma

  • anv: fix wrong rebase conflict resolution from VK_KHR_shader_fma MR

  • nak/sm20: fix immediate encoding for F2I and F2F

  • nak/hw_tests: add F2I test for NaN behavior

  • nir: use function foreach helpers inside nir_cleanup_functions

  • nir: add nir_shader_fully_linked helper

  • rusticl: return Result instead of Option from convert_spirv_to_nir

  • rusticl: abort compilation if the nir shader is not fully linked

  • gallium: add pipe_caps::hw_clear_buffer_sizes

  • rusticl/util: implement Debug for CLVec

  • rusticl/kernel: convert Queue parameter to Device in launch

  • rusticl/kernel: add interface to launch kernel with arguments without binding them

  • rusticl/kernel: add offset to buffer bindings

  • rusticl/meta: add builtin kernel support

  • rusticl/meta: add builtin kernels for buffer fills

  • rusticl/mem: make Image::fill return a closure

  • rusticl/mem: use meta for clEnqueueFillBuffer and clEnqueueSVMMemFill

  • rusticl/mem: implement 1Dbuffer fills on top of a plain buffer fill

  • asahi: update agx_get_cl_cts_version for submission 471

  • mesa_clc: support 32 bit targets

  • meson/rusticl: fix typo in depfile for builtin shaders

  • rusticl/meta: mark builtin kernels SPIR-V as 64 bit

  • rusticl/mesa: compile 32 bit version

  • rusticl/program: add Program::from_spirv_with_devs

  • rusticl/meta: support 32 bit devices

  • rusticl/meta: split out ulong kernels

  • rusticl/memory: return 0 for CL_IMAGE_SLICE_PITCH also for 2d images

  • rusticl/kernel: add libclc source hash to kernel shader keys

  • vtn/opencl: fix libclc needing fp16 lowering to fp32

  • nouveau: Fix return of dangling pointer in nouveau_fence_new

  • clc: make libclc optional for configs not needing it

  • clc: use our downstream fork of libclc

  • rusticl: warn if we load not our own libclc fork

Ken Cunningham (1):

  • llvmpipe: fix arch of LLVM JIT when cross compiling on Apple

Ken Xue (1):

  • radv: remove checking on the gralloc handle->numFds

Kenneth Graunke (113):

  • jay: Add missing ROR case

  • jay: Don’t forget UACCUM!

  • iris: Implement force_dual_color_blend_by_location via NIR

  • iris: Call elk_nir_lower_fs_outputs for Gen8 RT reads, not brw

  • nir: Set FRAG_RESULT_DUAL_SRC_BLEND in outputs_written when lowering

  • brw: Switch FS outputs to semantic IO and FRAG_RESULT_DUAL_SRC_BLEND

  • brw: Set prog_data::dual_src_blend from NIR outputs written bitfield

  • brw: Drop dead code from dispatch limit check for dual source blending

  • brw: Limit SIMD width based on NIR rather than first backend compile

  • nir: Allow bias for nir_texop_sparse_residency_intel

  • nir: Lower SSBO helper writes too

  • intel/nir: Only add an explicit LOD 0 when lod/bias don’t already exist

  • anv: Delete anv_instance::mesh_conv_prim_attrs_to_vert_attrs

  • anv: Use device->info.has_mesh_shading in key->mesh_input check

  • jay: Include depth and stencil on all MRT stores

  • jay: Add a TODO for coarse pixel shading

  • jay: Gripe more clearly about dual source blending

  • brw: Lower sample_pos for non-per-sample shaders in NIR

  • jay: Move render target store payload/descriptor construction to backend

  • jay: Implement fragment shader stencil writes

  • jay: Implement sample mask writes

  • jay: Add comments summarizing the PS thread payload layout

  • jay: Set Dispatch GRF Start Register in jay_setup_payload()

  • jay: Add a GPR_FROM_UGPRS opcode

  • jay: Implement sample position

  • jay, nir: Make a dispatch_mask_intel intrinsic

  • jay: Implement coverage mask

  • jay: Implement load_fs_config_intel

  • jay: Prohibit JAY_STRIDE_8 for EXPAND_QUAD

  • jay: Call constant folding before collecting FS outputs

  • jay: add a hack until we munge barycentrics dynamically

  • jay: Don’t skip sampler payload copies for 2 or fewer sources

  • anv: Drop TES dispatch mode asserts

  • brw: Fix URB read length for tessellation evaluation shaders

  • brw: Ensure entire input load fits in push data

  • brw: Refactor urb_read_length setting for TES

  • brw: Fix mistake in brw_nir_lower_deferred_urb_writes

  • brw: Fold constants after nir_lower_io for VS/GS/TES outputs

  • jay: Add URB load support

  • jay: Fix scratch surface address save/restore

  • jay: Remember sp_delta_B when rematerializing stack pointer lane 0

  • jay: Generalize EXTRACT_LAYER to take an arbitrary mask

  • jay: Handle facing that differs across subspans

  • jay: Implement viewport index FS input

  • jay: Fix null render target writes

  • jay: Drop render target stores with unconditional discards

  • jay: Implement dual color blending (but require SIMD16)

  • jay: Don’t swap FS interpolation .yz deltas

  • jay: Implement fragment shader barycentrics

  • jay: Pass proper simd_width to brw_nir_apply_key for fragment shaders

  • jay: Add tessellation evaluation shader support

  • jay: Add an INTEL_JAY=all option

  • jay: Unroll loops before lowering deferred URB writes

  • jay: Assert FS input deltas exist

  • jay: Fix hard coded number of FS inputs

  • jay: Ignore RT store condition if there are no outputs

  • jay: Improve unconditional discard removal

  • jay: Fix rewrite_without_flags for SEL with other flag sources

  • anv: Fix shader stats when using jay for non-compute stages

  • jay: Still predicate Null RT store if everything is discarded

  • jay: Implement load_subgroup_size

  • intel/nir: Improve address reuse in brw_nir_lower_immediate_offsets

  • intel/nir: Turn load_global_constant into load_global_intel too

  • jay: Call intel_nir_lower_shading_rate_output earlier

  • jay: implement load_frag_shading_rate

  • jay: Store the FS config def

  • jay: Store a test of the dynamic “is coarse?” FS config bit.

  • jay: Set prog_data->uses_fs_config when coarse pixel shading is dynamic

  • jay: Set the coarse pixel render target descriptor bit

  • jay: Don’t run the entire optimization loop before prog data

  • jay: Use nir_lower_frag_coord_to_pixel_coord

  • jay: Lower to pixel_coord_intel and frag_coord_w_rcp after prog data

  • jay: fix frag coord .z lowering with coarse pixel shading

  • jay: Allow BFN on U16 types

  • jay: Make a builder local in setup_fragment_payload

  • jay: Implement coarse pixel coordinate calculations

  • jay: Fix stack smashing with more than 16 FS inputs

  • jay: Rewrite FS output gathering

  • jay: Run brw_nir_lower_alpha_to_coverage earlier

  • jay: Optimize out noop samplemask writes

  • jay: Add missing HF conversion stride restrictions

  • jay: Tighten mixed stride restrictions

  • jay: Make a jay_clobbers_address_reg() helper

  • jay: Add a new VECTOR_EXTRACT opcode for indirect moves

  • jay: Implement indirect push constant loads for 32-bit

  • jay: Implement indirect push constant loads for 8/16-bit sizes

  • jay: Move brw_nir_apply_key call to be shared among stages

  • nir: Early out in nir_opt_shrink_vectors if all components are read

  • nir: Don’t shrink intrinsics and undefs to vec5s

  • jay: Speed up shuffles and vector extracts with uniform offsets

  • nir: Add an option for whether TCS invocation_id should be uniform

  • intel/compiler: Set nir_divergence_tcs_invocation_id_uniform

  • jay: Emit a noop URB write for EOT if there isn’t one to reuse

  • jay: Increase JAY_NUM_LAST_USE_BITS to 64

  • jay: Implement tessellation control shaders

  • jay: Use LOOP_ONCE if a loop ends in HALT too, not just BREAK

  • brw: Fix GS EOTs to not have an empty channel mask on LSC platforms

  • brw: Update comment that’s so old it makes no sense

  • brw: Inline brw_do_emit_fb_writes

  • brw: Assert that repclears aren’t used on Gfx12+

  • brw: Drop brw_compile_fs_params::allow_spilling

  • jay: Use INTEL_SIMD_DEBUG=cs for compute shaders, not fs

  • intel: Drop INTEL_DEBUG=no{8,16,32} flags

  • intel: Fix multipolygon flags in INTEL_SIMD_DEBUG default handling

  • intel: Refactor SIMD selection’s debug flag handling

  • intel: Replace INTEL_DEBUG=do32 with INTEL_SIMD_DEBUG

  • brw: Switch to INTEL_SIMD_FORCE for multipolygon modes

  • brw: Respect subgroup size requirements even with INTEL_SIMD_DEBUG

  • brw: Allow spilling and other poor decisions when using INTEL_SIMD_DEBUG

  • brw: Fix INTEL_SIMD_DEBUG=fs32 to work at all

  • brw: Rework FS SIMD selection to follow requirements over debug flags

  • jay: Add u16 and f16 support to CSEL

  • intel: Temporarily disable madvise on iris on xe.ko

Koch, Pawel (1):

  • Update docs regarding anv shader dumps Reviewed-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com>

Konstantin (3):

  • vulkan/cmd_queue: Handle struct copies that are not pointers

  • lavapipe: Re-emit push constants if the size changed

  • lavapipe: Fix push_constant_size for shader objects

Konstantin Seurer (64):

  • vulkan/radix_sort: Add support for 96-bit keys

  • vulkan: Rename radix_sort to radix_sort_u64

  • vulkan: Rename key_id_pair to key32_id_pair

  • vulkan: Implement 64-bit morton codes

  • radv/rt: Use 64-bit keys for gfx11-

  • util/u_trace: Add an option to emit additional code

  • util/u_trace: Rework resource management

  • util/u_trace: Print tracepoints with indentation

  • vulkan: Fixes for a spec update

  • vulkan,spirv: Update spec to 1.4.352

  • radv: Move debug options to radv_instance.h

  • radv: Move a whole bunch of debug/profiling related into a subdir

  • radv/tools: Rename radv_debug to radv_debug_hang

  • radv: Move radv_find_memory_index to radv_debug.c

  • radv: Add and use helpers for managing internal allocations

  • llvmpipe: Use i1 for sparce residency and expand as needed

  • llvmpipe: Implement sparse residency feedback for buffers

  • llvmpipe: Fix sparse binding large areas

  • lavapipe: Bump maxRayDispatchInvocationCount to the min requirement

  • nir: Duplicate the name in nir_def_set_name

  • llvmpipe: Remove lp_llvm_descriptor_base

  • lavapipe: Reduce descriptor sizes even further

  • lavapipe: Implement VK_KHR_shader_untyped_pointers

  • lavapipe: Add lvp_nir_lower_push_constants

  • llvmpipe: Fix memory leak when allocating sample functions

  • lavapipe: Perform shader object compatibility check early

  • lavapipe: Ignore src_plane for samplers

  • tools: Update imgui to the docking branch and add backends

  • meson: Add some include directories

  • vulkan/bvh: Add defines for acceleration structure types

  • radv: Use a separate BLAS pointer copy pass for BVH4

  • radv: Rename copy_blas_addrs to copy_addrs

  • radv: Rename serialization fields in radv_accel_struct_header

  • util: Add RTI file format definitions

  • radv/tools: Add RTI file dumping

  • rti: Initial commit

  • vulkan: Handle arbitrary build flag counts in vk_build_stage

  • vulkan: Make vk_build_stage non-static

  • vulkan: Filter for updates in vk_build_stage

  • vulkan: Move build_flags to vk_build_config

  • vulkan: Add vk_accel_struct_cmd_begin_debug_marker

  • vulkan: Use vk_build_stage for encode/update passes

  • vulkan: Shuffle around bvh build code

  • meson: Add mesa_python_path

  • vulkan: Move capture_key_pressed to vk_device

  • util/u_trace: Add an option for accumulating tracepoint ranges

  • util/u_trace: Do not generate empty structs

  • radv: Add u_trace support

  • util/u_trace: Include payload in the range accumulation key

  • radv: Include build_flags in the range key

  • util/u_trace: Release memory for reused timestamps

  • radv: Ignore entrypoints inside meta OPs for utrace

  • radv,anv: Enable BVH updates

  • tool: Rename RTI to gamma

  • radv: Fix generating ray history code

  • gamma: Set the window title to gamma

  • u_trace: Initialize fuzzy_* callbacks correctly

  • vulkan: Fix ROOT_FLAGS_OFFSET_ID offset

  • radv: Store root_flags for BVH8

  • radv: Use 64bit keys on GFX12

  • gamma: Do not render minimized viewports

  • gamma/radv: Display the ray launch ID

  • radv: Delay lowering printf

  • radv/bvh: Fix updating acceleration structures containing AABBs

Kovac, Krunoslav (3):

  • amd/vpelib: fix custom color space handling

  • amd/vpelib: Enable VPE cap for 3DLUT

  • amd/vpelib: Fixes for external lut compound

Lakshman Chandu Kondreddy (3):

  • zink: Query external memory handle type compatibility

  • freedreno: Add support for A704

  • zink: Set can_do_invalid_linear_modifier workaround for QCOM blob driver

Lars-Ivar Hesselberg Simonsen (9):

  • pan/genxml: Print shader hex in trace for Valhall

  • pan/va/disasm: Print 64 bit src/dest regs as reg pairs

  • pan/va/disasm: Align FAU printing

  • pan/va/disasm: Align indentation

  • panvk: Fix debug flag overlap

  • panvk/v10+: Align allocations >= 64k to 64k

  • panvk: Ensure 64k alignment for sparse images

  • pan/format: Prefer 16X16_BLOCK_U over INTERLEAVED_64K

  • panvk/v10+: Fix size gt -> gte for 64k alignment

Leandro Dorileo (1):

  • intel/executor: inform oa not available if that’s the case

Leder, Brendan Steve (Brendan) (1):

  • radeonsi/vpe: Update DCC API and programming

Lei Huang (1):

  • amd/virtio: enable Android amdgpu-virtio build option

Leon Perianu (1):

  • pvr: enable VK_EXT_device_memory_report

Lin, Ricky (1):

  • amd/vpelib: DPM detect first frame action

LingMan (6):

  • rusticl: Drop custom `addr` implementation

  • mr-label-maker: Add `Rust` label for `src/compiler/rust`

  • mr-label-maker: Drop rule applying `Rust` to all .rs files

  • mr-label-maker: Apply `Rust` label to `clippy.toml` and `build-rust.sh`

  • mr-label-maker: Apply `Rust` label to `rustfmt.toml`

  • mr-label-maker: Apply `Rust` label whenever crate dependencies are changed

Lionel Landwerlin (160):

  • intel/dev: fixup intel_needs_workaround() macro

  • anv: avoid C23

  • anv: fix compute push constant allocations on pre Gfx12.5 platforms

  • anv: fix invalid value for push block index

  • anv: fix debug printfs on hang

  • anv: fixup compute queue detection

  • anv: rework debug flag

  • anv: switch from INTEL_DEBUG to ANV_DEBUG for shader-print

  • anv: remove unused defines

  • anv: fix relocations into internal shaders

  • anv: simplify inline uniform descriptor loads

  • nir: expose nir_opt_dce_impl

  • anv: run a single impl loop for apply_pipeline_layout

  • anv/apply_layout: move some helpers around

  • ci/zink/intel: disable TGL demo-v2 trace

  • brw: track push constants shader stats

  • anv: promote push constant pointers to push buffers

  • anv: add a pass to realign global loads on DX CBV resources

  • intel/ci: update expectation for RPL

  • vulkan: add tracking for VK_EXT_primitive_restart_index

  • anv: implement VK_EXT_primitive_restart_index

  • anv: expose VK_KHR_shader_constant_data

  • anv: fix null pointer access

  • anv: stop using queue priority KHR aliases

  • anv: remove a bunch of KHR alias uses

  • anv/docs: update environment variable docs

  • anv: reorder debug options

  • anv: add a shader-dump debug option

  • imgui: update copy and port all tools using it

  • intel/tools: add eu stall viewer

  • brw/lower_texel_address: add heap support

  • anv: split sampler state packing from API object creation

  • brw: add heap support to brw_lower_storage_image

  • intel: add resource intrinsic support for heaps

  • anv: add lowering of descriptor heap intrinsics

  • anv: implement EXT_descriptor_heap entry points

  • anv: add descriptor heap binding support

  • docs: document ANV_DEBUG=desc-dirty

  • anv: enable EXT_descriptor_heap

  • anv: fix arc artifacts on Farming simulator 2022

  • anv: print out the content of the printf buffer at vkDestroyDevice

  • anv/brw/nir: fix wa_18019110168

  • anv: expose non binding-table/push-pointer flushing

  • anv: expose RT state flushing

  • anv: enable compute state flushing with indirect state

  • anv: move a bunch of structures to anv_types.h

  • anv/intel: add device generated commands shaders

  • anv/apply_layout: use the resource index to compute descriptor buffer addresses

  • anv: add apply_layout support for device bindable shaders/pipelines

  • anv: add a helper to flush the descriptors for indirect compute execution

  • anv: program relative push set offset for descriptor buffers device bindable shaders

  • vulkan: add pipeline helper to retrieve scratch-size/ray-queries

  • anv: add support for indirect execution set

  • anv: add indirect command layout support

  • anv: add unspecified internal kernel send count support

  • anv: allow simple shader spilling for complex ones

  • anv: enable generation shader calls

  • anv: handle descriptor binding with DGC

  • anv: implement generated preprocess & execute

  • anv: add barrier flags handling for preprocess buffers

  • anv: handle preprocess buffer creation on <= Gfx12.0

  • anv: track generated commands work with perfetto

  • anv: expose VK_EXT_device_generated_commands by default on Gfx12.5+

  • anv: add a device generated command debug option

  • anv: add Gfx9 support VK_EXT_device_generated_commands

  • anv: expose VK_KHR_maintenance11

  • anv: group all performance drirc together

  • anv/iris: stop using 3DSTATE_PUSH_CONSTANT_ALLOC_PS on Gfx12.5

  • anv: rename push constant allocation helper

  • anv: add an option to disable push constant space reallocation

  • vulkan/runtime: fix invalid address flags value for CmdCopyBufferToImage2

  • anv: fixup null address check

  • anv: implement VK_KHR_device_address_commands

  • anv: remove old entrypoints

  • anv: implement missing device image property compression filtering

  • vulkan/wsi: write VkImageCompressionControlEXT from swapchain to image creation

  • anv: enable VK_EXT_swapchain_compression_control when possible

  • blorp: stop requesting the fp64 shader for ELK

  • blorp: only request fp64 shader on when required

  • anv: sweep the NIR fp64 shader before keeping it on the device

  • anv: only load fp64 software shader when needed

  • anv: add an option to disable allocation over subscription

  • brw: simplify VF component packing code

  • anv: add SIMD32 requirement heuristic for Dragon Dogma 2

  • brw/jay: move some coarse lowering to NIR

  • brw/jay: move sample_mask_in handling to NIR

  • docs/features: updates for Anv

  • anv: temporarily reenable scratch page by default

  • anv: bump max compute workgroup count

  • anv: further optimize dirty state after secondary emission

  • anv: only reprogram line-stipple if enabled

  • util: add a script to auto-generate a drirc infrascture per driver

  • util/drirc_gen: enable validation for a specific driver

  • hasvk: rename a couple of drirc options

  • hasvk: add a driver section for drirc

  • drirc: remove non Anv option in the Anv section

  • anv: use the new generation script for drirc

  • anv: fix missing bindless flag hashing

  • anv: fix render target remapping tracking at the beginning of render passes

  • brw: avoid requiring a valid render target for empty fragment shaders

  • spirv: fixup infinite recursion with shader replacement

  • anv: use shader source hash rather than cmd_buffer fields

  • intel: switch shader hash to 64bit value

  • mi_builder: mi_umax2 tests

  • anv: rename drirc script

  • anv: move fake_sparse drirc to feature category

  • anv: move compression control drirc to feature section

  • anv: fake VK_EXT_image_compression_control on Xe2+

  • iris: only call brw_nir_fs_needs_null_rt() with no render targets

  • brw: fix null render target decision

  • anv: fix assert/crash in import of compressed local memory on xe2+

  • anv: align storage texel buffer support on image support

  • anv: add missing condition to update 3DSTATE_RASTER

  • anv: fix 3DSTATE_SF line width programming with Bresenham lines

  • anv: don’t forget dataport flush for ANV_DEBUG=dgc-dump

  • elk: assert always/never on some of the FS config flags

  • brw: remove always true condition

  • brw: remove interpolator coarse bit setting

  • brw/jay: track usage of fs_config by backend

  • brw: only check for shader_info::fs.uses_sample_shading

  • anv/brw/jay: de-dynamify per-sample interpolation

  • anv: hash binding tables for EXT_descriptor_heap too

  • brw: add shader key to enable robust SLM accesses

  • anv: add hitman2 workaround for SLM load vectorization

  • anv: fix descriptor heap indexing of YCbCr embedded samplers

  • vulkan/runtime: fixup group building with shaders from libraries

  • spirv: add parsing of vkd3d-proton shader hashes

  • drirc: add a callback mechanism do deal with shader hash & options

  • anv: enable VK_EXT_descriptor_heap by default

  • vulkan: condition cmd_queue initialization to driver need

  • anv: fix push constant address emission for gfx commands

  • anv: add missing handling of push pointers in gfx dgc

  • anv: use vkd3d-proton provided shader hashes if available

  • anv: add infrastructure to deal with missing barriers in applications

  • anv/brw: fixup 64bit array image accesses

  • anv: fix 64bit image atomic emulation with EXT_descriptor_heap

  • anv: more dgc push constant fix

  • anv: fix push pointer optimization with DGC

  • anv: add memory heap budget tracking across VkInstance

  • anv/brw: limit push constant promotion in vertex shaders

  • anv: introduce an option to disable disk cache

  • anv: fixup RT building barrier

  • anv: fix compression control reporting on xe2+

  • anv/ci: turn on astc emulation testing

  • anv: fill min_array_element with indirect descriptors

  • anv: fixup max push data delivered to shaders

  • anv: add workaround for atomics on R11G11B10 images

  • anv: remove previous Horizon Forbidden West workaround

  • anv: flush accumulated barriers for top of TOP_OF_PIPE

  • anv: fix Wa_18040903259

  • vulkan/runtime: fixup vk_shader leak on RT group recompile

  • anv: fix push buffer descriptor address relocation

  • brw: add missing INTEL_FS_CONFIG_PER_PRIMITIVE_REMAPPING handling

  • brw: fix wa_18019110168 lowering

  • anv: fix leak in RT binding point

  • anv: fix barrier for Wa_1508744258 / Wa_14024015672

  • iris: fix barrier for Wa_1508744258 / Wa_14024015672

  • jay: copy resource_intel surface handle value

  • anv: fixup the logic dealing with STATE_BYTE_STRIDE

  • anv: only consider active view-capable queues for image views

Lishin (3):

  • mesa/st: relax shader_has_one_variant checks for GLES2

  • broadcom/qpu: add V3D 7.1 disasm tests

  • v3d/v3dv: use common compute limits

Liu, Mengyang (2):

  • aco: fix broken VGPRs reservation for 64-bit attributes in VS prologs

  • amd: disable reset_filter_cam for mec

Lone_Wolf (2):

  • ac/llvm: fix build with LLVM 23 (MCSubtargetInfo)

  • clc: fix build with LLVM23 (TargetRegistry::lookupTarget)

Lorenzo Rossi (104):

  • nir: Extract float_is_half tests in common code

  • nir/opt_algebraic: optimize fadd/fmul with 16-bit source and constant

  • pan/compiler: Allow 16-bit alpha for atest_pan

  • pan/compiler: Fix WaRaR hazard in pressure scheduler

  • pan/compiler: Lower unaligned scratch memory accesses

  • pan/compiler: Handle ssbo_atomics in lower_vs_atomics

  • nir/lower_point_size: Handle 16-bit point sizes

  • nir/opt_sink: Add pan-specific load_input

  • panfrost: Constant-fold io locations after lowering

  • pan/compiler: Sort preprocess

  • panvk/jm: Fix tls_size overwrite in indirect draws

  • pan/compiler: Rework scratch memory strategy

  • panvk,panfrost: Pass inputs and info to postprocess

  • panvk: Remove pan_optimize_nir call

  • pan/compiler: Rename bifrost_optimize_nir

  • pan/compiler: Sort postprocess

  • pan/compiler: Collect nopersp varyings in lower_noperspective_fs

  • pan/compiler: Add better documentation for second lower_int64

  • pan/compiler/lower_fs_inputs: Do not trust slot->alu_type

  • panfrost: Split default key creation in helper function

  • panfrost: Plumb VS varying_layout in FS

  • pan/bi: Vectorize f2f16 on v10 and earlier

  • pan/bi: Switch old license texts to SPDX

  • pan/mid/fuse_io_cvt: Disable fusion on highp

  • pan/bi: Add nir_fuse_io pass

  • pan/bi: Add a printing helper for pan_varying_layout

  • pan/valhall: fuse_cmp skip when fusing the same instruction

  • pan/bifrost: Make CSE independent of liveliness labels

  • nir/opt_algebraic: Optimize mediump fadd/fmul done in highp

  • pan/bifrost: Fix 16-bit demote_if

  • kraid: Fix out-of-tree build issue

  • kraid/tests: Edit meson to help rust-analyzer provide IDE suggetsions

  • panfrost: Separate the compiler from libpanfrost

  • panfrost: Reorder meson definitions

  • compiler/rust/lower_bounded: Add FromIterator impl

  • kraid/swizzle: Add fold_u64

  • kraid/ir: Add SrcMod::fold_u64

  • kraid/ir: Add FauRef UserPage creation utilities

  • kraid: Add alloc_vec utility

  • kraid: Fix FauRef Display bug

  • kraid/ops: Add a small crate documentation for conventions

  • kraid: Add OpIMul

  • kraid/model: Ensure dyn Model is Send + Sync

  • kraid: Add hw_runner

  • kraid: Add basic hw_tests

  • kraid: Add Foldable and initial tests for OpShiftLop

  • kraid/hw_runner: Unmap buffers on drop

  • kraid: Sort opcodes and keep them sorted

  • kraid: Replace alloc_vec with alloc_ref

  • compiler/rust/float16: Implement total_cmp as present in f32

  • kraid/encode_v9: Implement accumulator ops for FCmp

  • kraid/ops: Fix Display for OpFCmp

  • kraid: Add tests for OpFCmp

  • kraid: Add Foldable impl for OpCSel

  • kraid/encode_v9: Implement accumulator ops in ICmp

  • kraid: Add tests for OpICmp

  • kraid: Add tests for OpIAdd

  • kraid: Add tests for OpIMul

  • kraid: Add OpISub with tests

  • kraid: Add OpClz with tests

  • kraid: Add OpIToF32

  • kraid/nir: Add IMul

  • kraid: Add ufind_msb

  • kraid: Add ShaderInfo

  • kraid: Add register preloading

  • kraid: Add load_push_constant

  • kraid/hw_tests: Use preloaded registers

  • pan,nir: Add Panfrost image intrinsics

  • pan/bi: Lower image load/store/lea in NIR and fix OOB access

  • pan/bi: Remove unused backend lowering

  • panfrost/model: Add var,cvt,sfu rates for Valhall architectures

  • panfrost/compiler: Properly compute Valhall ALU bound

  • drm-shim.py: Add more panfrost models

  • panfrost/drm-shim: Fix assertion on v12+

  • kraid: Legalize immediates

  • kraid: Add OpIAbs and plumb it through

  • kraid/nir: Fix nir_op_extract*16

  • kraid: Add BitRev and wire it up

  • kraid: Add OpPopCount and wire it up

  • kraid: Add support for clamp and round to OpFAdd

  • pan/bi: Add kraid-specific algebraic rules

  • kraid: Add F32ToI32 and wire it up

  • kraid/algebraic: Lower b2i conversions

  • kraid/algebraic: Lower nir_op_pack_uvecX_to_uint

  • kraid: Wire up fneg

  • kraid: Add OpFma and wire it up

  • kraid: Add OpFlush

  • kraid: Add FRound and wire it up

  • kraid: Add Frexp and wire it up

  • kraid: Add OpFMin/OpFMax and wire them up

  • kraid: Add OpLdExp and wire it up

  • kraid: Add a pass macro for validation and debug

  • kraid: dead-code elimination pass

  • pan/bi: Fix f2f16(a@16) in shader-db run on v13

  • pan/bi: Move pan_nir_fuse_io in bi_optimize_late

  • pan/nir_fuse_io_cvt: Add texture cvt fusion

  • pan/compiler: Don’t widen unaligned push-constants too much

  • panfrost: Unify FAU constants and relocation handling

  • panfrost: Promote constants to FAU

  • panfrost/midgard: Fix fau max not initialized

  • panfrost: Fix wrong layout reuse in user clip planes

  • pan/nir: Fix header static inline function

  • pan/compiler/stats: Fix ALU not being used in instruction bounds

  • pan/bi: Propagate swizzle in bi_optimizer_result_type

Louis Montagne (2):

  • zink: relax build-id length assertion for Mach-O

  • meson: allow DRI on darwin to enable Zink + EGL builds

Loïc Molinari (17):

  • pan/crc: Restrict CRC buffer creation to 1st RT mipmap level

  • pan/crc: Introduce pan_fb_info_is_fully_covered()

  • pan/crc: Check AFBC renderblock size on v5 and v6 too

  • pan/crc: Check CRC buffer validity and coverage on v5 and v6 too

  • pan/crc: Simplify CRC buffer selection logic

  • pan/crc: Use RT selection loop in single RT case

  • pan/crc: Check CRC requirements in dedicated function

  • pan/crc: Disallow CRC on sparse AFBC images

  • pan/crc: Simplify CRC buffer initialization

  • pan/crc: Cache temporary CRC info

  • pan/crc: allow setting a NULL pointer to the CRC validity state

  • pan/crc: Store CRC state in a struct

  • pan/crc: Enable CRC for multiple RTs on v6

  • pan/crc: Disable CRC on v4

  • pan/crc: Enable Empty Tile Elimination

  • pan/crc: Optimize clear color hashing

  • panfrost/ci: Mark “spec@!opengl 1.4@copy-pixels” as flake

Lucas Francisco Fryzek (1):

  • util/u_trace: Don’t use empty initializer list

Lucas Fryzek (1):

  • Modify x11_xcb_display_supports_xshm to get xshm opcode

Lucas Stach (4):

  • etnaviv: clean up index buffer handling code a bit

  • etnaviv: move index buffer handling in draw_vbo after derived state handling

  • etnaviv: reserve state emission space early in draw_vbo

  • etnaviv: move sampler source update before draw space reservation

Luigi Santivetti (3):

  • pvr: de-dup strncmp in pvrsrvkm winsys

  • pvr: add missing multi-arch support for pipeline exec and stats

  • pvr: re-use texture state words for each load op

Lukas Zapolskas (2):

  • pan/pps: Move PanfrostDevice to a separate file

  • pps: Add the Primitive, Instruction, Pixel and Fragment unit types

Maaz Mombasawala (3):

  • Revert “ci: vmware farm is offline, stop using it”

  • Revert “ci-farms/vmware: Disable vmware tests for now”

  • svga: Update CI expectations.

Marc Alcala Prieto (41):

  • pan/genxml: Add performance-trilinear enum values

  • pan/genxml: Add missing enum values on v9-v13

  • pan/genxml: Add v14 definition

  • pan/genxml: Implement RUN_FRAGMENT2

  • pan/decode: Remove progress-related decoding logic

  • pan/genxml: Build libpanfrost_decode for v14

  • pan/clc: Build for v14

  • pan/fb: Implement pan_emit_fb_desc for v14+

  • pan/desc: Implement pan_emit_fbd for v14+

  • pan/texture: Add v14+ YUV pipe format mappings

  • pan/format: Add v14+ YUV pipe format mappings

  • pan/afbc: Add v14+ AFBC YUV compression mappings

  • pan/afrc: Add v14+ AFRC YUV compression mappings

  • pan/lib: Build for v14

  • panvk: Implement RUN_FRAGMENT2

  • panvk: Handle provoking vertex and simultaneous reuse on v14

  • panvk: Build for v14

  • pan: Add v14 support

  • pan/va: Fix packing test for LdVarBufImmF16 on v11

  • pan/bi,va: Use dedicated LD_VAR_BUF_FLAT* opcodes on v14+

  • panfrost: Implement RUN_FRAGMENT2 on the Gallium driver

  • panfrost: Build the Gallium driver for v14

  • panfrost: Advertize Mali-G1-Pro support

  • docs/panfrost: Advertize Mali-G1-Pro support

  • pan/decode: Support INTERLEAVED_64K Z/S target dumps

  • pan: Layer offset is not longer available starting on v14

  • panvk/csf: Allow 256 layers per tiler descriptor on v14+

  • panfrost: Advertise Mali-G1-Premium and Mali-G1-Ultra support

  • panfrost: Remove duplicated flushes before RUN_FRAGMENT[2]

  • panvk/csf: Emit fragment layer state just before RUN_FRAGMENT2

  • panvk/csf: Implement incremental rendering on v14+

  • pan/csf: Fix incremental rendering on v14+

  • pan/bi: Load vertex view index from preload on v14+

  • panvk: Fix multiview support on v14+

  • pan: Add helper for max multiview view count and rise it to 16 on v14+

  • pan/compiler: Rename multiview to per_view_outputs

  • pan/va: Fix serialization of atomic operations using BI_ATOM_OPC_AUMIN

  • pan/va: Unit test BI_ATOM_OPC_AUMIN

  • pan/ci: Remove GLES shader image load/store atomic flake

  • panvk: Fix DRM format modifiers for multi-planar YUV formats

  • panvk/csf: Avoid poisoning read-only fragment SRs

Marek Olšák (167):

  • nir: add back color0/1 system values and VARYING_SLOT_PARAM_GEN_AMD

  • ac/nir: add ac_nir_get_io_driver_location as replacement for IO bases

  • ac,radeonsi: don’t use nir_intrinsic_base for FS outputs

  • radeonsi: don’t recompute IO bases for FS outputs

  • radeonsi: stop setting si_shader_info::output_semantic for FS

  • radeonsi: stop using si_shader_info::output_semantic for passthrough TCS

  • radeonsi: stop using output_semantic[] for LS outputs passed via VGPRs

  • radeonsi: remove si_shader_info::output_semantic[]

  • radeonsi: remove si_shader_info::num_outputs

  • ac,radeonsi: stop using nir_intrinsic_base for TCS inputs passed via VGPRs

  • ac/llvm: correctly load 16-bit TCS inputs from VGPRs and simplify

  • ac/llvm: reorder/remove variables in visit_load_input

  • radeonsi: update shader info in si_nir_lower_color_flatshade_twoside

  • ac,radv,radeonsi: don’t use nir_intrinsic_base for FS inputs

  • radv: remove radv_recompute_fs_input_bases

  • radeonsi: compute si_shader_info::color_attr_index without input_semantic[]

  • radeonsi: compute si_shader_info::inputs_read without input_semantic[]

  • radeonsi: remove si_shader_info::input_semantic[]

  • radeonsi: don’t call nir_recompute_io_bases for FS

  • radeonsi: set num_vs_inputs from nir->num_inputs and use it more

  • radeonsi: just get si_shader_info::num_inputs from NIR

  • amd: remove unnecessary and transitive #includes

  • ac/nir: add ac_nir_assign_fs_input_locations to set PS input locations in stone

  • nir/opt_licm: add a private state structure for the pass

  • nir/opt_licm: use nir_metadata_control_flow

  • nir/opt_licm: hoist instructions across multiple levels of nested loops

  • radeonsi/ci: remove the fixed XFB test from fails/flakes

  • radeonsi/ci/build: also fetch video decode/encode sample for VK CTS

  • nir/opt_dce: factor out dead instruction removal into a helper

  • nir/opt_dce: add shader_info::assert_inputs_not_dead

  • ac/nir: factor out ac_nir_lower_tex_coords from ac_nir_lower_image_tex

  • ac/nir/lower_tex_coords: move input loads instead of cloning them

  • aco/tests: update ACO tests for ac_nir_lower_tex_coords refactoring

  • glsl,gallium: add pipe_caps::glsl_bindless_handles_are_32bit

  • radeonsi: set glsl_bindless_handles_are_32bit

  • nir: add frag_coord_xy

  • nir/lower_wpos_ytransform: handle frag_coord_xy

  • nir/opt_frag_coord_to_pixel_coord: handle frag_coord_xy

  • nir: add direct lowered frag_coord building to replace lowering passes

  • nir: use nir_build_frag_coord everywhere

  • amd: add a tool that prints tiling layouts for all shim devices

  • winsys/amdgpu: revert invalid changes from CS functions

  • winsys/amdgpu: fix memory leaks when amdgpu_cs_create fails

  • amd/tools: rewrite ac_print_tiling_layouts to print all layouts, including XORs

  • nir/opt_licm: add filter callback

  • nir/tests: add nir_opt_licm tests

  • nir/licm: allow speculative hoisting across terminate if the filter is set

  • nir: add an option to ignore INTERP_MODE_NONE in nir_shader_gather_info

  • radeonsi: fix a typo in si_shader_update_spi_shader_formats

  • ac,radeonsi: add a helper to print PS input VGPR layout

  • ac,radeonsi: add helpers to print SPI_SHADER_COL/Z_FORMAT

  • aco,radeonsi: use enums for color barycentrics instead of input VGPR indices

  • radeonsi: remove dead get_frag_coord_from_pixel_coord optimization

  • radeonsi: use shader_info::fs::uses_sample_shading for ac_nir_lower_ps_early

  • ac: add ac_shader_args::line_stipple_tex_ena

  • radeonsi: move SI_SPI_PS_INPUT_ADDR_FOR_PROLOG into a helper function

  • aco,radeonsi: don’t forward LINE_STIPPLE_TEX_ENA VGPR from the PS prolog

  • aco,radeonsi: declare prolog CENTROID VGPRs only if used

  • radeonsi: declare prolog ANCILLARY & SAMPLE_COVERAGE VGPRs only if used

  • radeonsi: declare prolog LINE_STIPPLE_TEX_ENA VGPR only if needed

  • radeonsi: declare prolog LINEAR_SAMPLE/CENTER VGPRs only if used

  • radeonsi: simplify get_interp_info_from_input_load

  • radeonsi: stop using TGSI definitions for interpolation

  • radeonsi: handle any size of shader args in the LLVM PS prolog

  • nir/opt_move_to_top: add an option to exclude moving at_offset/at_sample loads

  • nir: generalize nir_vertex_divergence_analysis -> nir_custom_divergence_analysis

  • nir/opt_varyings: use workgroup divergence to identify convergent mesh outputs

  • util/set: add helper _mesa_set_equal

  • nir/opt_varyings: rewrite elimination of duplicated outputs

  • nir/tests: don’t leave “namespace {” unclosed in nir_opt_varyings_tests.h

  • nir/tests: test new output deduplication cases

  • nir/opt_varyings: always report progress when calling nir_remove_varying

  • nir/tests: use ASSERT_EQ instead of ASSERT_TRUE in nir_opt_varyings tests

  • nir: add missing SYSTEM_VALUE_FRAG_COORD_W_RCP

  • nir: handle load_frag_coord_w_rcp in multiple passes, same as non-rcp

  • nir: change nir_frag_coord_form options to a bitmask

  • nir: add nir_frag_coord_use_pixel_coord for OpenGL

  • radv: switch to nir_frag_coord_xy_z_w_separate with w_rcp

  • radeonsi: switch to nir_frag_coord_xy_z_w_separate with w_rcp

  • radeonsi: enable nir_frag_coord_use_pixel_coord

  • radeonsi: don’t treat sample_pos as using frag_coord

  • ac,radeonsi: remove all frag_coord_xy code

  • nir/opt_frag_coord_to_pixel_coord: factor out helper nir_all_uses_of_float_are_integer

  • radv: add a pass that selects either frag_coord_xy or pixel_coord, but not both

  • radv: remove dead load_sample_pos code

  • radv: move SPI_PS_INPUT_ENA emission into radv_emit_ps_state

  • radv: select frag_coord_xy and pixel_coord conditionally based on dynamic state

  • nir/opt_algebraic: add more ffract/ffloor/ftrunc/f2u/f2i patterns

  • radeonsi/tests: add an ordered append bandwidth test

  • radv: ignore color attachment samples for ps_iter_samples

  • ac/nir/lower_ps_early: remove obsolete comment

  • ac/nir/lower_ps_early: assume frag_coord_is_center is always true

  • ac/nir: add a new pass ac_nir_lower_sample_mask_in

  • radv: switch to ac_nir_lower_sample_mask_in

  • radv: enable SAMPLE_COVERAGE PS VGPR dynamically

  • radv: fix an inefficiency where the ANCILLARY PS VGPR was enabled but unused

  • radeonsi: use ac_nir_lower_sample_mask_in

  • ac/nir/lower_ps_early: remove now-unused lowering of sample_mask_in

  • radv: make RAST_SAMPLES_STATE dirty in CmdBeginRendering only on gfx12+

  • radv: emit_rast_samples_state uses uses_vrs_attachment only on gfx11+

  • radv: don’t leave SPI_PS_INPUT_ENA uninitialized with NULL PS to fix a hang

  • ac/surface: print the modifier in ac_surface_print_info

  • ac: add basic HTILE dword printing

  • radv: bump the sparse alignment requirement to 64K

  • radv: fix VK_MEMORY_PROPERTY_DEVICE_COHERENT_BIT_AMD with sparse buffers

  • radv,radeonsi: disallow VRS flat shading if SubgroupInvocationID is used

  • radv: rename vrs_coarse_shading -> vrs_flat_shading

  • nir/opt_idiv_const: a / uint_max -> b2i(a == uint_max)

  • radv: stop using set_sh_reg_idx(3) to reduce CP overhead

  • radv: fix setting COMPUTE_DISPATCH_INTERLEAVE on the gfx queue

  • radeonsi: remove unnecessary and indirect #includes

  • radeonsi/ci: allow glcts to be in the cts directory

  • radeonsi/ci: change DEQP_TARGET to default for Wayland

  • ac,radeonsi: remove uses_kernel_cu_mask and associated code

  • nir/opt_varyings: rewrite indirect IO tracking and dead IO elimination

  • nir/opt_varyings: split tidy_up_indirect_varyings

  • nir/opt_varyings: shrink pathological varying arrays to 1 element

  • nir: fix ibfe handling in ssa_def_bits_used

  • nir: extend ssa_def_bits_used to allow getting bits for any src component

  • nir: add a comp parameter into nir_def_bits_used

  • nir: change nir_all_uses_of_float_are_integer to return type masks and bits used

  • nir: change nir_def_bits_used to accept nir_scalar

  • nir/opt_idiv_const: a / b (where b > uint_max / 2) -> b2i(a >= b)

  • radv/lower_opt_fs_frag_pos: optimize f2u32(frag_coord_xy) & 0x1 to subgroup ops

  • radv: use quad_pos for pixel_coord conditionally based on dynamic state

  • radv: disable AMD_device_coherent_memory on gfx12 due to out of order behavior

  • ac/nir: fix incorrect upper bound for view_index

  • radv: lower view_index to a user SGPR instead of layer_id

  • radv: use PKT3_SET_SH_REG_PAIRS for setting multiple view_index SGPRs on gfx12

  • nir/tests: restructure opt_varyings_tests_bicm_sysval

  • nir/opt_varyings: move (c ? interp_input0 : interp_input1) into the prev shader

  • nir: add shader_info::sample_mask_in_declared because it has side effects

  • radv: fix a rare crash with NULL PS and force_vrs_per_vertex

  • radv: use radeon_opt_set_context_reg for PA_CL_VRS_CNTL to fix random behavior

  • radv: use VRS flat shading even if other VRS state is enabled

  • radv: remove the VRS rate output if VRS flat shading overrides it

  • radv: cosmetic VRS changes

  • radv: don’t use PS_ITER_SAMPLE to force VRS 1x1, use SC/DB VRS override instead

  • radv: remove the VRS rate output if VRS is force-disabled by FS

  • radv: don’t execute pre-rast shader info code for FS

  • radv: remove no-op code from radv_consider_force_vrs for POPS

  • radv: disallow force_vrs_per_vertex with FragCoord when using GPL & ESO

  • radv: move force_vrs_per_vertex to emit_fsr_state to make it robust (rewrite)

  • radv: remove redundant PA_CL_VRS_CNTL setting from the initial state

  • radv: reduce duplication in gfx103_emit_vrs_override_state

  • radv: disable the VRS image on gfx11.x if the VRS rate is overridden

  • radv: fold gfx103_pipeline_vrs_flat_shading into its only use

  • radv: always set EN_VRS_RATE=1 because GE_VRS_RATE can also disable it

  • radv: inline radv_is_vrs_enabled

  • radv: use shader_info::fs::sample_mask_in_declared

  • radv: ignore VRS state for sample_mask_in lowering and optimizations

  • radv: take sample_mask_in_declared into account when lowering to pixel_coord

  • radv: set key.ps.force_vrs_enabled and key.vrs_may_be_enabled more accurately

  • bin/drm-shim: forward the error code from the command to the user

  • radv: fix low pixel throughput with NULL DS on GFX11.x

  • radv: don’t set DB_Z_INFO.NUM_SAMPLES = 3 on gfx12

  • radv/nir_trim_fs_color_exports: use a state structure to pass parameters

  • radv/nir_trim_fs_color_exports: remove mrt0.w if alpha_to_one makes it dead

  • ac: fix a GPU hang with LLVM due to incorrect VGPRS decoding of LLVM output

  • ac/llvm: rename ac_parse_shader_binary_config -> ac_parse_llvm_binary_config

  • radeonsi: use wave64_vgpr_encode_granularity

  • radv: don’t expose memory types from AMD_device_coherent_memory without the ext

  • radv: fix determining the raster prim for guardband

  • radv: fix determining the raster prim for line mode

  • radv: fix determining the dynamic raster prim for FS barycentrics

  • radv: fix determining the static raster prim for FS barycentrics and front_face

  • nir/opt_varyings: fix incorrect counting of emit_vertex within a block

Mario Kleiner (22):

  • wsi/display: Expose VK_FORMAT_B8G8R8A8_UNORM before VK_FORMAT_B8G8R8A8_SRGB

  • wsi/display: Improve connector->last_nsec timestamping.

  • wsi/display: Add workaround for all-zero valued pageflip events.

  • wsi/display: Deal with vblank-less systems for VK_EXT_present_timing.

  • wsi/common: Small compliance fixes for VK_EXT_present_timing.

  • wsi/common: Allow VK_EXT_present_timing present without presentStageQueries.

  • wsi/common: Allow to return queue_done_time in host time domain.

  • wsi/wayland: Unconditionally assign present_timing.time_domain.

  • wsi/common: Add VK_GOOGLE_display_timing support for KHR_display.

  • wsi: Don’t try to create a timestamp query pool without driver support.

  • vulkan/wsi: Optionally expose VK_GOOGLE_display_timing on wsi wayland+x11.

  • wsi/wayland: Always use clock monotonic domain for GOOGLE_display_timing.

  • vulkan/wsi: Add hk, nvk as VK_GOOGLE_display_timing supported drivers.

  • docs/features: Add missing VK_EXT_present_timing enabled for X11.

  • wsi/display: Actually fix vblank-less systems for VK_EXT_present_timing.

  • pvr: Expose VK_KHR_present_id and VK_KHR_present_wait.

  • hasvk: Expose VK_KHR_present_id2 and VK_KHR_present_wait2.

  • hasvk: Expose VK_KHR_calibrated_timestamps.

  • hasvk: Expose VK_EXT_present_timing and VK_GOOGLE_display_timing.

  • wsi/display: Don’t update connector last_frame/nsec in vkGetSwapchainCounterEXT.

  • wsi/x11: Skip next_present_ust_lower_bound assignment in certain FRR mode.

  • wsi/x11: Refine VRR vs. FRR detection a bit.

Martin Roukala (né Peres) (18):

  • zink/ci: mark blender-demo-cube_diorama as flaky on gfx1201

  • turnip/ci: document recent flakes

  • ci: disable the valve-kws farm

  • Revert “ci: disable the valve-kws farm”

  • freedreno/ci: reduce the parallelism of the a750-vk job

  • freedreno/ci: document more failures for the a750-gl-cl job

  • radv/ci: reduce parallelism for radv-gfx1201-vkcts

  • radv/ci: document more flakes

  • radeonsi/ci: document new flakes

  • amd/ci: tighten the timeouts of the Valve jobs

  • zink/ci: document a recent regression on navi10

  • zink/ci: document more flakes

  • zink/ci: tighten the timeouts of the valve jobs

  • nvk/ci: tighten the timeouts of the valve jobs

  • radv/ci: bump the timeout of the valve vkd3d-asan jobs

  • ci: allow controlling which hw test jobs to create at pipeline creation

  • panfrost/ci: turn bifrost / valhall rules into per-kernel driver

  • radv/ci: document another WSI flake in radv-renoir-vkcts-full

Mary Guillemard (45):

  • nvk: Use SET_REFERENCE in nvk_CmdResetQueryPool

  • nvk: use MME shadow RAM in nvk_meta begin/end

  • nvk: Move nv_push closer to their uses in nvk_cmd_begin_end_query

  • nvk: Clear counters at the begin of a query

  • nvk: Remove delta handling from query pool

  • nvk: Conditionally enable counters when needed

  • nvk: Move report offset to reports_start for nvk_CmdCopyQueryPoolResults

  • nvk: Handle zero queries in CmdCopyQueryPoolResults and CmdResetQueryPool

  • nvk: Store available and timestamps packed together

  • nir/lower_bit_size: Preserve float controls when lowering alu ops

  • nvk: Handle foreign queue dependencies

  • nvk: Handle host accesses barrier

  • nvk: Multiply by local_size for CS invocations in DGC codepath

  • nak: Allow YY swizzle for SM20 and SM32 asserts

  • nir/nir_format_convert: Add missing u2f32 in nir_format_unpack_r9g9b9e5

  • nir,nak: Add match_any_nv

  • nak: Add a lowering pass for shared memory atomics in mesh stages

  • nvk: Prepare nvk_shader for GS header upload for mesh shaders

  • nvk: Prepare cbuf for mesh shader support

  • nvk: Add support for mesh and task shader binding

  • nvk: Implement mesh draw commands

  • nak: Implement mesh and task shader stages

  • nvk: Do not set lower_cs_local_index_to_id

  • nvk: Only lower shared memory for compute shaders

  • nvk: Lower mesh and task shaders

  • nvk: Advertises VK_EXT_mesh_shader

  • docs/nvk: Add some notes about mesh shading and ISBE layout

  • nvk: Do not report task and mesh stages as supported on pre-Turing

  • nvk/nvkmd: Do not merge bind operations across VA mappings

  • nvk: Implement support for non graphics timestamp

  • nouveau/mme: Add some simple MME shadow RAM dumper

  • nouveau/mme: Add a test for MME Shadow RAM behavior

  • nvk: Increase maxStorageBufferRange and maxBufferSize

  • nvk: Default to output primitives as lines for tesselation parameters

  • nvk/ci: Update expectations and document failures

  • nvk: Use I2M in CmdUpdateBuffer when possible

  • panvk: Split cmd_prepare_push_uniforms logic

  • drm-shim/nouveau: Report proper values in DRM_NOUVEAU_GET_ZCULL_INFO

  • drm-shim/nouveau: Stop using nouveau gallium names for classes

  • drm-shim/nouveau: Add Ada A to Blackwell B support

  • nvk: add a build option to override the build ID

  • nvk: Only increment CS counters when query is active in CmdDispatchBase

  • nvk: Do not enable remap in nvk_copy_indirect

  • nvk: Do not take base into account when lowering emulated attributes

  • nvk: Reenable compression support on Turing with nouveau 1.4.3

Matt Turner (15):

  • intel/elk: Remove some dead code

  • intel/elk: Remove dead TXL_LZ/TXF_LZ opcodes

  • radv: fix UB in radv_format_pack_clear_color for snorm formats

  • radv/perfcounter: guard select1 access in radv_emit_select

  • radv/perfcounter: add GFX11 performance counter selectors

  • radv: expose VK_KHR_performance_query on GFX11

  • util, llvmpipe: flush subnormals to zero on ARM/AArch64

  • nir: fix dedup_entry memcmp on structs with padding

  • gallivm: fix lp_build_round on altivec/VSX

  • gallivm: fix small_unorm -> unorm8 fetch path on big-endian

  • nir/tests: allow relative error in compare_inexact

  • nir: fix f2u/f2i constant folding to poison NaN and out-of-range inputs

  • nir: use i2f32 for patterns with signed-extraction opcodes

  • nir: fix pack_uvec4_to_uint to mask input components to 8 bits

  • nir/tests: fall back to integer comparison when float interpretation is NaN

Matthieu Oechslin (7):

  • r600: Fix crash on R600/R700 with custom border color

  • r600: Improve and document R600_TRACE

  • r600: Stop emitting relocs with virtual address enabled

  • r600: Workaround GPU hang with compute shaders when VA is enbaled

  • r600: Calculate address at emit time for SSBOs

  • r600: Fix MSAA 2D view from array with VA enabled

  • r600: Document RADEON_VA and remove SB options references

Mauro Rossi (5):

  • radv: Fix gnu-empty-initializer errors in 480a94fb

  • radv: Fix gnu-empty-initializer errors in 8c10eab1

  • radv: Fix gnu-empty-initializer errors in ca9191a8

  • intel/common: remove fallthrough annotation in unreachable code

  • pan/perf: fix building error due to ‘Mali G1.xml’ file name with space

Maíra Canal (2):

  • etnaviv/ml: derive stride-2 destriding offsets from padding

  • v3dv: Drop legacy comments about single-sync support

Mel Henning (36):

  • nak: Use shader_info->var_copies_lowered

  • nak: Use NIR_LOOP_PASS

  • nvk: Split out nvk_cmd_fill_memory

  • nvk: Allocate a zcull save region in fewer cases

  • nvk: Zero zcull data in layout transition

  • nvk: Don’t LOAD_ZCULL w/ VK_RENDERING_RESUMING_BIT

  • nvk: Re-enable zcull save/restore

  • nvk: Add a wfi for blackwell in CmdDispatchIndirect

  • nvk: Disable compression on Turing

  • compiler/rust: Fix inline wrapper include dir

  • nak/nvdisasm_tests: Fix expected value of F16v2

  • nak: Fix encoding of f16x2 min/max on sm90+

  • vk/meta: Move get_uint_format_for_blk_size to common

  • nvk: Make nvk_cmd_buffer_queue_flags non-static

  • nvk: Split out aligned_for_linear_attachment

  • nvk: Use meta for image copies where possible

  • nil: Pass ImageDim to Tiling::choose()

  • nil: Pick tiling params closer to proprietary

  • nvk: Interp frag_coord at centroid for min_sample_shading

  • nvk: Fix DGC localsize computation

  • nvk: Serialize shaders with asm

  • nvk/meta: Rename begin/end with a _gfx suffix

  • nvk/meta: Implement save/restore for compute

  • nvk/meta: Add save_generic helpers

  • nvk: Use compute meta for some vkCmdCopyImage2

  • nvk: Use meta for vkCmdCopyImageToBuffer2

  • nvk: Use meta for vkCmdCopyBufferToImage2

  • nvk: Use meta for vkCmdCopyBuffer2

  • nvk: Add _ce suffix to nvk_cmd_fill_memory

  • nvk: Use meta for vkCmdFillBuffer

  • nvk: Handle large indirect stride pre-Turing

  • nvk: Prepare indirect draws for 64-bit stride

  • nvk: Convert draw/dispatch to device_address_commands

  • nvk: Don’t re-align ssbo size/address

  • nvk: Move ssbo_4b_align to drirc

  • nvk: Use ?: in nvk_physical_device_compiler_flags

Michael Cheng (11):

  • intel/ds: Add end_event_dyn() and CREATE_DUAL_EVENT_CALLBACK_DYN macro

  • intel/ds: Label compute events with dispatch dimensions in Perfetto

  • intel/ds: Label selected draw events with vertex count

  • brw: Fix ordered dependency exec_all handling on Xe2+

  • intel/brw: allow baking more SBID dependencies into instructions on Xe2+

  • intel/brw: Don’t bake a long-pipe RegDist with an SBID dependency

  • intel/brw: Factor out combinable_ordered_pipe() helper

  • nir/opt_gcm: add option to keep texture ops in large loops

  • intel/brw: keep texture ops in large loops

  • intel: Fix DEBUG_FS_SIMD mask

  • intel: Fix operator precedence in intel_simd_overridden

Michal Krol (7):

  • gallium: add pipe_sampler_view::min_lod_clamp

  • lavapipe: implement VK_EXT_image_view_min_lod with fractional minLod

  • gallivm/llvmpipe: fix VK_EXT_image_view_min_lod via texture handle path

  • lavapipe: lower array-deref-of-vec for mesh shader outputs

  • lavapipe: fix format properties for R10X6G10X6B10X6A10X6_UNORM_4PACK16

  • gallivm: honour exec mask in EmitMeshTasksEXT

  • gallivm: don’t deref a NULL buffer descriptor with an empty exec mask

Michel Dänzer (14):

  • winsys/amdgpu: Use render node only as fallback

  • mr-label-maker: Label src/gallium/winsys/amdgpu as radeonsi

  • mr-label-maker: Label src/gallium/winsys/radeon as r300, r600 & radeonsi

  • egl/gbm: Do not destroy BO of current front buffer

  • egl/gbm: Use local variable for better readability

  • egl/gbm: Eliminate max_age local variable

  • egl/gbm: Ignore buffers with no BO for destroying excess BOs

  • egl/gbm: Ignore current front buffer in get_back_bo

  • egl/gbm: Use local variable for better readability in get_back_bo

  • egl/gbm: Eliminate local variable “age” in get_back_bo

  • egl/gbm: Use continue instead of nested block

  • egl/gbm: Eliminate local variable “max_age” in get_back_bo

  • dri3: Increment draw->send_sbc after waiting for last presentation

  • dri3: Simplify target_msc calculation in loader_dri3_swap_buffers_msc

Mike Blumenkrantz (104):

  • lavapipe: KHR_device_address_commands

  • radv: add RADV_QUEUE_DISABLE env var for selectively disabling queues

  • llvmpipe: fix min_samples + A2C

  • lavapipe: fix indirect memory copies

  • lavapipe: fix pushconst data updating

  • lavapipe: null out local var to avoid uninit warning

  • util/format: support 256-bit formats in util_format_get_tilesize()

  • lavapipe: use the right type for DGC mesh draws

  • lavapipe: rework immutable samplers

  • lavapipe: allow fbfetch with shader objects

  • vk/cmd_queue: always ceil() param lens

  • vulkan: update spec to 1.4.350

  • lavapipe: maintenance11

  • llvmpipe: always set view_index for linear rasterizer

  • llvmpipe: unify setting raster_state for thread data

  • lavapipe: update cbuf count when remapping attachments

  • lavapipe: unset attachment remap state if pColorAttachmentLocations==NULL

  • lavapipe: fix setting colormasks when attachments get remapped

  • ci: stop skipping HIC tests on lavapipe

  • zink: use maintenance5 to more effectively set storage texel usage for bufferviews

  • zink: delete zink_resource_object::storage_buffer

  • zink: remove remaining maint5 checks

  • aux/trace: silence -Waddress warnings in macros

  • zink: delete unused descriptor variable

  • zink: fix mixing of mesh descriptor bindings with gfx bindings

  • meson: fix renderdoc integration define

  • vulkan: move vk_shader_stages_from_bind_point() to vk_util

  • zink: disable implicit sync handling for qcom proprietary

  • zink: rework custom sample locations

  • lavapipe: enable some forgotten ds3 states

  • zink: fix unbinding vertex buffers from null VS state

  • zink: add another anv/adl flake

  • zink: create views for samplers lazily

  • lavapipe: correctly disable depth/stencil in secondaries

  • vk/cmd_queue: simplify gross struct duplication

  • zink: use custom sample locations to (mostly) handle multisample=disabled

  • zink: link up vs COLx vars -> fs BFCx

  • zink: be more conservative about query pool sizing

  • lavapipe: stop using pipeline layouts in some places

  • lavapipe: Implement VK_EXT_descriptor_heap

  • zink: handle uint wrapping with batch submit count

  • zink/bo: reduce wasted memory due to the size tolerance in pb_cache

  • zink/bo: add an enum to disambiguate bo types

  • zink/bo: stop using pb_buffer vtable for destroy

  • zink/bo: use only a single layer of slabs

  • zink/bo: use pb_buffer_lean to save a little mem

  • zink/bo: check for usage before completion when reclaiming bos

  • zink: use maint11 for sso shader object compile

  • zink/clear: fix full_clear condition in texture clear

  • zink/clear: handle texture clears on current fb texture

  • llvmpipe: create a zeroed payload for use without task shaders

  • zink: always return DMA_BUF type handles from resource_get_handle

  • zink: tag tc info update in a few more places

  • util/tc: iterate the rp info more accurately during batch execution

  • aux/tc: enforce strict resolve semantics

  • vulkan/wsi: pass VkSurfaceCapabilities2KHR to get_capabilities

  • vulkan/wsi: add VK_IMAGE_CREATE_MULTISAMPLED_RENDER_TO_SINGLE_SAMPLED_BIT_EXT where supported

  • lavapipe: EXT_multisampled_render_to_swapchain

  • zink: fix import2d sampler view creation

  • zink: when triggering zink_blit_barriers() for src==dst, apply separate barriers

  • zink: stop forcing barriers if previous access was write

  • zink: properly invalidate fb attachments on dontcare stores

  • zink: proactively apply transfer sync when tracking renderpasses

  • zink: don’t invalidate cbufs without inlined resolve

  • zink: add some ci flakes

  • tu: handle partially set resolve attachment info without crashing

  • util/tc: store resolve geometry to rp info

  • zink: use tc info to handle partial resolves

  • util/tc: unset TC_RESOLVE_STRICT

  • zink: set NO_TASK_SHADER for pipeline layouts with shader objects

  • zink: a618 ci updates

  • zink: always use src stages when flushing glMemoryBarrier calls

  • zink: always flush specified memory access for glMemoryBarrier calls

  • zink: reset usage following SHADER_WRITE access

  • zink: drop imageless_framebuffer requirement

  • zink: add a vb param to vertex buffer binding

  • zink: move vb binding out of c++

  • zink: hook up VK_KHR_device_address_commands

  • zink: use DAC for vertex binding

  • zink: fix a missing case of zink_batch_submit_count_diff()

  • zink: stop unsetting resource usage on batch reset

  • zink: free nir if cs program create fails

  • zink: enable signed vbs

  • st/pbo_compute: account for drivers failing to create cs shaders

  • zink: split more read/write barriers

  • zink: stop adding usage with last-ref tracking

  • zink: unset unordered access on ordered transfer ops

  • zink: use bigger hammer to force sync between unordered->main cmdbufs

  • zink: add api for disabling reordered read/write

  • zink: unset ordered_access_is_copied when disabling unordered access

  • zink: noop per-resource synchronization for unordered->ordered access

  • gallium/cso: make unbind_context an explicit call

  • lavapipe: handle depth blit aspect masking

  • st/context: unbind gs shader before deleting hw select gs shaders

  • zink: always un-suspend queries on end

  • lavapipe: advertise dynamicRenderingLocalReadDepthStencilAttachments

  • zink: fix the fix for ZINK_RENDERDOC=all

  • zink: don’t increment unique_id for reused bos

  • zink: translate depth write ALWAYS to GE/LE if possible

  • util/blitter: fix blitting multiple array layers

  • zink: start ZINK_DEBUG=perfinfo

  • zink: revert cached mem handling for staging uploads

  • zink: add anv ci flake

  • zink/ci: switch zink/anv jobs to surfaceless+i915

Mohamed Ahmed (8):

  • nil/modifiers: Clarify drm_format_mods_for_format rejecting modifiers for unsupported color formats

  • nvk: Calculate and stash the plane offset and alignment at create time

  • nvk: Extend tiled_shadow to be multiplanar

  • nvk: Defer tiled shadow plane memory allocation to draw time

  • nvk: Enable multiplanar YCbCr linear modifiers

  • nvk: Use the pre-calculated offsets for sparse binds

  • nvk: Remove nvk_image_plane_size_align_B()

  • nil: enable PLC for compressed data

Nanley Chery (23):

  • intel/blorp: Halve max bpp for some redescribed blits

  • anv: Add transfer_src usage for ANDROID_external_format_resolve

  • anv: Improve the fast clear layout perf-warn

  • anv: Improve the CCS_E-incompatible perf-warn

  • anv: Avoid aux-disabling paths for block-compression

  • anv: Allow CCS on more storage images for gfx12.5

  • anv: Move storage check out of CCS-compat helper

  • anv: Flush previous aux-mode changes

  • intel/isl: Define a CMF for ASTC formats

  • intel/isl: Fix the initial state HiZ state for Xe2+

  • anv: Dedent a closing curly brace

  • anv: Skip some CCS performance warnings on gfx9-11

  • anv: Allow partial depth fast clears on gfx12+

  • anv: Set TRANSFER_DST_BIT for HiZ operations

  • intel: Add and use ISL_AUX_USAGE_ZCS

  • hasvk: Delete enum anv_depth_reg_mode

  • iris: Rework HiZ plane optimization disabling

  • iris: Disable HiZ planes for some read-only tests

  • anv: Track the depth buffer aux usage

  • anv: Enable overriding HiZ in depth stencil state

  • anv: Rework HiZ plane optimization disabling

  • anv: Bypass HiZ planes for read-only depth tests

  • anv: Drop the anv_disable_hiz drirc option

Natalie Vock (9):

  • radv/rt: Don’t overwrite bvh_base at the start of the traversal loop

  • radv: Dump printf buffer after detecting a GPU hang

  • radv/rt: Cache stack sizes of ahit/isec shaders from imported NIR

  • radv: Work around midpoint sorting issues instead of disabling

  • radv: Fix destroying address binding reports

  • mailmap: Update my email

  • nir/opt_loop: Don’t peel header blocks that jump

  • radv/nir: Clean up descriptor index lowering

  • radv: Expose mutable acceleration structure descriptors

Nataraj Deshpande (1):

  • intel/perf: map ray tracing counters to RAYTRACING group

Neha Bhende (1):

  • svga: fix shared memory index for svga driver

Nemallapudi, Jaikrishna (1):

  • intel/dev: fix timebase_scale ticks-to-ns precision loss across 2^32

Nick Hamilton (5):

  • pco: fix clamping the array index when shaderImageGatherExtended is enabled

  • pvr: Enable shaderImageGatherExtended

  • pvr: Revert don’t csb emit multi-layer clear attachments without rta support

  • pvr: Fix load-op shader when loading from a 2d image view of a 3d image

  • glthread: fix check for unroll draws using user VBOs when the ctx supports GLES

Okenczyc, Andrzej (1):

  • amd/vpelib: Report unsupported status if streams target rect equals 0

Olivia Lee (23):

  • pan/bi: fix memory access alignment

  • pan/genxml: add definitions for adjacency draw modes

  • panvk: add support for adjacency primitive topologies

  • panvk: fix executable properties handling for IDVS varying shaders

  • pan/csf: rename immediate CS add builder functions

  • pan/v13: add CS builder functions for reg/reg add and sub instructions

  • pan/v13: add CS builder functions for shift instructions

  • pan/v13: implement constant integer multiplication CS helper

  • pan/v13: implement CS udiv

  • panvk: remove redundant invalid primitive topology cases

  • panvk/csf: allow SYNC_WAIT-style synchronization in launch_gfx_cs

  • panvk: add create_shader helper for compiling full meta shaders

  • panvk: return both gpu and cpu pointers from cmd_prepare_*_push_uniforms

  • poly: allow VS outputs with <32 bits per component

  • poly: refactor GS lowering to store output and selected variables together

  • poly: preserve src_type in GS rast shader outputs

  • poly: preserve output component counts in GS

  • poly: move hk passthrough GS code to libpoly

  • poly: allow specifying output types in passthrough GS

  • nir: add nir_slot_num_components helper

  • poly: allow specifying component count in passthrough GS key

  • poly: clarify assertion failure message in lower_store_to_var

  • panvk/csf: flush primitives generated query writes from CSF

Omar Rashwan (2):

  • intel: Fix bit width of int literal in eu stall viewer

  • intel: define type for std::max in eu stall viewer

Patrick Lerda (19):

  • r600: refactor r600_shader_buffer_info_sel

  • r600: refactor eg_setup_buffer_constants

  • r600: rename sh_txs_cube_array_comp to sh_resinfo_via_uniform

  • r600: cypress resinfo buffer size workaround

  • r600: fix alpha-to-coverage and alpha-to-one used together

  • r600: add sample_lz and sample_c_lz opcodes compatibility

  • r600: update r600 nir for sample_lz and sample_c_lz

  • r600: enable EXT_texture_shadow_lod

  • docs/features: add GL_EXT_texture_shadow_lod

  • r600: remove r600_get_hw_atomic_count

  • r600: fix atomic buffer offset

  • r600: implement tes and tcs instanced gl_PrimitiveID support

  • r600: make r600_copy_region_with_blit global

  • r600: implement msaa 2d view from array

  • r600: update vertex emit_varying_pos

  • r600: fix atomic_counter_post_dec

  • r600: update memory barrier operations

  • i915: fix emit_hw_vertex() unbounded memory access

  • r600: update muladd support configuration

Paulo Zanoni (30):

  • intel/isl: fix assert when surf->size_B is > UINT_MAX

  • intel/isl: warn about excessive num_elements only once

  • anv: don’t silently convert view ranges from u64 to u32 then u64

  • docs/envvars: remove ANV_SPARSE and ANV_SPARSE_USE_TRTT

  • docs/envvars: document ANV_SYS_MEM_LIMIT

  • docs/envvars: update the ANV_DEBUG documentation

  • intel/nir: fix sparse shadow comparison for BRW

  • anv/sparse: bring back our (limited) support for depth/stencil

  • brw: evict memory for workgroup scope in Xe2 and newer

  • intel/mi_builder: add mi_ixor()

  • intel/mi_builder: add mi_umax2()

  • intel/blorp: prepare for usage of mi_builder.h

  • libcl/vk: add aligned(4) to VkCopyMemoryIndirectCommandKHR

  • libcl/vk: add VkCopyMemoryToImageIndirectCommandKHR and its members

  • anv: implement VK_KHR_copy_memory_indirect

  • intel/tools: fix stall_csv_filename maybe-unitialized error

  • intel/brw: move cache_mode assignment to after send->sfid choice

  • brw: split cache mode selection into atomic, load and store modes

  • brw: have a single if-ladder to pick cache_modes

  • brw: control cache_mode through bypass_{l1,l3} variables

  • intel/blorp: don’t silently ignore compilation failures

  • intel/blorp: fix blorp base key initialization

  • intel/blorp: move struct blorp_blit_prog_key to blorp_blit.c

  • intel/blorp: don’t include “util/format_rgb9e5.h”

  • anv: give anv_ensure_fp64_shader() a chance to be called

  • brw: don’t preprocess software doubles if opts->softfp64 is not set

  • anv: don’t put clear colors for aliased images in private bindings

  • intel/blorp: rearrange struct blorp_blit_prog_key

  • intel/blorp: pack every blorp key struct

  • intel/blorp: memset(0) blorp keys during initialization

Pavel Ondračka (111):

  • r300,i915/ci: update expectations

  • r300/ci: update expectations

  • i915/ci: update expectations

  • r300: fix MSAA resolve COLORPITCH tiling after pipe_surface de-pointerization

  • r300: dirty VS state when switching variants

  • dri3: add big-endian 8888 fourccs to dri3_cpp_for_fourcc

  • dri: add big-endian 8888 entries to dri2_format_table

  • dri: add big-endian 8888 entries to driImageFormatToSizedInternalGLFormat

  • nir: fix partial loop unroll OOB check for loops not starting at 0

  • nir/tests: add helpers for counting used/unused instructions

  • nir/tests: add partial unroll OOB tests

  • r300: drop unused input arrays from ntr

  • r300: drop multiple ubo support from ntr

  • r300: drop GS/tess and load_draw_id support from ntr

  • r300: drop framebuffer fetch handling from ntr

  • r300: drop GLSL 4.x texture ops from ntr

  • r300: drop unsupported sampler dimensions from ntr

  • r300: drop the i915g vertex_id/instance_id U2F branch from ntr

  • r300: drop GLSL 4.x interpolation intrinsics from ntr

  • r300: drop opcode paths lowered before emission from ntr

  • r300: drop TEX2/TXB2/TXL2 dead path from ntr

  • r300/ci: run EGL deqp tests

  • r300/ci: update expectations

  • i915/ci: update expectations

  • r300: remove unused LIT opcode

  • r300: pack immediates more aggressively to avoid running out of constant slots

  • r300: reuse positive and negative immediate values

  • r300: remove extra newline for compiler errors

  • r300: remove the redundant control flow checks

  • r300: stop dumping TGSI

  • r300: always convert to NIR and move ntr later

  • r300: collect input/output info directly from NIR

  • r300: move compiler init earlier and to a helper

  • r300: fix use-after-free of remap_table in rc_remove_unused_constants

  • r300: build the dummy fragment shader directly as NIR

  • r300: emit one full vec4 immediate per NIR load_const

  • r300: move r300_transform_*_trig_input out of nir_to_rc

  • r300: extract TGSI->RC translation helpers into nir_to_rc.h

  • r300: emit RC instructions directly from nir_to_rc

  • r300: drop the TGSI opcode middle step from ntr

  • r300: allocate FS outputs from NIR locations

  • r300: use NIR varying locations directly in ntr

  • r300: stop declaring samplers with ureg in ntr

  • r300: lower sysvals to varyings

  • r300: use backend texture targets directly in ntr

  • r300: drop ureg shader properties in ntr

  • r300: drop ureg_DECL_temporary in ntr

  • r300: drop ureg_DECL_vs_input in ntr

  • r300: drop dead tg4_offsets and query_levels paths in ntr

  • r300: drop ureg_DECL_address

  • r300: get rid of user_src and ureg_dst

  • r300: stop using ureg_dst_undef in ntr

  • r300: get rid of ureg_program in ntr

  • r300: get rid of ureg_src_undef in ntr

  • r300: get rid of ntr_emit_load_output

  • r300: get rid of the precise modifier

  • r300: use RC registers directly in nir_to_rc

  • r300: drop more dead ntr code

  • nir/algebraic: prevent ffract optimization on lowered ffloor

  • r300: fix R300_VAP_TCL_BYPASS state leak

  • r300: fix vs->first leak in swtcl delete path

  • r300: add NIR pass to append the wpos output

  • r300: add NIR pass to add required color outputs

  • r300: convert swtcl vertex shader setup to NIR

  • r300: remove dead first-time build path from r300_pick_vertex_shader

  • r300: clean up some dead draw/TGSI leftovers

  • r300: remove draw support on big endian

  • meson: require r300 LLVM draw only on x86

  • gallivm: add NIR pass to lower load_ubo_vec4 to load_ubo

  • gallivm: add no_integers intrinsic fixup pass

  • gallivm: add NIR pass to lower float if conditions

  • gallivm: add algebraic NIR pass for no_integers

  • gallivm: pass base NIR ALU types to cast_type

  • r300: enable VS instance ID in draw

  • r300: prepare for the the draw NIR path

  • draw: use gallivm NIR for no_integers vertex shaders

  • i915: use gallivm NIR for vertex shaders

  • r300: use signed index offset for index translation

  • r300: always use 32-bit indices on big endian

  • r300: use R32_FLOAT as 32-bit dummy vertex format

  • r300: fix occlusion query results on big endian

  • r300: fix BE 8888 render-to-texture endian state

  • r300: fix BE RGB565/RGB5 render-to-texture formats

  • r300: fix BE constant blend color for colorbuffer formats

  • r300: fix BE depth/stencil raw transfer endian state

  • r300: clean up endian swap selection

  • r300: add NIR LICM pass to enable removal of residual loops

  • i915: add NIR LICM pass to enable removal of residual loops

  • r300: remove dead state constants

  • r300: add private NIR state constant plumbing

  • r300: handle texture destination constraints in nir_to_rc translation

  • r300: lower backend texture coordinate handling in NIR

  • r300: lower fragment position in NIR

  • r300: lower VS seq/sne in NIR

  • r300: lower FS alpha-to-one in NIR instead of backend

  • r300: handle FS depth output channel mapping at nir_to_rc time

  • r300: lower r300 FS derivative stubs in r300_optimize_nir

  • r300: move r500 FS derivative fixup to nir_to_rc translation

  • r300: extend the wined3d A0 rounding pattern recognition

  • r300: some post-int/bool lowering optimizations

  • r300: don’t split ALU instructions on R5xx

  • r300: run rgb alpha conversion after RC optimize

  • r300: penalize presubtract NOP in pair scheduling

  • r300: run late CSE after lowering vectors to registers

  • r300: fix swtcl per-vertex point size

  • r300: remove redundant FACE input workaround

  • r300: keep NIR output count consistent

  • r300: unify WPOS output handling between swtcl and hwtcl

  • r300: share vertex shader variant handling in hwtcl and swtcl paths

  • r300: emulate gl_FrontFacing on R3xx/R4xx

  • nv50_ir_ra: align B96 spill slots to vec4

Peng_Lx (1):

  • turnip/kgsl: close the dma-buf fd of our own allocations

Peyton Lee (14):

  • amd/vpelib: add alpha fill support check

  • amd/vpelib: Support vpe 2.0

  • amd/gmlib: add tm_generate_formatted_3DLut

  • radeonsi/vpe: add VPE 2.0 support

  • frontends/va: add ABGR format mappings

  • amd: validate and expose VPE 2.0.0

  • radeonsi: gate format and rotate/flip support by VPE version

  • amd/vpelib: support vpe 2.2

  • ac/gpu_info: add VPE_2_2 support

  • radeonsi/vpe: adjust message

  • amd/vpelib: tighten external LUT compound color pipeline updates

  • amd/vpelib: refine coding style

  • amd/vpelib: Fix Color Corruption AV1 Issue

  • amd/vpelib: Replace hardcoded bg format table size with sizeof

Philipp Zabel (2):

  • etnaviv/isa: Fix Meson warning about etnaviv_isa_rs dummy library

  • meson: add rusticl to with_driver_using_cl

Physics Enthusiast (1):

  • venus: allow to use vtest as a fallback for virtgpu

Pierre-Eric Pelloux-Prayer (40):

  • radeonsi: clamp cp prefetch size

  • ac/info: add gfx12.1 identification

  • radeonsi/tests: update expectations

  • amd/virtio: use AMDGPU_VA_MGR_RESERVE_HALF_VA_FOR_PRT

  • amd/virtio: fix amdgpu_sw_info_address_prt_wa_control_bit handling

  • radeonsi: handle NULL return value from amdgpu_cs

  • gallium/vl: only release created sampler views

  • radeonsi: delay aux context initialization to first use

  • radeonsi: add has_gfx_compute property to si_screen

  • radeonsi: don’t use staging texture when we can’t blit

  • radeonsi/vce: deal with has_gfx_compute being false

  • radeonsi: create a mm subfolder for multimedia code

  • radeonsi: add si_init_screen_nir_options

  • radeonsi: add gfx subfolder

  • radeonsi: move shader cache code to new file

  • radeonsi: extract si_init_gfx_caps from si_init_screen_caps

  • radeonsi: add si_resource_copy_buffer

  • radeonsi/gfx: add si_gfx_screen.c

  • radeonsi/gfx: move code from si_get to si_gfx_screen

  • radeonsi: add si_gfx_context.c and move code from si_pipe.c

  • radeonsi: add si_context.c

  • radeonsi: move all multimedia files to mm

  • radeonsi: move more code to gfx subfolder

  • radeonsi: move function prototypes from si_pipe.h to si_gfx.h

  • radeonsi/gfx: remove unnecessary u_stub usage

  • radeonsi/gfx: move static inline helpers to si_gfx.h

  • radeonsi: add tests subfolder and move AMD_TEST code inside

  • gallium/u_blitter: remove unused skip_viewport_restore

  • radeonsi: fix sqtt setup

  • radeonsi/sqtt: hash only the relevant part of the shader key

  • radv, radeonsi: do sqtt buffer_size calc using uint64

  • ac/sqtt: add ac_sqtt_update_bo_size

  • radeonsi: fix sdma copy for gfx10

  • radeonsi: consolidate aux context creation into si_get_aux_context

  • radeonsi: use aux context locks in si_destroy_screen

  • ac/parse_ib: initialize data variables to 0

  • radeonsi: delay si_disk_create_cache call

  • radv/rra: bump rt_driver_interface_version

  • radv: restore RRA capture support

  • radeonsi: fix typo in si_copy_from_staging_texture

Pohsiang (John) Hsu (15):

  • mediafoundation: periodic clang-format

  • mediafoundation: code clean up

  • meson: Make with_gfx_compute depend on video encode support (mediafoundation)

  • d3d12: add av1 handling to d3d12_video_encoder_get_encode_headers and d3d12_video_encoder_update_current_encoder_config_state_av1

  • mediafoundation: extract code to ProcessDX12EncodeContext

  • mediafoundation: fix a few minor variant bool handling

  • mediafoundation: detach xThreadProc frame processing from apiLock to unblock concurrentt ProcessOutput calls

  • d3d12: fix infinite gop handling in d3d12_video_enc_av1.cpp

  • mediafoundation: define AVC_LOG2_MAX_FRAME_NUM_MINUS4, HEVC_LOG2_MAX_PIC_ORDER_CNT_LSB_MINUS4 instead of using number.

  • mediafoundation: preserve low latency ping pong behavior between ProcessInput and ProcessOutput

  • mediafoundation: change default value for HEVC_LOG2_MAX_PIC_ORDER_CNT_LSB_MINUS4 to 12

  • mediafoundation: initial av1 dx12 hmft prototype

  • d3d12: add support to output temporal delimiter for AV1 via raw_header

  • mediafoundation: ask for temporal delimiter for AV1 via raw header

  • d3d12: fix msvc build warning C4819

Qiang Yu (2):

  • ac,radeonsi,radv: fix print IB assertion fail for reserved fields

  • ac,radeonsi,radv: use V_581A_* engine sel for non-pws acquire_mem packet

QwertyChouskie (3):

  • docs/features: Remove mentions of r300 and nv30

  • docs/features: Fix typo

  • docs/features: Mark VK_EXT_descriptor_heap done for anv

Radu Costas (6):

  • pco: Set register classes for vec refs

  • pco: Move preproc_vecs out of loop

  • pco: Add debug variables for RA

  • pco: Move RA context handling to state-based

  • pco: Try allocating with optimal temp registers

  • pvr, ci: Update axe and bxs failure list

Rahul Mahantappa Bhadrashette (1):

  • radeonsi: emit MESHLET registers for mesh shaders in gfx10_emit_shader_ngg

Raviraj Uppal (2):

  • driconf: disable allow_rgb16_configs for SPECviewperf

  • radv: app workaround implemented using internal layers for GFXBench 5.0

Rhys Perry (122):

  • ac/nir_lower_global_access: perform range analysis if useful

  • ac: add gfx11.7 enums

  • aco/gfx11.7: add opcode numbers

  • aco: adjust some gfx_level checks for gfx11.7

  • aco/gfx11.7: don’t create v_dot2c_f32_f16

  • aco/gfx11.7: don’t use v_pack_b32_f16 in do_pack_2x16

  • aco/gfx11.7: allow any src VGPR for VOPD with two v_dual_mov_B32

  • ac/gpu_info/gfx11.7: enable has_point_sample_accel

  • aco/gfx11.7: claim support

  • radv/gfx11.7: take GFX12 paths in radv_nir_lower_cooperative_matrix

  • radv/gfx11.7: enable float8

  • radv/gfx11.7: enable shaderMixedFloatDotProductFloat8AccFloat32

  • radv/gfx11.7: don’t advertise shaderImageFloat32AtomicMinMax

  • aco: refactor spiller to use spills_needed variable

  • aco: prefer spilling smaller temporaries if it finishes spilling

  • aco: use RegisterDemand::operator[] more

  • radv: move ac_nir_lower_indirect_derefs to end of radv_shader_spirv_to_nir

  • radv: lower indirect derefs after linking

  • radv: don’t use radv_optimize_nir after lowering indirect derefs for RT

  • ac: move lds_size_per_workgroup to ac_compiler_info

  • ac: move has_cs_regalloc_hang_bug to ac_compiler_info

  • radv: assert there is no padding in cache keys

  • radv: initialize nir_shader_compiler_options directly in compiler info

  • radv: move load_grid_size_from_user_sgpr to radv_physical_device

  • radv: move use_llvm to radv_compiler_info::key

  • radv: move fields to radv_compiler_info::key

  • radv: add fields to radv_compiler_info from radv_physical_device_cache_key

  • radv: remove radv_compiler_info::cache_key

  • radv: hash radv_compiler_info::key into the cache key

  • radv: remove most fields from radv_physical_device_cache_key

  • radv: remove radv_physical_device_cache_key

  • radv: remove radv_device_cache_key

  • ac/llvm: fix isub image atomic

  • radv: inline shader_compile()

  • radv: remove radv_aco_convert_opts

  • radv: replace radv_nir_compiler_options with a LLVM one

  • radv: split radv_compiler_info’s family into debug::family and key::family

  • aco/gfx11.7: fix v_pk_min_f16/v_pk_max_f16 opcode numbers

  • nir/search: fix nir_algebraic_automaton after constant folding op(bcsel)

  • nir: rename nir_src_parent_instr to nir_src_use_instr

  • radv,ac: make rembrandt and vangogh cache compatible

  • radv: don’t pass GPU name to disk_cache_create

  • radv: remove family from cache key

  • aco/validate: fix some RA validator error messages

  • aco/ra: fix v3b VALU at byte>0

  • aco/ra: test the register file in get_reg_specified() when necessary

  • aco: add helpers to get instruction subdword capabilities

  • aco: rework subdword definition RA validation a bit

  • aco/ra: fix compact_relocate_vars path for get_reg_for_operand

  • aco/ra: fix fill() with certain subdword cases

  • aco: fix regclasses for spill/reload subdword temporaries

  • aco/ra: don’t rename phi operands in get_reg_phi()

  • aco/ra: use phi_dummy instead of is_phi()

  • aco/ra: remove precolored checks in get_reg_impl()

  • nir/algebraic: optimize ishl(iadd(ishl, ishl))

  • nir/algebraic: optimize ishl(iadd(iadd(iadd(a, #b), c), d), #e)

  • nir: make cmat_muladd_amd a subgroup intrinsic

  • nir: add load_deref_transpose_amd

  • nir: add load_global_transpose_amd

  • nir,ac/nir,aco: add load_global_tr_amd

  • radv: track cooperative matrix robustness

  • radv: use load_deref_transpose_amd for transposed cooperative matrix loads

  • radv: fix usage of radv_nir_cmat_length

  • aco: add cost estimation of s_barrier

  • aco: don’t emit workgroup-scope p_barrier for single-wave workgroups

  • aco/waitcnt: always use uint32_t for event masks

  • aco: fix printing of primitive exports

  • aco: optimize redundant s_wait_alu vm_vsrc(0) during waitcnt insertion

  • aco: only assume load/store with semantic_atomic is atomic

  • aco: don’t emit waitcnts before subgroup-scope execution barriers

  • aco: add split barrier instructions

  • aco: use split barrier instructions

  • aco: schedule split barriers

  • ac/gpu_info: add has_smem_partial_oob_access_bug

  • radv: workaround has_smem_partial_oob_access_bug

  • nir/opt_undef: fix prefer_nan

  • aco: don’t increase barrier exec scope to subgroup

  • radv/bvh: use atomic load/store in update_gfx12.comp

  • ac/lower_global_access: set cursor earlier

  • ac/lower_global_access: rewrite try_extract_additions

  • ac/lower_global_access: parse u2u64 even if *out_offset!=NULL

  • ac/lower_global_access: extract constants after ishl/imul

  • ac/lower_global_access: combine multiple 32-bit offsets

  • radv: use radv_shader_stage_key::keep_{statistic,executable}_info more

  • radv: move nir_debug_info from debug to key

  • radv: don’t create nir_string if dump_shader=true

  • radv: simplify radv_declare_shader_args parameters

  • radv: add radv_shader_stage_key::keep_shader_arg_info

  • radv: cache shader IR, asm and spir-v

  • radv: parse stats from binary in radv_parse_binary_debug_info

  • radv: make raytracing radv_shader_stage_key array per-shader

  • radv: make raytracing radv_shader_stage_key initialization per-shader

  • radv: merge radv_shader_stage_key for combined ahit/isec shaders

  • radv: rework creation of traveral radv_shader_stage_key

  • radv: inline some helpers used in radv_pipeline_get_shader_key

  • radv: use nir_opt_uub

  • radv: repeat loop in radv_optimize_nir_algebraic_early more

  • radv: do nir_opt_algebraic last in radv_optimize_nir_algebraic_early

  • radv: fix barriers in decompress shaders

  • drm-shim: implement most readlink() without initializing the shim

  • drm-shim: skip init_shim() if drm_shim_fd_lookup() would be NULL

  • nir: add NIR_MEMORY_CONTROL_ARRIVE and NIR_MEMORY_CONTROL_WAIT

  • nir: add nir_lower_disordered_control_barriers

  • aco: implement NIR_MEMORY_CONTROL_ARRIVE and NIR_MEMORY_CONTROL_WAIT

  • vtn: implement SplitBarrierEXT

  • radv: implement VK_EXT_shader_split_barrier

  • radv: do radv_parse_binary_debug_info in radv_shader_dump_asm

  • ac/nir: skip SMEM fixup for more 32-bit load_global addresses

  • radv: set RADV_CMD_DIRTY_GFX12_HIZ_WA_STATE around attachment clears

  • vtn,nir: print shader filename and spec constants

  • radv: include debug information in bvh shaders

  • radv: shorten internal shader names

  • vulkan/bvh: fix to_emulated_float(-0.0)

  • vulkan/bvh: don’t update min/max_bounds with inactive nodes

  • vulkan/bvh: enable SignedZeroInfNanPreserve in leaf.h

  • vulkan/bvh: limit valid 32-bit node keys to 0xfffffe00

  • ac/nir/ngg: don’t shrink device-scope memory barriers

  • ac/nir/ngg: track uniformity with multiple set_vertex_and_primitive_count

  • nir/load_store_vectorize: rework barriers to use entries

  • nir/load_store_vectorize: recreate entry key when adding from predecessor

  • vtn: don’t fail at uniformity decorations on variables and OpBufferPointer

  • radv: drop support for cooperativeMatrixRobustBufferAccess

Rob Clark (65):

  • tu: Remove use of fd_perfcntr_type

  • freedreno: Remove use of fd_perfcntr_type/result_type

  • freedreno/perfcntr: Remove type and result_type

  • freedreno/registers: Add json to describe perfctr groups

  • freedreno/registers: Generate perfcntr tables

  • freedreno/perfcntrs: Switch to generated perfcntr tables

  • freedreno/registers: Small reg32 vs reg64 fixes

  • freedreno/registers: Sync back xml changes from kernel

  • freedreno/registers: Add pipe to perfcntr group

  • freedreno/registers: Add gen8 perfcntr support

  • freedreno/registers: Correct register name

  • freedreno/registers: Add gen8 perfcntrs

  • pps: Re-emit time clock_sync more regularly

  • freedreno/ds: Use gpu timestamps

  • freedreno/ds: Split a6xx/a7xx counters out

  • tu: Fix preemption latency selector values

  • freedreno/a6xx: Expose subgroup ops

  • rusticl: Support/ignore -qcom-accelerate-16-bit

  • freedreno/common: Fix X2-90, add X2-85

  • freedreno/registers: Skip deprecated warns for kernel

  • freedreno/registers: Add a6xx CMP counter group

  • freedreno/registers: Gen8 perfcntr fixes

  • drm-uapi: Sync msm_drm.h

  • freedreno/common: Add ioctl ptr helpers

  • freedreno/fdperf: Move where we setup counter groups

  • freedreno/fdperf: Prepare for partial-counter usage

  • freedreno/fdperf: Add PERFCNTR_CONFIG support

  • freedreno/ds: PERFCNTR_CONFIG support

  • freedreno/ds: Add a8xx derived counters

  • freedreno/perfcntrs: Add helpers to resolve group and countable

  • freedreno/perfcntrs: Add helper to assign counters

  • tu: Use counter allocation helper

  • freedreno/a6xx: Use counter allocation helper

  • freedreno/perfcntrs: Refactor derived counter setup

  • freedreno/perfcntrs: Use helper for derived counters

  • freedreno: Skip BV perfcntrs

  • tu: Disable preemption for counters on gen8

  • tu/gen8: Program slice selector regs

  • freedreno/a6xx: Program gen8+ slice SEL regs

  • freedreno/perfcntrs: Expose gen8 counters

  • freedreno/a6xx: Push RB_A2D_PIXEL_CNTL magic into blitter

  • freedreno/registers: Improve A2D docs

  • freedreno/registers: Add RB_RESOLVE_CNTL_0.YUV_PLANE_ID

  • freedreno/a6xx: Un-open-code RB_A2D_PIXEL_CNTL

  • tu: Un-open-code RB_A2D_PIXEL_CNTL

  • perfetto: Add API to flush track events

  • util/thread: Flush traces at thread exit

  • util/queue: Flush perfetto before blocking

  • rusticl: Flush perfetto track events

  • perfetto: Increase SMB size

  • perfetto: Use BufferExhaustedPolicy::kStall

  • freedreno/perfetto: Add non-draw stage

  • freedreno/perfetto: serialize clk snapshots

  • freedreno/perfetto: Use sequence-scoped clk

  • freedreno/crashdec: Update gpu revision parsing

  • freedreno/crashdec: Add additional HFI queue

  • freedreno/a6xx: Don’t clamp 32b clear values

  • freedreno/a6xx: Fallback for conditional blits

  • freedreno/a6xx: Handle R9G9B9E5 blits as R32_UINT

  • freedreno/a6xx: Set HALF_PRECISION for R11G11B10_FLOAT

  • freedreno/a6xx: Don’t forget UBO driver params

  • freedreno/decode: Fix shader stats in summary mode

  • freedreno/ci: Update a660-vk-traces-restricted checksums

  • nir/convert_address_format: Split convert_def into two passes

  • nir/convert_address_format: Handle non-deref sources

Rob Herring (Arm) (34):

  • ethosu: Make quantization shift signed

  • ethosu: Add a common initializer for struct ethosu_operation

  • ethosu: Store ethosu_tensor struct ptr in feature map

  • ethosu: Move stride calculation to lowering

  • ethosu: Fix concatenation OFM scaling

  • ethosu: Support axis 1 concatention

  • ethosu: Add fully-connected operation

  • ethosu: Rename ethosu_lower_add to ethosu_lower_eltwise

  • ethosu: Support element wise op with constant IFM buffer

  • teflon: Add multiply operation

  • ethosu: Add multiply operation support

  • teflon: Add TANH operation support

  • ethosu: Add logistic and TANH operations

  • teflon: Add hard swish operation

  • ethosu: Add hard swish operation

  • teflon: Add LeakyRelu operation

  • ethosu: Add LeakyRelu operation

  • teflon: Add quantize operation

  • ethosu: Add quantize operation

  • ethosu: Add reshape operation

  • teflon: Add minimum and maximum operations

  • ethosu: Add minimum and maximum operators

  • ethosu: Add performance counter debug output

  • teflon: Ensure all TfLiteRegistration fields are 0

  • meson: Skip NIR tests with headers-only NIR

  • teflon/tests: Use reference kernels

  • ethosu: Compute elementwise broadcasts from OFM shape

  • ethosu: Fix depthwise conv layout for IFM depth 1

  • ethosu: Flatten fully connected inputs

  • ethosu: Fix command dependency tracking

  • ethosu: Improve equal-cost block selection

  • ethosu: Use full scale for NOP pooling

  • ethosu: Preserve spatial dimensions for FC lowering

  • ethosu: Preserve fused pad extents

Robert Mazur (6):

  • ci: update firmware tag to ff46ce35

  • ci: update kernel tag to v6.19-mesa-712d

  • imagination/ci: Use standard CI-tron gfx-ci/linux kernel

  • pvr: switch core count mesa_logw() to pvr_finishme()

  • pvr: introduce PVR_IGNORE_FINISHME_WARNINGS envvar

  • pvr/ci: enable PVR_IGNORE_FINISHME_WARNINGS

Rohit Athavale (1):

  • mediafoundation: Test compile steps v/s step , and set build flag

Roland Scheidegger (1):

  • gallivm: fix subtle filtering issue with different min/mag filter for cube maps

Roman Stratiienko (3):

  • v3dv/android: Add deferred ANB allocation support

  • v3dv: move noop_job creation to device scope

  • v3dv: Emulate multi-queue support via vk_queue for Android

Romaric Jodin (1):

  • anv: Declare 00-mesa-defaults.conf as an input to anv_dricrc_gen.py

Rouf, Farhan (4):

  • amd/vpelib: Introduced reset to frontend

  • amd/vpelib: Changed cmd_info input for background segment

  • amd/vpelib: Chroma coefficient select for sampled formats corrected

  • amd/vpelib: Refactoring Reset Pipes Function

Rudraksha Gupta (1):

  • freedreno: add Adreno 225

Ryan Houdek (1):

  • turnip: Add an override to uncached memory type

Ryan Mckeever (2):

  • pan/bi: check if preds are dominated by header in bi_find_loop_blocks

  • docs: advertise VK_KHR_multiview support for Bifrost

Ryan Neph (1):

  • anv/xe: prevent WaitIdle optimization for fences with exported sync_fd

Ryan Zhang (4):

  • panvk: add VK_IMAGE_LAYOUT_DEPTH_READ_ONLY_OPTIMAL to host copy layouts

  • gfxstream/platform: add missing inc_include to platform_virtgpu build

  • panvk: set cfg cull status according to primitive topology

  • panvk: Drop empty SYNC_ONLY bind queue ops

Saeed, Ghamr (1):

  • amd/vpelib: variable was accumulating size and not reset properly

Sagar Ghuge (49):

  • anv/rt: Copy 16bytes at once instead of copying 8bytes

  • anv: Fix Wa_14021821874, Wa_14018813551, Wa_14026600921

  • brw: Pass write back register for ray query messages

  • intel/genxml: Update xml for dynamic stack ID control fields

  • anv: Enable dynamic stack ID control on Xe3+

  • intel/genxml: Disable compute walker mid-thread preemption

  • util: Increase array size to 20

  • intel/genxml: Added dispatch timeout counter extended field

  • anv: Update values for DispatchTimeoutCounter

  • anv: Set execution mask based on SIMD size

  • brw/rt: Commit hit even if we are skipping closest hit shader

  • brw/rt: Update committed hit leaf type properly

  • brw/rt: Use BLAS(Object) level to get the ray address

  • jay: Implement halt

  • anv/rt: Skip invalid node in child block count

  • anv: Pass vk_acceleration_structure_build_state as param

  • anv/rt: Extract common code in separate header

  • anv/rt: Use constant BVH offset instead of pushing

  • anv: Track parent-child map for BVH update

  • anv: Track leaf block offset map

  • intel: Add debug option to dump out parent-child map

  • anv: Implement update BVH

  • intel: Add debug hook to dump out BVH after update

  • ci-farms/vmware: Disable vmware tests for now

  • anv: Allocate lookup maps for update based on mode and flag

  • intel: Add drirc option to write lookup maps unconditionally

  • Revert “anv: Fix Wa_14021821874, Wa_14018813551, Wa_14026600921”

  • anv: Workaround game bug for Witcher3

  • jay: Extend CS payload to handle BTD stack IDs

  • jay: Handle nir_intrinsic_load_btd_stack_id_intel intrinsic

  • jay: Factor out RT message header build part

  • jay: Implement btd_retire instrinsic

  • jay: Implement BTD Spawn intrinsic

  • anv: Compile init RT shader with Jay

  • anv: Bump subgroup size for histogram and prefix shader

  • vulkan: use center/extent form for instance node AABB transform

  • jay: Control cache_mode through bypass_{l1,l3} variables

  • jay: Setup bindless thread payload

  • jay: Handle nir_intrinsic_load_btd_global_arg_addr_intel

  • jay: Handle nir_intrinsic_load_btd_local_arg_addr_intel

  • jay: Set simd width for bindless shader

  • jay: Handle EOT for RT shader

  • jay: Init header with zero for all components

  • jay: Stuff stackIDs for trace_ray message

  • brw: Track if CS uses fences

  • intel: Fix async compute thread limit

  • anv: No need to flush RT cache if we update buffer via CS

  • jay: Track max stack size for bindless shaders

  • jay: Add loop_once_halt opcode

Sahitya Kandru (1):

  • freedreno: Modify reg_size_vec4 for a608 and a612 to 32

Sam James (4):

  • glx: append extra_ld_args_libgl, not clobber

  • src: add -Wl,–no-fatal-rwx-sections for two libraries

  • gallium/dri: fix redundant Meson condition

  • util: remove bogus const attribute

Samuel Pitoiset (263):

  • vulkan: add an option to lower SHADER_RECORD_INDEX to non-uniform

  • radv: lower SHADER_RECORD_INDEX to non-uniform

  • radv/ci: document some HIC failures since addrlib uprev for GFX11.7

  • radv: add enable_mrt_output_nan_fixup to the physical cache key

  • ac/surface: add stencil-only support for host mem->surf copies

  • radv: add depth+stencil formats support with host image copy

  • radv: allow depth+stencil formats with host image copy

  • amd: allow addrlib to enable SIMD if possible

  • radv: advertise VK_EXT_host_image_copy by default on GFX10.3+

  • radv/ci: document more HIC regressions on NAVI10

  • vulkan: refactor vk_pipeline_robustness_state_fill() slightly

  • vulkan: pre-compute the default robustness state in the device

  • vulkan,treewide: stop passing vk_device to vk_pipeline_robustness_state_fill()

  • spirv,treewide: rework specialization constant

  • radv: fix GPU hangs with PS epilogs and secondaries properly

  • radv/rt: pass more parameters to radv_rt_nir_to_asm()

  • radv: add a radv_compiler_info object

  • radv: use radv_compiler_info everywhere during compilation

  • spirv: add support for SPV_KHR_constant_data

  • radv: advertise VK_KHR_shader_constant_data

  • radv: move queue related cmd buffer state to a new struct

  • radv: move uses_perf_counters to radv_cmd_buffer_queue_state

  • radv: move shader_upload_seq to radv_cmd_buffer_queue_state

  • radv: remove redundant initialization when beginning a cmdbuf

  • radv: zero-initialize radv_cmd_state only when a cmdbuf is reset

  • radv: pass radv_compiler_info to radv_pipeline_get_shader_key()

  • radv: store the number of PS params heuristic to radv_compiler_info

  • radv: re-introduce DGC+multiview support and enable it for vkd3d-proton only

  • radv: fix a potential NULL pointer dereference when emitting VBOs

  • radv: remove an useless check when emitting the index buffer

  • radv: only emit the “normal” index buffer when needed with DGC

  • radv: stop dirtying some states after DGC execute

  • radv: cleanup invalidating vertex draw state

  • radv: replace use_ngg_streamout by gfx_level checks

  • ac,radv,radeonsi: replace mesh_fast_launch_2 by gfx_level checks

  • vulkan: add missing VkMemoryRangeBarriersInfoKHR support

  • radv: add missing VkMemoryRangeBarriersInfoKHR from DAC

  • radv: simplify resetting pipeline state for ESO

  • radv: rename RADV_CMD_DIRTY_PIPELINE to RADV_CMD_DIRTY_GRAPHICS_PIPELINE

  • radv: stop tracking the last emitted graphics pipeline

  • radv: add RADV_CMD_DIRTY_COMPUTE_PIPELINE

  • radv: add RADV_CMD_DIRTY_RAY_TRACING_PIPELINE

  • radv: remove useless tracking about non-coherent RBs with secondaries

  • radv: slightly rework initializing the default graphics state

  • ci: bump libdrm to 2.4.133

  • meson: bump required libdrm to 2.4.133 for AMDGPU

  • radv: move suspend_streamout to radv_streamout_state

  • radv: move streamout bindings to radv_streamout_state

  • radv: remove unnecessary radv_cmd_state::mesh_shading

  • radv: move index buffer state to radv_index_buffer_state

  • radv: cleanup suspending/resuming cond rendering with DGC

  • radv: move conditional rendering state to radv_cond_render_state

  • radv: move vertex buffer state to radv_cmd_state

  • radv: re-organize radv_cmd_state slightly

  • radv/ci: bump timeouts for radv-{navi21,gfx1201}-vkcts-full

  • ac/gpu_info: store more addr space info

  • ac/gpu_info: add has_smem_with_null_prt_bug

  • ac/gpu_info: query the PRT workaround control bit from libdrm

  • ac/nir: add a pass to fixup SMEM loads with NULL PRT pages

  • radv: run the pass to fixup SMEM loads with NULL PRT pages

  • radv: use the “LOW” address space for UBOs

  • radv/amdgpu: emulate sparse residency for the SMEM loads with NULL PRT workaround

  • radv: set RADEON_FLAG_EMULATE_SPARSE_RESIDENCY for sparse SSBO/UBO buffers

  • radv/ci: update list of skipped tests

  • ci: uprev vkd3d

  • radv: fix printing image format with RADV_DEBUG=img

  • radv/meta: fix expanding HTILE on compute with multisampling

  • docs: describe the contributions workflow for RADV

  • radv: bump VkConformanceVersion to 1.4.5.3

  • radv: fix determining needed dynamic states when rasterization is disabled

  • radv: make optimalTilingLayoutUUID driver and chip specific

  • ac/surface: allow to select hybrid/block memcpy path for host copies

  • radv: take advantage of VK_HOST_IMAGE_COPY_MEMCPY_BIT

  • vulkan: replace VK_SHADER_CREATE_INDEPENDENT_SETS_BIT_MESA with the maint11 flag

  • vulkan: stop forcing independent sets for shader object

  • ci: uprev vkd3d

  • radv/tests: add tests for global pipeline keys compatibility

  • radv: allow DGC+multiview by default

  • radv: fix an assertion with RADV_DEBUG=fullsync on GFX11+

  • radv: do not fallback to compute for image->buffer copies with emulated formats

  • spirv: preserve the explicit stride for untyped pointers with matrices

  • radv: add support for VK_SHADER_CREATE_INDEPENDENT_SETS_BIT_KHR

  • radv: adjust minImageTransferGranularity for transfer queue

  • radv: advertise VK_KHR_maintenance11

  • radv: fix another case of VRS with mipmaps on GFX10.3

  • radv: remove a TODO about layeredShadingRateAttachments

  • radv: invalidate command buffer state after executing secondaries

  • radv/meta: adjust an assertion for HTILE expand on SDMA with compute fallback

  • radv: clear the follower gang semaphore when a cmdbuf is reset

  • radv: destroy the gang CS when a cmdbuf is reset

  • radv: fix copying acceleration structure with DAC

  • nir: fix shuffling local IDs for quad derivatives with larger workgroup sizes

  • radv: enable radv_wait_for_vm_map_updates for Forza Horizon 6

  • radv: advertise VK_EXT_device_fault by default

  • radv: remove an outdated comment in radv_GetDeviceFaultInfoEXT()

  • radv: move radv_GetDeviceFaultInfoEXT() to radv_device.c

  • util: pass a struct to driParseConfigFiles()

  • util: do not generate drirc options that shouldn’t be parsed

  • util: fix declaring drirc options as string

  • radv: rename few drirc options for consistency

  • radv: use the new generation script for drirc

  • radv/ci: cleanup list of expected failures

  • util: add very basic way to validate drirc files

  • radv: validate drirc option names at compile time

  • radv: close the local fd immediately after the winsys is created

  • radv: rename radv_zero_vram to vk_zero_vram

  • radv: use radv_device::ws directly for quering sync payloads

  • radv: pre-compute a mask of supported global queue priorities

  • radv: add a separate function to query allocated/usage for each heap

  • radv: remove declared but unused create_null_physical_device()

  • radv/amdgpu: simplify syncobj verifications during submissions

  • radv: determine supported syncobj types directly in the physical device

  • radv/amdgpu: fix releasing the mutex for virtio and RADV_PERFTEST=localbos

  • nir: add new intrinsics for SPV_KHR_abort

  • spirv: implement SPV_KHR_abort

  • nir: add nir_lower_abort

  • radv: close the local fd slightly later when enumerating physical devices

  • radv: remove useless checks when creating a physical_device

  • radv: rename master_fd to wsi_master_fd

  • util: share the DOCTYPE for all driconf files

  • util: add a separate file for Zink drirc

  • util: add a separate file for RadeonSI drirc

  • util: add a separate file for turnip drirc

  • util: add a separate file for ANV drirc

  • util: add a separate file for NVK drirc

  • util: add a separate file for r300 drirc

  • util: add a separate file for iris drirc

  • util: add a separate file for asahi drirc

  • util: add a separate file for asahi vulkan drirc

  • util: add a separate file for panvk drirc

  • util: add a separate file for panfrost drirc

  • util: add a separate file for crocus drirc

  • util: add a separate file for dozen drirc

  • util: add a separate file for virgl drirc

  • util: add a separate file for r600 drirc

  • util: add a separate file for msm drirc

  • util: add a separate file for d3d12 drirc

  • util: add a separate file for vmgfx drirc

  • util: add a separate file for v3d drirc

  • util: add a separate file for hasvk drirc

  • util: remove useless comments in 00-mesa-defaults.conf

  • util,turnip: move drirc entries with vk_dont_care_as_load to Turnip

  • util,asahi: move drirc entries with no_fp16 to asahi

  • util: remove declared but unused drirc options

  • ci: adjust time-trace.sh to not exceed the limit of 255 chars

  • radv/ci: fix list of expected failures

  • radv/amdgpu: rework tracking allocated memory for budget

  • radv/amdgpu: stop deduplicating winsys

  • ac/nir,radv: lower task payload to zeroes when the mesh shader has no task

  • radv: enable radv_force_64_byte_sampled_image for Crimson Desert

  • radv: cleanup conditional header includes

  • radv: implement VK_KHR_device_fault

  • radv: advertise VK_KHR_device_fault

  • radv: fix DGC with conditional rendering and task+mesh shaders

  • radv/amdgpu: allow RADV_PERFTEST=localbos with virtio

  • radv: cleanup occurrences of radeon_info::has_vm_always_valid

  • radv: return VK_ERROR_INITIALIZATION_FAILED if VM_ALWAYS_VALID isn’t supported

  • ci: uprev vkd3d

  • util: remove declared but unused DRIC_CONF_VK_REQUIRE_ASTC

  • util/drirc_gen: add a function to declare commmon VK options

  • radv: declare common VK drirc options using the helper

  • anv: declare common VK drirc options using the helper

  • turnip: declare common VK drirc options using the helper

  • radv/rt: fix a memory leak with hash tables

  • radv/rt: fix a memory leak with the RT prolog NIR

  • radv/rt: fix a memory leak with ahit/isec group

  • radv: fix a memory leak with perfcounters

  • radv/amdgpu: destroy the BO for the NULL PRT workaround earlier

  • vulkan: Update spec to 1.4.353

  • ci: uprev vkd3d

  • ci/vkd3d: add support for running with ASAN

  • radv/ci: run vkd3d jobs with ASAN by default

  • radv: add the mesh scratch ring BO to the preambles BO list

  • radv: handle errors correctly when creating gang waits

  • aco: emit nir_jump_halt

  • radv: implement VK_KHR_shader_abort

  • radv: advertise VK_KHR_shader_abort

  • radv/amdgpu: defer allocating the NULL PRT BO

  • radv/ci: skip all WSI tests on GFX1201

  • util/drirc_gen: change the driconf DTD to not require one app/engine entry

  • util/drirc_gen: fix generating 64-bit driconf options

  • util/drirc_gen: prevent generating empty structs

  • util/drirc_gen: add heap_memory_percent to common VK options

  • util/drirc_gen: allow to override the defaults VK WSI common options

  • dzn: use drirc_gen

  • pvr: use drirc_gen

  • panvk: use drirc_gen

  • hk: use drirc_gen

  • v3dv: use drirc_gen

  • venus: use drirc_gen

  • radv,anv: remove useless includes for drirc stuff

  • util/drirc: remove the driver option in drirc_validate

  • ci: add a new option called profile in ci_run_n_monitor.py

  • radv/ci: skip all WSI tests also on NAVI21/NAVI31

  • util: remove useless entries for Intel hasvk

  • hasvk: use drirc_gen

  • ac/video: drop an useless drm_minor check

  • ac/descriptors: fix setting CB_COLOR_ATTRIB3.RESOURCE_LEVEL

  • ac/gpu_info: only initialize has_desc_resource_level on GFX10+

  • nir/print: add a missing UNREACHABLE for unknown jump instructions

  • nir: add a new nir_jump_abort

  • nir,aco: use nir_jump_abort instead of nir_jump_halt for abort

  • radv/ci: add more flakes for RAPHAEL

  • radv/ci: update the list of expected failures for NAVI10

  • radv: prevent closing the render node fd twice for AMD_FORCE_VPIPE=1

  • radv: allow to query GPU info without creating a winsys

  • radv/amdgpu: add a function to query heap info

  • radv: query heap info without using the winsys

  • radv: duplicate the fd used for syncobj with KHR_display

  • radv: create one winsys for each logical device

  • radv: fix REPLAYED shader arena blocks not being marked as holes on free

  • radv/amdgpu: fix padding by one VM page

  • vulkan: fix lowering untyped accel struct with descriptor heap

  • spirv: mark UBO/SSBO array accesses as always in-bounds

  • radv: fix a synchronization issue with taskmesh and pending cache flushes

  • radv: fix a synchronization bug with DGC preprocess and taskmesh

  • radv: clear gang cache flushes when the command buffer is reset

  • vulkan: fix incorrect sType for VkDebugUtilsObjectTagInfoEXT

  • radv: workaround game bugs with Sniper Elite 5

  • spirv: allow mapping readonly buffers with struct members

  • radv: fix clearing the streamout state on GFX12

  • radv: remove redundant memory initialization on the CPU

  • vulkan: fix initializing address flags

  • vulkan: add vk_buffer_usage_flags()

  • radv: cleanup pCreateInfo uses for VkBuffer

  • radv: cleanup pCreateInfo uses for VkImage

  • radv: allow ptr to be NULL in radv_cmd_buffer_upload_alloc()

  • radv/amdgpu: fix computing allocated VRAM for imported BOs from fd

  • radv: fix a memleak with embedded samplers and descriptor heap

  • radv: store copying embedded samplers for heap to the shader layout

  • radv: enable VK_EXT_descriptor_heap by default

  • radv: disable VRS with MSAA 8x on GFX11-11.7 to prevent GPU hangs

  • radv: disable VRS for flat shading with MSAA 8X to prevent GPU hangs on GFX11

  • radv: disable VRS with MSAA 8x also on GFX10.3

  • radv: remove the deprecated warning for RADV_FORCE_FAMILY

  • docs,radv: auto-generate driconf documentation from drirc_gen

  • util: remove declared but unused driconf vulkan-related options

  • vulkan,anv,radv: do not crash when querying descriptor size for unsupported type

  • zink: fix a memleak in zink_init_format_props()

  • glsl: fix a memleak in link_assign_subroutine_types()

  • vulkan: search VkImageViewUsageCreateInfo in pNext

  • vulkan: implement VK_KHR_extended_flags

  • vulkan/wsi: implement VK_KHR_extended_flags

  • zink/ci: update lists for RADV

  • loader: fix a memleak

  • kopper: fix a memleak

  • zink: fix a memleak with fences

  • zink: fix a memleak with the emulated GS NIR shader

  • zink: fix a memleak with sampler state

  • radv: use the image view usage for MSRSTT transient iviews

  • radv: implement VK_KHR_extended_flags

  • radv: advertise VK_KHR_extended_flags

  • ci: apply patches to fix memleaks for GL/GLES CTS

  • zink/ci: add a new job for NAVI31 with ASAN enabled

  • pipe-loader: fix a global-buffer-overflow ASAN error when getting driconf

  • radv/meta: fix restoring descriptor heaps

  • Revert “spirv: allow mapping readonly buffers with struct members”

  • radv: always consider some outputs as invariant

  • radv: remove radv_invariant_geom driconf option

  • zink/ci: update trace checksums

  • radv: force late-Z with fragment shaders that use fbfetch

  • spirv: fix handling OpAbortKHR

  • radv: fix invalid assertions in DGC when queues aren’t enabled

Serdar Kocdemir (14):

  • gfxstream: add gitignore for generated code

  • gfxstream: Add VK_EXT_pipeline_protected_access

  • gfxstream: allow VK_KHR_maintenance extensions

  • gfxstream: some cleanup on device extension allow list

  • Set driver ID for gfxstream

  • gfxstream: allow VK_GOOGLE_display_timing

  • gfxstream: remove android conditioning for sampler extensions

  • gfxstream: use VK_DRIVER_ID_MESA_GFXSTREAM as driver id

  • gfxstream: update codegen for host side vulkan header update to v1.4.350

  • gfxstream: Fix codegen causing missing vulkan structures

  • gfxstream: disallow maintenance6 extension due to serialization bugs

  • gfxstream: correctly ignore timeline semaphore info

  • gfxstream: allow VK_EXT_border_color_swizzle

  • gfxstream: check in auto-generated guest code

Sergi Blanch Torne (15):

  • ci: disable Collabora’s farm due to maintenance

  • Revert “ci: disable Collabora’s farm due to maintenance”

  • ci: disable Collabora’s farm due to maintenance

  • Revert “ci: disable Collabora’s farm due to maintenance”

  • xfiles: update before uprev

  • ci: disable Collabora’s farm due to maintenance

  • Revert “ci: disable Collabora’s farm due to maintenance”

  • ci: review initial ANGLE flakes

  • ci,crnm: information from pipeline url

  • ci,crnm: search MR pipelines in forks

  • ci,crnm: handle exception when auth fail

  • xfiles: update before uprev Piglit

  • xfiles: update expectations based on 2026-7-8 nightly

  • ci: disable Collabora’s farm due to maintenance

  • Revert “ci: disable Collabora’s farm due to maintenance”

Sergi Blanch-Torne (1):

  • ci,crnm: bugfix project default

Sergio Sanchez Valencia (1):

  • d3d12/wgl: reclaim deferred BOs before ResizeBuffers

Shih, Jude (4):

  • amd/vpelib: Alpha blending enhancement

  • amd/vpelib: Refactor DPP function table layout

  • amd/vpelib: Fix Compiler Warnings

  • amd/vpelib: Realign DPP callback initialization with the updated interface layout

Sid Pranjale (12):

  • nvk: Implement VK_EXT_shader_atomic_float

  • vulkan: implement VK_EXT_debug_marker

  • v3dv: drop legacy CPU queue fallback paths

  • nak/nir: lower f16vec2 shared atomics

  • v3dv: replace single-field options struct with bool

  • v3dv: directly use v3d_has_feature instead of caps struct

  • broadcom/common: add multisync helpers

  • gallium/v3d: use common multisync code

  • v3dv: use common multisync code

  • v3dv: implement CPU-side fence merging for queue signaling

  • v3dv: simplify queue submission

  • v3dv: remove unused no-op job allocation setup

Silvio Vilerino (18):

  • d3d12: Create PIPE_BIND_SHARED resources with D3D12_RESOURCE_FLAG_ALLOW_SIMULTANEOUS_ACCESS

  • mediafoundation: Create readable dpb buffers with PIPE_BIND_RENDER_TARGET and PIPE_BIND_SHARED for DX11 sharing

  • Revert “d3d12: Video sliced encode: Use same ID3D12Fence/different per slice values as optimization”

  • d3d12: Flush stale video encode wait registrations when reusing ID3D12Fence objects

  • d3d12: Support video encode AUTO slice/tile only capable hardware

  • mediafoundation: check for AUTO slice/tile only capable hardware

  • d3d12: Use sequential video enc subregion signaling

  • d3d12: d3d12_create_fence_raw to lazily register fence event on waits

  • d3d12/video: fix comparison-with-wider-type warnings

  • d3d12: avoid signed integer overflow in copy staging box setup

  • util: u_trace.c: Fix error C4189: buffer_count: local variable is initialized but not referenced

  • d3d12: Fix NULL dereference check in d3d12_video_buffer_destroy

  • d3d12: Use res_device to import resource from different device via handle

  • pipe: Expose new fence_wait_multiple operation

  • d3d12: Implement fence_wait_multiple with SetEventOnMultipleFenceCompletion

  • mediafoundation: Use eventless fence_wait_multiple instead of WaitForMultipleObjects

  • d3d12: Remove event cleanup since d3d12_fence now uses lazy SEOC

  • d3d12: Check fence values before SEOC in d3d12_fence_wait_multiple

Simon Perretta (29):

  • pco: reserve additional outputs for trilinear sampled coeffs

  • pco: amend tg4 lowering

  • pco: track how many tg4/raw sample comps are needed

  • pvr: consider barriers when calculating compute instances

  • pvr, pco: add support for spilling shared memory to global memory

  • pvr, pco: store device runtime info in compiler context

  • pco: conditionally spill shared memory to global memory

  • pco, pvr: finish and enable VK_KHR_workgroup_memory_explicit_layout

  • pco: drop global path for null descriptor checking

  • pco: add mappings for setl, savl ops

  • pvr, pco: add “real” basic subgroup support

  • pco: handle mov offset special regs

  • pco: add support for read_invocation via shared memory

  • pco: add subgroup ballot support via shared memory

  • pvr: advertise VK_EXT_shader_subgroup_ballot and ballot feature

  • pco: add br.skip_next op

  • pco: commonize execution mask counter ref helper function

  • pco: add support for subgroup vote_{all,any} ops

  • pvr: advertise VK_EXT_shader_subgroup_vote and vote feature

  • pco: add support for reduce/scan ops with cluster awareness

  • pvr: advertise subgroup arithmetic and clustered features

  • pco: add support for subgroup shuffle ops

  • pvr: advertise subgroup shuffle and shuffle relative features

  • pvr, pco: add support for VK_KHR_shader_subgroup_rotate

  • pvr: advertise VK_KHR_shader_subgroup_uniform_control_flow

  • pvr, pco: advertise support for VK_EXT_subgroup_size_control

  • pco: allow non-pure integer formats for image xchg atomics

  • pco: lower sysvals early for fragment shaders

  • pco: allow fence ops to be legalized if they come last in a block

Skyth (1):

  • spirv2dxil: Replace UAV_FENCE_THREAD_GROUP usage with UAV_FENCE_GLOBAL.

Sonny Jiang (2):

  • radeonsi: always set is_format_supported in screen create

  • radeonsi/vcn: Add vcn_5_0_2 support

Stijn Tintel (1):

  • rocket: fix mmap leak in buffer map/unmap

Stéphane Cerveau (1):

  • vulkan/video: Reject interlaced picture layout for H.264 baseline profile

Suresh Guttula (1):

  • ac: Add vcn_5_3_0 support

Sushma Venkatesh Reddy (2):

  • intel/perf: Add WCL OA support

  • intel/dev: Clamp PTL+ CS workgroup threads to 32

Tacodiva (1):

  • vulkan/runtime: Fix bad assumption in GetPipelineBinaryDataKHR

Tanner Van De Walle (5):

  • draw: add lower-bound assert on shader_stage

  • gallium/u_blitter: add lower-bound assert on target

  • util/format: add lower-bound assert on format

  • dzn: silence PREfast C33010 warnings

  • nir/nir_builder: inline dst_bit_size calculation in assert

Tapani Pälli (13):

  • intel/compiler: implement macl part of Wa_18035690555

  • drirc: use anv_disable_drm_ccs_modifiers for any GTK version

  • drirc/anv: add flag to disable VK_EXT_subgroup_size_control

  • drirc: set anv_disable_subgroup_size_control for bg3

  • anv: do not use resource barrier with split barriers

  • intel/dev: update mesa_defs.json from workaround database

  • iris: use INTEL_NEEDS_WA_14025112257 define for workaround

  • anv: use INTEL_NEEDS_WA_14025112257 define for workaround

  • anv: allocate tile sized temporary copy instead of whole size

  • anv: skip writing xfb buffer if we get null information

  • iris: align down the max_shader_buffer_size

  • anv: fix a null pointer access with isl_mod_info

  • anv: optimization for Wa_14025112257 case

Thomas H.P. Andersen (5):

  • nvk: set queryResultStatusSupport

  • nvk: use the new generation script for drirc

  • nouveau/cubin: use libelf 64 bit instead of gelf

  • nvk: hide NVX_binary_import behind NVK_EXPERIMENTAL=dlss env var

  • nvk: add env var to allow backwards compat in dlss

Thong Thai (25):

  • util: move u_stub to src/util, add u_stub_gfx_compute.h

  • util: allow for overriding u_stub tail

  • meson: update default build option for libva subproject

  • meson: check if video encoding support is to be built

  • frontends/va: decode only stubs

  • radeonsi: move si_get video functions to si_video

  • amd: make ac_ib_parser an amd tool build option

  • gallium/auxiliary/vl: Fix typo in cs_create_shader pseudo-code comment

  • pipe: Add PIPE_VIDEO_VPP_BLEND_MODE_PREMULTIPLIED_ALPHA

  • vl/video_buffer: Set alpha swizzle if format has alpha

  • gallium/video: Add enabled flag to vpp interface

  • gallium/vl: Implement compositor shader-based alpha blending

  • frontends/va: Enable shader-based alpha blending

  • amd: Build nir files only when with_gfx_compute

  • radeonsi: Remove ACO dependency for non-GFX/compute builds

  • nir: Only build NIR headers when with_gfx_compute is false

  • gallium/auxiliary: Simplify auxiliary for non-gfx/compute builds

  • meson: Make with_gfx_compute depend on video encode support

  • meson: Don’t require libelf for radeonsi when with_gfx_compute is false

  • radeonsi: Allow call to stub’d si_init_gfx_context to continue

  • radeonsi: Store SQTT cb_id

  • radeonsi: Handle SQTT timestamps

  • radeonsi: Store SQTT device_id

  • radeonsi: Implement SQTT CB_START and CB_END

  • radeonsi: Setup SQTT sampling clocks

Timothy Arceri (17):

  • glcpp: update out of date comment

  • glcpp: fix paste within macro function expansion

  • amd/radeonsi: dont clamp packed user varyings

  • mesa: fix typo in validation string

  • ac/nir/lower_tex_coord: update cursor when moving wqm coordinates

  • ac/nir/lower_tex_coord: basic lower tex coord test

  • mesa: flush bitmap cache when scissor box changes

  • nir: use the correct induction var when guessing loop iterations

  • glsl: allow uniform block layout qualifiers when SSBO enabled

  • util/u_range_remap: allow insert to truncate range

  • glsl: treat temp globals wrappers as roots when resolving function calls

  • zink: fix swap interval changes being dropped

  • nir/opt_dead_write_vars: handle memcpy_deref as reads

  • util: add Blockland workaround for crash

  • util/mesa: add workaround to zero invalidated buffers

  • util: add workaround for Riddick using round() in glsl 1.20

  • llvmpipe: emit FS input vertex attributes in driver location order

Timur Kristóf (10):

  • nir/divergence: Consider ACCESS_SMEM_AMD divergence across subgroups

  • nir/divergence: Consider uniformity of read_invocation accross subgroups

  • nir/divergence: Consider ttmp_register_amd and load_scalar_arg_amd as workgroup divergent

  • ac/nir: When loading an arg, assert that it’s used

  • ac/nir: Fix SMEM workaround with emulated RT

  • radv: Wait for idle after every submission on GFX6-7

  • radeonsi: Wait for shaders and flush L2 after every submission on GFX6-7

  • Revert “radv: Mitigate GPU hang on Hawaii in Dota 2 and RotTR”

  • ac/nir/ngg: Remember if a mesh shader has non-API waves.

  • ac/nir/ngg: Use workgroup divergence analysis for mesh output counts.

Tomeu Vizoso (5):

  • teflon/tests: avoid loading build-tree tensorflow-lite stub at runtime

  • teflon/tests: make tflite stubs fail loudly with diagnostics

  • teflon: remove synthetic model generation and flatbuffers dependency

  • ci: Remove flatbuffers from builds

  • teflon/tests: Remove leftover files from synthetic tests

Toshinari Morikawa (2):

  • virgl: fix memory leak on shader translation

  • egl: avoid calling loader_get_driver_for_fd with fd = -1

Trigger Huang (12):

  • radv: supports protected memory allocation

  • radv: allow creation of protected queues

  • radv: support secure submission

  • radv: add protected type bits for memory requirements

  • radv: enable protected memory

  • radv: emulate MSRTSS via implicit MSAA resolve

  • radv/meta: thread separate src/dst sample counts through gfx copy

  • radv/meta: derive gfx copy dst sample count from the destination

  • radv/meta: add MSRTSS attachment replicate helper

  • radv: replicate MSRTSS attachments on LOAD_OP_LOAD

  • radv: handle VkSubpassResolvePerformanceQueryEXT

  • radv: enable VK_EXT_multisampled_render_to_single_sampled

UMU618 (1):

  • venus: fix typo in vn_queue_submit_2_to_1

UMUTech (1):

  • wsi: correct the erroneous assertion

Utku Iseri (1):

  • v3dv: close display_fd on incompatible_driver path

Val Packett (3):

  • util: rust: align API with real eventfd capabilities

  • util: rust: Support detecting socket file descriptors

  • util: rust: Add a way to create a Tube from an existing OwnedFd

Valentine Burley (102):

  • mr-label-maker: Label Collabora farm with driver tags

  • ci/zink/intel: Disable flaky TGL canvas_moire-v2 trace

  • zink/ci: Document recent flakes

  • anv/ci: Revert ADL VKCTS job to stable 6.17 kernel

  • zink/ci: Move Turnip flakes to correct list

  • tu/drm/virtio: Fix tu_wait_fence timeout handling

  • freedreno/drm/virtio: Fix wait_fence ret ordering

  • zink/ci: Remove Cezanne job

  • radv/ci: Add more ASAN VKCTS jobs on Cezanne

  • vulkan/android: Add deferred image helper

  • panvk: Use vk_android deferred image helper

  • vulkan/android: Add vk_android_import_anb_memory helper

  • vulkan: Query memory requirements in vk_android_import_anb_memory

  • tu: Implement deferred image creation for ANB and AHB

  • ci/crosvm: Sanitize CROSVM_RET in crosvm-runner.sh

  • tu: Fix D16 depth clear rounding mismatch in sysmem mode

  • panfrost/ci: Update kernel to pick up ZSTD support for ZRAM

  • venus/ci: Skip more robustness tests on ANV

  • tu: Move Android extensions into main list

  • tu: Add shared image support on Android

  • panfrost/ci: Move t860 jobs to nightly

  • panfrost/ci: Document recent g610 flake

  • ci/android: Remove SurfaceFlinger wait in get_surfaceflinger_pid

  • ci/android: Fix intermittent adb root failures

  • ci/android: Update Cuttlefish build

  • ci/deqp: Add Android WSI support

  • lavapipe/ci: Enable WSI testing on Android

  • turnip/ci: Enable WSI testing on Android

  • venus/ci: Enable WSI testing on Android

  • ci/android: Remove CtsDeqpTestCases from Android CTS

  • ci/deqp: Backport host_image_copy fix

  • ci/lava: Reduce LAVA job timeout to 20 minutes for Marge

  • mr-label-maker: Add rule for new trace replay config files

  • ci: Add missing rule for new trace replay config files

  • tu/autotune: Clear active_batches before history objects are freed

  • ci/deqp: Backport validation error fix

  • ci/deqp: Backport landed patch

  • ci/deqp: Rewrite headless Android WSI patch

  • venus/ci: Skip more even more robustness and synchronization2 tests on ANV

  • ci: Disable debian-riscv64

  • panvk/ci: Mark dEQP-VK.subgroups.* as flaky on G925

  • ci: Bump ci-deb-repo revision to update aapt

  • ci/android: Update Android CTS to android-cts-16.0_r5

  • ci/android: Add arm64 support for Android CTS

  • turnip/ci: Add nightly Android CTS job

  • tu: Disable -Wmisleading-indentation when compiling with GCC

  • panvk: Fix ignored qualifier warnings

  • meson: Add Soong compatibility compiler flags to Vulkan drivers

  • tu/kgsl: Fix memory type support detection for unsupported flags

  • tu: Merge tu_image_init and tu_image_update_layout

  • venus/ci: Widen the ANV skips

  • tu: Advertise VK_KHR_internally_synchronized_queues

  • turnip/ci: Update ANGLE trace checksum

  • panfrost/ci: Switch traces over to gpu-trace-perf

  • tu: Fix vk_queue leak on submitqueue creation failure

  • vulkan/queue: Add common queue emulation support

  • tu: Emulate second graphics queue for skiavk on Android

  • tu/ci: Add coverage for emulated second graphics queue

  • ci/android: Update Cuttlefish build

  • anv/ci: Disable anv-adl-vk job

  • venus/ci: Move pre-merge ANV coverage from Comet Lake to Alder Lake

  • venus/ci: Retire Intel Comet Lake runner

  • venus/ci: Revert ADL jobs to stable 6.17 kernel

  • intel/gen: Explicitly declare gen_opcodes_private.h dependency

  • perfetto: Centralize perfetto header include in u_perfetto.h

  • bin: Expose drm-shim in meson devenv

  • doc/ci: Add drm-shim CI reproduction guide

  • zink/ci: Increase zink-lavapipe parallelism

  • virgl/ci: Retire disabled virgl-iris jobs

  • virgl/ci: Retire disabled android-virgl-llvmpipe

  • pipe-loader: Enable null winsys on Android

  • zink/ci: Remove zink-anv-cml-asan job

  • intel/ci: Remove nightly CML jobs, retire runner

  • intel/ci: Increase iris-apl-egl parallelism

  • zink/ci: Drop old VVL filters

  • panfrost/ci: Fix typo in .panfrost-vk-manual-panthor-rules template

  • tu: Fix uninitialized gmem_offset when a GMEM layout is impossible

  • ci: Update kernel to Linux 7.1.2

  • panfrost/ci: Use Linux 7.1 kernel for more jobs

  • tu: Fix capture/replay with sampler custom border color

  • ci/lava: Uprev lava-job-submitter

  • drm-shim/freedreno: Add support for Adreno 610

  • drm-shim/freedreno: Shim perf counter config ioctl

  • bin/drm-shim: Add more freedreno GPUs

  • tu: Disable VK_EXT_extended_dynamic_state2 patch control points on A702

  • zink: Gate tess/geom barrier stages on feature support

  • ci: Bump ci-deb-repo revision to update vulkan-loader

  • ci/deqp: Backport landed Android WSI patch

  • ci/deqp: Update VK CTS to 1.4.6.1

  • ci/deqp: Backport -frounding-math default for GCC builds

  • tu: Implement VK_EXT_primitive_restart_index

  • ir3: Fix ballot_components for subgroups smaller than 32

  • ir3: Derive max_variable_workgroup_size from device limits

  • ir3: Compute subgroup_size from threadsize_base

  • tu: Use computed subgroup size

  • panvk/ci: Update expectations for g610-vk-asan following VK CTS uprev

  • panvk: Disable VK_KHR_internally_synchronized_queues on Vulkan 1.0

  • tu: Report correct maxFragmentInputComponents limits

  • zink: Use ShaderLayer capability for gl_Layer when available

  • freedreno: Increase reg_size_vec4 for A702

  • freedreno: Fix VPC_RAST_STREAM_CNTL register layouts

  • ci/piglit: Switch all trace jobs to surfaceless+gbm

Vincent Cloutier (2):

  • etnaviv: use buffer resource accessor for indirect draws

  • etnaviv: support native bitfield extract/reverse/count ops

Vinson Lee (14):

  • st/mesa: fix implicit conversion warning in st_atom_framebuffer

  • vulkan/screenshot-layer: initialize info to NULL

  • gfxstream: codegen: drop const from let-param scalar cast

  • mesa/main: cast GLhandleARB to unsigned int in api trace

  • ethosu/mlw_codec: silence warnings in the vendored Regor encoder

  • ethosu/mlw_codec: silence -Wunused-const-variable in vendored encoder

  • ethosu: use FALLTHROUGH macro in ethosu_emit_operation_accesses

  • radeonsi: remove duplicate ‘.bpp’ initializer in si_sdma_copy_image

  • gfxstream: link goldfish_address_space against perfetto

  • util/tests: replace sprintf with snprintf in cache tests

  • util/tests: fix unused variable warnings in cache List test

  • vulkan/screenshot-layer: replace itoa/sprintf with snprintf

  • vulkan/screenshot-layer: fix globalLock mutex leak

  • util/tests: silence unused iterator warning in sparse_bitset_test

Virgile Bello (3):

  • microsoft/compiler: sink load_invocation_id in TCS split even for single-use

  • microsoft/compiler, d3d12: flip tess winding at caller, not in nir_to_dxil

  • microsoft/compiler, d3d12: preserve TCS outputs and pad TES inputs for cross-stage signature matching

Vishnu Vardan (27):

  • mesa/st: remove redundant has_stencil_export from st_context

  • mesa/st: remove redundant astc_void_extents_need_denorm_flush from st_context

  • mesa/st: remove redundant has_shareable_shaders from st_context

  • mesa/st: remove redundant needs_texcoord_semantic from st_context

  • mesa/st: remove emulate_gl_clamp from st_context

  • mesa/st: remove has_time_elapsed from st_context

  • mesa/st: remove has_multi_draw_indirect from st_context

  • mesa/st: remove has_indirect_partial_stride from st_context

  • mesa/st: remove has_occlusion_query from st_context

  • mesa/st: remove has_single_pipe_stat from st_context

  • mesa/st: remove has_pipeline_stat from st_context

  • mesa/st: remove has_indep_blend_enable from st_context

  • mesa/st: remove has_indep_blend_func from st_context

  • mesa/st: remove can_dither from st_context

  • mesa/st: remove lower_flatshade from st_context

  • mesa/st: remove lower_alpha_test from st_context

  • mesa/st: remove lower_two_sided_color from st_context

  • mesa/st: remove lower_ucp from st_context

  • mesa/st: remove prefer_real_buffer_in_constbuf0 from st_context

  • mesa/st: remove has_conditional_render from st_context

  • mesa/st: remove lower_rect_tex from st_context

  • mesa/st: remove allow_st_finalize_nir_twice from st_context

  • mesa/st: remove can_bind_const_buffer_as_vertex from st_context

  • mesa/st: remove validate_all_dirty_states from st_context

  • mesa/st: remove can_null_texture from st_context

  • mesa/st: remove redundant has_hw_atomics from st_context

  • anv/rt: reorder encode_internal_node to only process valid children

Vlad Zahorodnii (1):

  • wsi/wayland: Add support for wl_fixes.ack_global_remove

Wig Cheng (2):

  • rocket: pad weight packing input channels to FEATURE_ATOMIC_SIZE

  • rocket: compute element-wise ADD requant instead of LUT

Wujian Sun (2):

  • mesa: Fix clipping order in _mesa_clip_blit()

  • mesa: Allow GL_SRGB_ALPHA_EXT as color-renderable when EXT_sRGB is supported

Xinju Li (1):

  • nir: resolve functions: only resolve functions that are reachable from main

Yannis Juglaret (1):

  • nouveau: fix data race in nouveau_fence_ref

Yiwei Zhang (146):

  • venus: adopt vk_android_init_deferred_image

  • venus: adopt vk_android_get_ahb_layout

  • venus: refactor vn_android_get_wsi_memory to return VkDeviceMemory

  • venus: adopt common vk_image::anb_memory

  • venus: adopt common ANB helpers

  • panvk: adopt common ANB helpers

  • lvp/android: use common ANB implementations

  • util/android_stub: drop legacy atrace

  • panvk: drop panvk_android_create_deferred_image

  • util/os_misc: use ndk api __system_property_get

  • egl/android: use ndk api __system_property_get

  • android_stub: drop cutils/properties dependency

  • CODEOWNERS: update owners for Android components

  • util/os_misc: use stable NDK __android_log_write helper

  • intel: use stable NDK __android_log_print helper

  • broadcom: remove unused Android log utils

  • android_stub: purge unused log utils

  • ci: uprev virglrenderer

  • android_stub: fix update-android-headers.sh for libbacktrace

  • android_stub: fix libhardware source include path

  • android_stub: avoid vending in unused headers

  • android_stub: sync Android 16 headers

  • pan/nir/tex: use unsigned type for texture op lod_or_fetch

  • venus: fix a renderer side queue timeline bound race

  • panvk: fix to report device memory with heapIndex

  • tu: fix to report device memory with heapIndex

  • venus: update create_from_device_memory to take a cmd payload

  • venus: let resource_create_blob wait for mem alloc

  • venus: fix unbound malloc leak in vn_ring_get_submits

  • anv: fix lock scope in anv_ensure_fp64_shader

  • anv: amend missing shader dump finish upon device destruction

  • venus: amend roundtrip between fence submit and wait idle

  • panvk: fix plane indexing for afbc image subres layout

  • pan: handle downscaling of plane view for multiplanar yuv textures

  • panvk: reject interleaved_64k for multiplanar yuv

  • panvk: avoid separate reconstruction filter for YUV texturing

  • panvk: use vk_component_mapping_to_pipe_swizzle

  • panvk: apply YUV swizzle to the view swizzle for YUV texturing

  • panvk: override default chroma siting for YUV texturing

  • panvk: enforce strict import for yuv images

  • panvk: add get_pan_image_props helper

  • panvk: use binding layout textures_per_desc to write image view descs

  • panvk: add pan_texture_get_payload_alignment to help with tex emit

  • panvk: add and use panvk_image_get_tex_count helper

  • panvk: properly set up image and view planes for YUV texturing

  • panvk: lower YUV texturing to do SW CSC

  • pan: add 8bit multi-planar 420 and 422 format for Vulkan

  • panvk: enable 8bit multiplanar YUV formats on v9+ to v13

  • panvk: add P010 native YUV support

  • vulkan/android: force linear for mutable format

  • venus/wsi: skip VkPresentRegionsKHR when pRegions is NULL

  • venus: avoid touching sfb dst slot upon resume

  • venus: always check device lost on sfb warn order

  • venus: rename vn_semaphore_feedback_cmd to vn_sync_feedback_cmd

  • venus: move sfb helpers into vn_feedback

  • venus: wrap sfb cmd preparation with vn_sync_feedback_command

  • venus: extract sync feedback host write and query

  • venus: move sfb suspend resume handling over to vn_feedback

  • venus: add vn_sync_feedback_enabled

  • venus: always check device lost on ffb warn order

  • venus: migrate ffb to use vn_sync_feedback

  • venus: recycle fence sfb in post submission

  • venus: drop vk_xwayland_wait_ready override

  • v3dv: drop vk_xwayland_wait_ready

  • hasvk: drop vk_xwayland_wait_ready

  • vulkan/wsi/util: purge vk_xwayland_wait_ready

  • venus/virtgpu: amend a missing sim mutex init

  • venus/virtgpu: drop obsolete SIMULATE_BO_SIZE_FIX

  • venus/virtgpu: implicit fencing is gone

  • venus/virtgpu: simplify to drop virtgpu_sync

  • venus: refactor vn_renderer_submit to only take a single batch

  • venus/virtgpu: merge SIMULATE_SUBMIT into SIMULATE_SYNCOBJ

  • venus/virtgpu: drop signaled_fd

  • venus/virtgpu: drop cpu sync timeout

  • venus/virtgpu: drop wait available

  • venus/virtgpu: use STACK_ARRAY for syncobj handles

  • venus/virtgpu: split out sim_syncobj

  • venus/virtgpu: refactor sim_submit

  • venus/virtgpu: use uAPI to signal syncobjs

  • venus/virtgpu: adopt u_sync_provider

  • venus/virtgpu: use virtgpu syncobj on supported kernels

  • venus/virtgpu: avoid pretending timeline sync support

  • venus/virtgpu: sim_syncobj to implement u_sync_provider

  • venus/virtgpu: simplify sim_syncobj

  • venus/virtgpu: refactor virtgpu_submit

  • venus/virtgpu: flatten all the virtgpu_ioctl_syncobj wrappers

  • venus: properly clean up driver internal sim syncobj allocs

  • venus: fix imported sync fence payload reset upon export

  • venus: avoid renderer semaphore wait upon temp payload export

  • venus: refactor external fence and semaphore advertisment

  • venus: drop obsolete zink performance workaround

  • venus: track can_feedback in struct vn_queue_submission

  • venus: purge sync feedback for sparse binding

  • venus: split fence/semaphore/event commands to vn_sync.(c|h)

  • venus: refactor imported semaphore check and wait

  • venus/wsi: refactor args of vn_wsi_fence_wait and vn_wsi_flush

  • venus: queue submit to take a single batch

  • venus: sparse binding to use vn_queue_submit

  • venus: simplify vn_queue_submission to handle single batch

  • venus: simplify feedback cmd setup

  • venus: simplify vn_queue_submission_alloc_storage

  • venus: extract pNext chain fixup out from feedback cmds init

  • venus: prepare to handle sparse binding batch interception

  • venus: properly drop imported semaphores from submission

  • venus: relax SYNC_FD semaphore import requirement for WSI

  • venus/renderer: improve renderer backend init logs

  • venus: recycle idle sfb cmds only after async wait

  • venus: properly check sync2 enablement

  • venus: host image copy to scrub present_src layout if needed

  • venus: clean up sync2 treatment leftovers

  • venus: simplify sync feedback tracking

  • venus: update pnext fix tracking

  • venus: explicitly track if need to fix batch

  • venus: flatten sync feedback cmd counting

  • venus: deprecate fence feedback

  • venus: split ring submission to vn_queue_submission_do_submit

  • venus: skip empty batch submission

  • venus: drop vn_sync_payload_external from submission tracking

  • venus: extract vn_timeout_to_poll_timeout

  • venus/virtgpu: only signal non-zero initial value

  • venus/virtgpu/vtest: drop initial value from sync reset

  • venus: prepare for VN_SYNC_TYPE_SYNC

  • venus: migrate fence over to VN_SYNC_TYPE_SYNC

  • venus: relax SYNC_FD fence export requirement

  • venus: deprecate queue idle wait workaround

  • venus: rename to be explicit about sync fd semaphore

  • venus: vn_semaphore_(is|wait)_sync_fd to support VN_SYNC_TYPE_SYNC

  • venus: rename existing queue submission wait semaphore tracking

  • venus: count and prepare storage to scrub SYNC_FD signal semaphore

  • venus: extract syncs and scrub SYNC_FD signal semaphores

  • venus: ensure renderer sync fence is submitted between queue batches

  • venus: migrate SYNC_FD semaphore over to VN_SYNC_TYPE_SYNC

  • venus: relax SYNC_FD semaphore export requirement

  • venus/virtgpu: hide has_timeline_sync behind a new perf option

  • venus/vtest: advertise timeline syncobj support

  • venus: drop vn_renderer_sync_flags

  • venus: add VN_SYNC_TYPE_TIMELINE_SYNC

  • venus: track queue internal array index within vn_device::queues

  • venus: count and init syncs from timeline semaphores

  • venus: implement host signal for TIMELINE_SYNC

  • venus: implement counter query for TIMELINE_SYNC

  • venus: add vn_wait_semaphores_legacy for legacy wait

  • venus: implement semaphore wait for TIMELINE_SYNC

  • venus: migrate timeline semaphore to VN_SYNC_TYPE_TIMELINE_SYNC

  • venus: document timeline semaphore implementation

  • venus: ensure cached vn_ring_submit batches are bounded

Yogesh Mohan Marimuthu (5):

  • ac,radeonsi,radv: add has_desc_resource_level var instead of gfx_level check

  • radv: Program RESOURCE_LEVEL bit in descriptor for dgc

  • amd: add initial code for gfx1156

  • ac: set has_smem_with_null_prt_bug to false for gfx1156

  • ac: set has_desc_resource_level to true for gfx1156

You, Min-Hsuan (1):

  • amd/vpelib: fix FROD alignment handling after interface change

Zan Dobersek (5):

  • fd: add a8xx perfcntr countables

  • tu: only support userspace-managed perfcounters on a7xx and earlier

  • tu/a8xx: remove enforced TU_DEBUG_FLUSHALL

  • tu/kgsl: initialize dump bo state in kgsl_bo_init sooner

  • fd: lrz_block in fdl6_lrz_layout_init() should not be static

Zeyang Lyu (1):

  • radv: Add base array layer to htile offset

Zhao, Jiali (2):

  • amd/vpelib: revert predication fix

  • amd/vpelib: fix HDR external monitor video black content

ZhengMing (1):

  • vulkan/wsi/win32: Prefer the more popular surface format on Windows

Zoltán Böszörményi (1):

  • radv: Advertise msrtss in features.txt

adrian baker (4):

  • jay: add bfloat16 support

  • jay: fix jay bf16 comment formatting

  • nir: remove incorrect algebraic properties from intel mixed bf ops

  • jay: add geometry shader support.

gyeyoung (2):

  • panvk: fix flags2-only bit leak in legacy format features

  • panvk: report DRM format modifiers through List2EXT

gyeyoung baek (1):

  • rocket: drop wrong assert(input_op_1) in ADD fuse path

hmtheboy154 (15):

  • pvr: add support for driconf for the Vulkan driver

  • driconf: Add an option to override Vulkan’s deviceName

  • anv: driconf: Add an option to override Vulkan’s deviceName

  • hasvk: driconf: Add an option to override Vulkan’s deviceName

  • nvk: driconf: Add an option to override Vulkan’s deviceName

  • radv: driconf: Add an option to override Vulkan’s deviceName

  • venus: driconf: Add an option to override Vulkan’s deviceName

  • v3dv: driconf: Add an option to override Vulkan’s deviceName

  • lvp: add support for driconf

  • lvp: driconf: Add an option to override Vulkan’s deviceName

  • tu: driconf: Add an option to override Vulkan’s deviceName

  • panvk: driconf: Add an option to override Vulkan’s deviceName

  • pvr: driconf: Add an option to override Vulkan’s deviceName

  • dzn: driconf: Add an option to override Vulkan’s deviceName

  • hk: driconf: Add an option to override Vulkan’s deviceName

hwandy (1):

  • Revert “intel/decoder: make libvulkan_intel to depend on stub decoder when buildtyle=release.”

inspector-ambitious (1):

  • loader: fix loader_open_render_node_platform_devices result allocation

jglrxavpok (1):

  • RADV: Add object names inside address binding report and vm_fault

jiajia Qian (6):

  • rusticl: extract tokenize() and fix UTF-8 handling in compile options

  • rusticl: add LinkOptions struct with validation

  • rusticl: validate build/compile options before passing to backend

  • rusticl: validate input_programs binary type in clLinkProgram

  • ci/panfrost: add piglit OpenCL testing for G610

  • rusticl/device: use OpenCL spec minimum for mem_base_addr_align

jinmiliu (2):

  • mesa/st: Set protected content context flag based on pipe context attributes

  • radeonsi: enable protected context support for Android

johniyoods (1):

  • egl/dri2: require valid render fd before advertising EGL_WL_bind_wayland_display

jyotiranjan (1):

  • radv/sqtt: forward zero-submit-count vkQueueSubmit2 for SQTT capture

ljohnson (1):

  • venus/wsi: deep copy pRectangles when cloning presentation info

llyyr (2):

  • radeonsi: don’t init screen state functions twice

  • vulkan/wsi/wayland: use mtx helpers in wait_for_present2

nyanmisaka (1):

  • intel/dev: update PTL device names

sergiuferentz (1):

  • gfxstream: Prevent LINUX_GUEST_BUILD from being added to android platforms

squidbus (89):

  • kk: Use device limits for buffers and compute shared memory.

  • kk: Enable VK_AMD_shader_image_load_store_lod

  • kk: Update dynamic depth stencil state regardless of set attachments.

  • kk: Add type inference for additional built-in intrinsics.

  • kk: Fix VK_CULL_MODE_FRONT_AND_BACK with points and lines.

  • asahi,nir: Move asahi dynamic clipz pass to common.

  • kk: Add support for VK_EXT_depth_clip_control.

  • kk: Fix issues with maximal reconvergence

  • kk: Enable VK_EXT_extended_dynamic_state3

  • kk: Enable VK_EXT_buffer_device_address

  • kk: Enable VK_(EXT/KHR)_global_priority and VK_EXT_global_priority_query

  • kk: Workaround for GPU capture under Rosetta 2.

  • nir: Only attempt subgroups lower_boolean_reduce for single component.

  • kk: Expand workaround 3 to cover general use of ballot/vote ops

  • kk: Fix emitting negative infinity

  • kk: Support subgroup rotate ops

  • kk: Enable remaining subgroup operations

  • kk: Fix geometry unroll for list primitives.

  • kk: Support VK_(KHR/EXT)_index_type_uint8

  • kk: Enable VK_EXT_multi_draw

  • kk: Split per-draw data to separate binding

  • kk: Fix handling of sample mask and sample rate shading

  • kk: Support VK_EXT_post_depth_coverage

  • kk: Support robustBufferAccess2

  • kk: Support nullDescriptor

  • kk: Enable VK_(EXT/KHR)_robustness2 and VK_EXT_pipeline_robustness

  • kk: Enable VK_(EXT/KHR)_line_rasterization

  • kk: Support shaderCullDistance

  • kk: Query device for supported sample counts

  • kk: Complete VK_EXT_memory_budget

  • kk: Create image layout from vk_image

  • kk: Implement index buffer robustness for BindIndexBuffer2

  • kk: Handle accurate OpSMod and default point size requirements

  • kk: Disable A8_UNORM format

  • kk: Support device without queue

  • kk: Fix image copies for depth/stencil<->color and differing subresources

  • kk: Support new query pool and dynamic rendering flags

  • kk: Enable maintenance extensions through VK_KHR_maintenance10

  • kk: Separate linear and GPU optimized image layout properties

  • kk: Support VK_EXT_host_image_copy

  • kk: Support VK_KHR_shader_fma

  • kk: Support attachment feedback loop extensions

  • kk: Support VK_KHR_unified_image_layouts

  • kk: Fix some missed NIR debug asserts

  • kk: Allocate temporary command memory from pool

  • kk: Fix pre-compiled compute grid size

  • kk: Enable code formatting enforcement

  • poly: Refactor poly_unroll_restart for general purpose unrolling

  • poly: Fix range used for index unroll bounds checks

  • kk: Fix compute system value and algebric lowering in pre-compiles

  • kk: De-duplicate geometry unroll logic

  • kk: Support VK_KHR_shader_untyped_pointers

  • kk: Refactor multi-draws and predicates into kk_draw_data

  • kk: Enable VK_EXT_nested_command_buffer

  • kk: Support VK_EXT_conditional_rendering

  • kk: Fix precomp data buffer alignment

  • kk: Sanitize image copy through buffer extents

  • kk: Do not use image-to-image copies for 1D compressed textures

  • kk: Handle index robustness for fully bound buffers manually

  • kk: Accurately declare supported samples in image format properties

  • kk: Support VK_EXT_vertex_attribute_robustness

  • kk: Support VK_EXT_blend_operation_advanced

  • kk: Support VK_EXT_custom_resolve

  • kk: Support VK_EXT_primitive_restart_index

  • kk: Support VK_EXT_primitive_topology_list_restart

  • kk: Fence read-write images after write

  • kk: Support VK_IMAGE_CREATE_BLOCK_TEXEL_VIEW_COMPATIBLE_BIT

  • kk: Support VK_EXT_external_memory_host

  • kk: Advertise additional tessellation dynamic state

  • kk: Perform sink-and-move of instructions

  • kk: Do not force render image view to all subresources

  • kk: Fix divide by 0 in non-indexed draw unroll

  • kk: Support VK_EXT_sample_locations

  • kk: Remove unused deprecated Metal APIs

  • kk: Ensure some vertex lowerings happen on hardware stage

  • kk: Enable shaderTessellationAndGeometryPointSize

  • kk: Migrate to Metal 4 pipelines

  • kk: Work around crash with multiple concurrent MTL4Compiler

  • wsi/metal: Support HDR10 color spaces

  • kk,wsi/metal: Support VK_EXT_hdr_metadata

  • kk,wsi/metal: Support VK_(KHR/EXT)_swapchain_maintenance1

  • kk: Respect precomp-compiler options when setting up kk_clc

  • kk: Implement draw-related commands using device addresses

  • kk: Work around Metal index robustness gaps

  • kk: Pre-declare texture SSA variables

  • kk: Use safe math for nir_fp_no_reassoc

  • kk: Enable shaderRoundingModeRTEFloat16/32

  • kk: Only support 1 sample for storage images

  • kk: Remove deprecated MTL4CommandQueueErrorDeviceRemoved

utzcoz (4):

  • gfxstream: Validate guest mapped-memory ranges in flush/invalidate

  • ci/amd: enable ACO validation on radeonsi jobs

  • radeonsi: convert gather_instruction to nir_function_instructions_pass

  • virtio: magma-gpu-rs: accept a null device in virtgpu_kumquat_finish

xueyuli2 (1):

  • amd/virtio: fix bo use-after-free race condition in amdvgpu_bo_free

yserrr (4):

  • llvmpipe: fix UB and incorrect value in compute caps shift

  • v3d: fix stencil blit layer selection

  • v3d: lower more 64-bit integer operations

  • v3d: remove duplicate util_blitter_save_so_targets() call