pytorch-backlog-intelligence / data /backlog_report.md
cjc0013's picture
Publish PyTorch backlog intelligence public demo
14ed4df verified
|
Raw
History Blame Contribute Delete
20.4 kB
# pytorch/pytorch Backlog Intelligence Packet
## Coverage
- Generated: `2026-07-09T03:47:02+00:00`
- Open issues represented: `15437` of `15437` listed
- Open PRs represented: `2770` of `2770` listed
- Full collection: `True`
- Work thread active coverage: `True`
- Issue collapse coverage: `15437` open issues in `2704` groups
- Issue collapse full coverage: `True`
- Issue review-unit reduction: `82.5`%
- Maintainer action bundles: `1579`
- Patch-plan bundles: `413`
- Existing PR decision bundles: `863`
- Full issue/PR bundle coverage: `True`
- No-write audit rows: `15803`
- No-write audit boundary: `GitHub writes: none`
- Relationship edges: `4000`
- Public surface: `read-only static report`
## Local Board
- Covered open issues: `15437`
- Covered open PRs: `2770`
- Local board boundary: `no GitHub writes`
- Resolution hash: `141b35358567a5186edb6d910eab9ccd2abd54eb84d9d80d0b16c50fa4844f6d`
- This local board shows what the tool would review, close, merge, park, or turn into a local patch plan, without doing any of it on GitHub.
## No-Write Audit Packet
- All open issue proof rows: `15437`
- Stale open PR proof rows: `366`
- Proof hash: `ec5b874ff029ecfa5d4ecd204d13f40398adb089965dbb6308b8caec6b268a61`
- These rows are attachable evidence of what the tool can triage, but they do not comment, label, close, merge, or upload anything.
## Maintainer Review Threads
- Multi-issue groups: `1957`
- PR-backed groups: `570`
- Singleton issue groups retained for audit: `747`
- `1 issue` - Review thread: 1 issue + 1 PR - pr #188801: [MPS] Leak MetalShaderLibrary bundled singleton to avoid exit-time destructor crashes - linked work thread - Review issues, active PRs, referenced context, checks, and review guidance together instead of reopening each issue separately. (#188812; active PRs: PR #188801)
- `2 issues` - Review thread: 2 issues + 1 PR - pr #189122: Fix sparse-dense mul dropping data when broadcasting a size-1 sparse dim - linked work thread - Review issues, active PRs, referenced context, checks, and review guidance together instead of reopening each issue separately. (#158861, #188900; active PRs: PR #189122)
- `2 issues` - Review thread: 2 issues + 1 PR - pr #185730: Fix dynamic shapes for variadic kwargs - linked work thread - Review issues, active PRs, referenced context, checks, and review guidance together instead of reopening each issue separately. (#150022, #150371; active PRs: PR #185730)
- `1 issue` - Review thread: 1 issue + 1 PR - pr #188948: Fix triangular_solve for sparse CPU tensors on non-MKL platforms - linked work thread - Review issues, active PRs, referenced context, checks, and review guidance together instead of reopening each issue separately. (#153410; active PRs: PR #188948)
- `3 issues` - Review thread: 3 issues + 2 PRs - pr #181720: [MPS] Make pin_memory return CPU-aliased storage backed by a unified MTLBuffer - linked work thread - Review issues, active PRs, referenced context, checks, and review guidance together instead of reopening each issue separately. (#180397, #181374, #188970; active PRs: PR #181720, PR #189256)
- `1 issue` - Review thread: 1 issue + 1 PR - pr #189043: Preload full bundled cuDNN set with RTLD_GLOBAL to prevent sublibrary version mismatch - linked work thread - Review issues, active PRs, referenced context, checks, and review guidance together instead of reopening each issue separately. (#188892; active PRs: PR #189043)
- `21 issues` - Review thread: 21 issues + 14 PRs - pr #181726: [xpu][1/4]Implement scaled_mm_v2 for MXFP8/MXFP4/NVFP4 on XPU - linked work thread - Review issues, active PRs, referenced context, checks, and review guidance together instead of reopening each issue separately. (#178040, #183988, #186348, #187988, #188477, #188675, #188704, #188706, #188707, #188708, #188709, #188710; active PRs: PR #181726, PR #181727, PR #181728, PR #183511, PR #184276, PR #187315, PR #187318, PR #187989)
- `3 issues` - Review thread: 3 issues + 3 PRs - pr #184481: Guard synthetic-base input alias layouts - linked work thread - Review issues, active PRs, referenced context, checks, and review guidance together instead of reopening each issue separately. (#93617, #178680, #188133; active PRs: PR #184481, PR #184694, PR #185891)
- `3 issues` - Review thread: 3 issues + 2 PRs - pr #185946: [CUDA] Warn instead of assert in `reportProcessMemoryInfo` to prevent crash on Tegra OOM - linked work thread - Review issues, active PRs, referenced context, checks, and review guidance together instead of reopening each issue separately. (#16706, #185240, #186374; active PRs: PR #185946, PR #186375)
- `2 issues` - Review thread: 2 issues + 3 PRs - pr #185846: [decomp] Match eager's device-dependent threshold comparison dtype - linked work thread - Review issues, active PRs, referenced context, checks, and review guidance together instead of reopening each issue separately. (#185470, #185484; active PRs: PR #185846, PR #186358, PR #187609)
- `3 issues` - Review thread: 3 issues + 1 PR - pr #189045: Fix NaN gradients in nn.MultiheadAttention for fully-masked rows (#41… - linked work thread - Review issues, active PRs, referenced context, checks, and review guidance together instead of reopening each issue separately. (#40932, #41508, #55056; active PRs: PR #189045)
- `3 issues` - Review thread: 3 issues + 1 PR - pr #188998: [inductor] Fix bf16 matmul accumulation precision in Triton templates (#188492) - linked work thread - Review issues, active PRs, referenced context, checks, and review guidance together instead of reopening each issue separately. (#188492, #188841, #189287; active PRs: PR #188998)
- `3 issues` - Review thread: 3 issues + 1 PR - pr #184621: Fix Dynamo torch.Size from tensor shape inputs - linked work thread - Review issues, active PRs, referenced context, checks, and review guidance together instead of reopening each issue separately. (#182649, #182651, #182652; active PRs: PR #184621)
- `2 issues` - Review thread: 2 issues + 2 PRs - pr #189294: [CI] Realign scipy after nvmath install to fix H100 smoke NumPy ABI crash - linked work thread - Review issues, active PRs, referenced context, checks, and review guidance together instead of reopening each issue separately. (#173059, #189034; active PRs: PR #189294, PR #189295)
- `2 issues` - Review thread: 2 issues + 2 PRs - pr #187354: Match Inductor i0/i1 infinities to eager CUDA - linked work thread - Review issues, active PRs, referenced context, checks, and review guidance together instead of reopening each issue separately. (#187332, #188545; active PRs: PR #187354, PR #188556)
- `2 issues` - Review thread: 2 issues + 2 PRs - pr #185756: [clamp] Fix float16 scalar overflow check inconsistency between CPU and GPU - linked work thread - Review issues, active PRs, referenced context, checks, and review guidance together instead of reopening each issue separately. (#171356, #187429; active PRs: PR #185756, PR #187908)
- `2 issues` - Review thread: 2 issues + 1 PR - pr #185496: Fix export autograd.grad saved tensor tracing - linked work thread - Review issues, active PRs, referenced context, checks, and review guidance together instead of reopening each issue separately. (#146719, #155044; active PRs: PR #185496)
- `2 issues` - Review thread: 2 issues + 1 PR - pr #184632: Fix CUDA fake strides for mixed dtype pointwise ops - linked work thread - Review issues, active PRs, referenced context, checks, and review guidance together instead of reopening each issue separately. (#182200, #184101; active PRs: PR #184632)
- `2 issues` - Review thread: 2 issues + 1 PR - pr #184628: Fix SingletonInt static guard evaluation - linked work thread - Review issues, active PRs, referenced context, checks, and review guidance together instead of reopening each issue separately. (#182217, #183369; active PRs: PR #184628)
- `2 issues` - Review thread: 2 issues + 1 PR - pr #184392: Fix ModularIndexing printer semantics - linked work thread - Review issues, active PRs, referenced context, checks, and review guidance together instead of reopening each issue separately. (#119883, #187027; active PRs: PR #184392)
## Maintainer Work Threads
- `151` pr #188801: [MPS] Leak MetalShaderLibrary bundled singleton to avoid exit-time destructor crashes - waiting on contributor - Wait for contributor update on PR #188801; keep related issue/PR context attached. (issue #188812, PR #188801)
- `139` pr #189122: Fix sparse-dense mul dropping data when broadcasting a size-1 sparse dim - ready for maintainer decision - Review active PRs and linked issues as one maintainer work thread. (issue #158861, issue #188900, PR #189122)
- `139` issue #186535: Windows, gloo: Access violation (0xC0000005) in ProcessGroupGloo::enqueue when calling allreduce on CUDA tensors — GlooAllreduceRegistry has no kCUDA creator - needs triage - needs triage (issue #186535)
- `139` issue #188323: [Inductor][CPU] dynamic=True convolution lowering crashes with ValueError: Exponent must be non-negative - needs triage - needs triage (issue #188323)
- `127` issue #157668: NCCL error caused due to use of NVLS in torch 2.7.1-cu128 on aarch64 gb200 cluster - needs triage - needs triage (issue #157668)
- `127` issue #187912: [CPU] Concurrent `cpublas::brgemm` calls can crash in the AMX path when the underlying oneDNN ukernel is shared - needs triage - needs triage (issue #187912)
- `121` pr #185730: Fix dynamic shapes for variadic kwargs - PR blocked - Review active PRs and linked issues as one maintainer work thread. (issue #150022, issue #150371, PR #185730)
- `121` issue #116254: C++ API `at::quantized_max_pool2d`: Heap-buffer-overflow - needs triage - needs triage (issue #116254, issue #162476)
- `121` issue #162422: Runtime failure when running torch.compile() & using GCC 11.5.0 on Neoverse V1 - needs triage - needs triage (issue #162422)
- `115` issue #154297: Hangs and timeouts on dist.reduce_scatter on B200 GPU - needs maintainer decision - needs maintainer decision (issue #154297, issue #162178, issue #162745, issue #162748, issue #162820, issue #162871, issue #162897, issue #162917, issue #162940, issue #163429, issue #165170, issue #165685, issue #187158, issue #189065)
- `115` issue #55655: Performance debugging / warning mode - stale/low urgency - stale/low urgency (issue #55655, issue #57118, issue #68768, issue #72948, issue #75725)
- `115` issue #66504: BatchNorm runtimeError: one of the variables needed for gradient computation has been modified by an inplace operation - needs triage - needs triage (issue #66504, issue #68407, issue #73332)
- `115` issue #144965: RuntimeError "global alloc not supported yet" when using TorchScript optimization. - needs triage - needs triage (issue #69078, issue #144965)
- `115` issue #106164: distributed.batch_isend_irecv() crash when send/recv refers to itself - needs triage - needs triage (issue #106164)
- `115` issue #116423: PyTorch Distributed Elastic Launch Segmentation Fault with Python 3.12 - needs triage - needs triage (issue #116423)
- `115` issue #119845: Segmentation fault in dataloader worker sub-process - needs triage - needs triage (issue #119845)
- `115` issue #145610: mmap fails on 64k page aarch64 systems for AOTI model loading - needs triage - needs triage (issue #145610)
- `115` issue #153517: [CI][CUDA][Distributed] test_non_blocking_with_eager_init timeout - needs triage - needs triage (issue #153517)
- `115` issue #162731: DTensor cached op propagation results can be mutated when propagating other ops - needs triage - needs triage (issue #162731)
- `115` issue #167693: DISABLED test_side_stream_backward_overlap_cuda (__main__.TestAutogradStreamSynchronizationCUDA) - needs triage - needs triage (issue #167693)
## Maintainer Attention Queue
- `151` issue #188812: [[MPS] MetalShaderLibrary::getBundledLibrary singleton can crash at process exit](https://github.com/pytorch/pytorch/issues/188812) - needs triage - confirm whether existing evidence is enough; otherwise wait for a reproducer
- `139` issue #188900: [`sparse.mul`: broadcasting a size-1 dimension of the sparse operand silently drops data](https://github.com/pytorch/pytorch/issues/188900) - needs triage - review early; this can change release or regression risk
- `139` issue #188323: [[Inductor][CPU] dynamic=True convolution lowering crashes with ValueError: Exponent must be non-negative](https://github.com/pytorch/pytorch/issues/188323) - needs triage - review early; this can change release or regression risk
- `139` issue #186535: [Windows, gloo: Access violation (0xC0000005) in ProcessGroupGloo::enqueue when calling allreduce on CUDA tensors — GlooAllreduceRegistry has no kCUDA creator](https://github.com/pytorch/pytorch/issues/186535) - needs triage - review early; this can change release or regression risk
- `127` issue #187912: [[CPU] Concurrent `cpublas::brgemm` calls can crash in the AMX path when the underlying oneDNN ukernel is shared](https://github.com/pytorch/pytorch/issues/187912) - needs triage - confirm whether existing evidence is enough; otherwise wait for a reproducer
- `127` issue #157668: [NCCL error caused due to use of NVLS in torch 2.7.1-cu128 on aarch64 gb200 cluster](https://github.com/pytorch/pytorch/issues/157668) - needs triage - confirm whether existing evidence is enough; otherwise wait for a reproducer
- `121` issue #150022: [Dynamic Shapes with **kwargs](https://github.com/pytorch/pytorch/issues/150022) - needs triage - review early; this can change release or regression risk
- `121` issue #162422: [Runtime failure when running torch.compile() & using GCC 11.5.0 on Neoverse V1](https://github.com/pytorch/pytorch/issues/162422) - needs triage - review early; this can change release or regression risk
- `121` issue #162476: [heap-buffer-overflow in torch.quantized_max_pool2d via Python API](https://github.com/pytorch/pytorch/issues/162476) - needs triage - review early; this can change release or regression risk
- `115` issue #189194: [[ROCm] torch 2.13 wheel: "Can't detect vectorized ISA for CPU" in torch.compile smoke test on non-ROCm image (regression vs 2.12.1)](https://github.com/pytorch/pytorch/issues/189194) - needs triage - review early; this can change release or regression risk
- `115` issue #189150: [SIGSEGV / cudaErrorIllegalAddress when replaying a CUDA graph that captured multiple training iterations of two trainers over deeply unrolled recurrent modules](https://github.com/pytorch/pytorch/issues/189150) - needs triage - review early; this can change release or regression risk
- `115` issue #189282: [[CPU] operator_benchmark: backward-pass (add, + batchnorm on aarch64) ~200-500x slower since ~Nov 2025 (x86 + aarch64; possible measurement artifact)](https://github.com/pytorch/pytorch/issues/189282) - needs triage - review early; this can change release or regression risk
- `115` issue #189281: [[CPU] operator_benchmark: embedding / embeddingbag ~70-130x slower since ~Nov 2025 (x86 + aarch64; possible measurement issue)](https://github.com/pytorch/pytorch/issues/189281) - needs triage - review early; this can change release or regression risk
- `115` issue #189239: [[Inductor]Unable to apply layout optimization on convolution backward causes training performance regression.](https://github.com/pytorch/pytorch/issues/189239) - needs triage - review early; this can change release or regression risk
- `115` issue #189144: [[XPU][B580] `flex-attn-causal` performance drop with `last_level_cache_size` cache clear between runs](https://github.com/pytorch/pytorch/issues/189144) - needs triage - review early; this can change release or regression risk
- `115` issue #144965: [RuntimeError "global alloc not supported yet" when using TorchScript optimization.](https://github.com/pytorch/pytorch/issues/144965) - needs triage - review early; this can change release or regression risk
- `115` issue #189126: [Report a bug, as requested by the error message](https://github.com/pytorch/pytorch/issues/189126) - needs triage - review early; this can change release or regression risk
- `115` issue #189065: [CI for distributed tests on B200 fully red since ~2026-07-01: rank 0 "CUDA-capable device(s) busy or unavailable" -> NCCL id timeout -> 1170-min job timeout](https://github.com/pytorch/pytorch/issues/189065) - needs triage - review early; this can change release or regression risk
- `115` issue #170003: [batch_isend_irecv with nccl causes illegal memory access depending on P2P ordering](https://github.com/pytorch/pytorch/issues/170003) - needs triage - review early; this can change release or regression risk
- `115` issue #174288: [[distributed] Batched isend/irecv with NCCL backend hangs on high load](https://github.com/pytorch/pytorch/issues/174288) - needs triage - review early; this can change release or regression risk
## Issue Buckets
- stale/low urgency: `13044`
- needs triage: `1822`
- needs maintainer decision: `1332`
- has linked PR: `978`
- high-priority/blocker: `768`
- needs reproduction: `567`
- needs design decision: `353`
- good volunteer slice: `48`
## PR Buckets
- PR blocked: `1467`
- has linked issue: `1132`
- ready for maintainer decision: `558`
- draft/noise: `495`
- stale/low urgency: `366`
- waiting on contributor: `311`
- high-priority/blocker: `2`
## Top Clusters
- PR #188980 with linked issues: pr #188980, issue #188790. Review the issue and PR together instead of triaging both separately.
- PR #167224 with linked issues: pr #167224, issue #134385. Review the issue and PR together instead of triaging both separately.
- PR #189317 with linked issues: pr #189317, issue #181474. Review the issue and PR together instead of triaging both separately.
- PR #188632 with linked issues: pr #188632, issue #188544. Review the issue and PR together instead of triaging both separately.
- PR #186252 with linked issues: pr #186252, issue #184408. Review the issue and PR together instead of triaging both separately.
- PR #189313 with linked issues: pr #189313, issue #102948. Review the issue and PR together instead of triaging both separately.
- PR #189286 with linked issues: pr #189286, issue #189133. Review the issue and PR together instead of triaging both separately.
- PR #188006 with linked issues: pr #188006, issue #168868. Review the issue and PR together instead of triaging both separately.
- PR #186358 with linked issues: pr #186358, issue #185484. Review the issue and PR together instead of triaging both separately.
- PR #189129 with linked issues: pr #189129, issue #188890. Review the issue and PR together instead of triaging both separately.
- PR #186082 with linked issues: pr #186082, issue #140960. Review the issue and PR together instead of triaging both separately.
- PR #189294 with linked issues: pr #189294, issue #189034. Review the issue and PR together instead of triaging both separately.
- PR #189288 with linked issues: pr #189288, issue #189271. Review the issue and PR together instead of triaging both separately.
- PR #188931 with linked issues: pr #188931, issue #188891. Review the issue and PR together instead of triaging both separately.
- PR #188913 with linked issues: pr #188913, issue #188150. Review the issue and PR together instead of triaging both separately.
- PR #188998 with linked issues: pr #188998, issue #188492. Review the issue and PR together instead of triaging both separately.
- PR #189119 with linked issues: pr #189119, issue #189118. Review the issue and PR together instead of triaging both separately.
- PR #188996 with linked issues: pr #188996, issue #188711. Review the issue and PR together instead of triaging both separately.
- PR #189010 with linked issues: pr #189010, issue #188866. Review the issue and PR together instead of triaging both separately.
- PR #189158 with linked issues: pr #189158, issue #189157. Review the issue and PR together instead of triaging both separately.
## Limits
- This is an offline, read-only maintainer attention report.
- It does not comment, label, close, merge, or submit anything on GitHub.
- Public bundle exports are reduced rows and short excerpts; full raw API responses remain local unless explicitly approved.
- Duplicate clusters are suggestions for maintainer review, not automatic close decisions.