| issue #188812: [MPS] MetalShaderLibrary::getBundledLibrary singleton can crash at process exit | 151 | needs design decision | confirm whether existing evidence is enough; otherwise wait for a reproducer | 4 |
| issue #188900: `sparse.mul`: broadcasting a size-1 dimension of the sparse operand silently drops data | 139 | needs design decision | review early; this can change release or regression risk | 2 |
| issue #188323: [Inductor][CPU] dynamic=True convolution lowering crashes with ValueError: Exponent must be non-negative | 139 | needs design decision | review early; this can change release or regression risk | 8 |
| issue #186535: Windows, gloo: Access violation (0xC0000005) in ProcessGroupGloo::enqueue when calling allreduce on CUDA tensors — GlooAllreduceRegistry has no kCUDA creator | 139 | needs design decision | review early; this can change release or regression risk | 31 |
| issue #187912: [CPU] Concurrent `cpublas::brgemm` calls can crash in the AMX path when the underlying oneDNN ukernel is shared | 127 | needs design decision | confirm whether existing evidence is enough; otherwise wait for a reproducer | 2 |
| issue #157668: NCCL error caused due to use of NVLS in torch 2.7.1-cu128 on aarch64 gb200 cluster | 127 | needs design decision | confirm whether existing evidence is enough; otherwise wait for a reproducer | 29 |
| issue #150022: Dynamic Shapes with **kwargs | 121 | needs design decision | review early; this can change release or regression risk | 216 |
| issue #162422: Runtime failure when running torch.compile() & using GCC 11.5.0 on Neoverse V1 | 121 | needs design decision | review early; this can change release or regression risk | 220 |
| issue #162476: heap-buffer-overflow in torch.quantized_max_pool2d via Python API | 121 | needs design decision | review early; this can change release or regression risk | 300 |
| issue #189194: [ROCm] torch 2.13 wheel: "Can't detect vectorized ISA for CPU" in torch.compile smoke test on non-ROCm image (regression vs 2.12.1) | 115 | needs design decision | review early; this can change release or regression risk | 0 |
| issue #189150: SIGSEGV / cudaErrorIllegalAddress when replaying a CUDA graph that captured multiple training iterations of two trainers over deeply unrolled recurrent modules | 115 | needs design decision | review early; this can change release or regression risk | 0 |
| issue #189282: [CPU] operator_benchmark: backward-pass (add, + batchnorm on aarch64) ~200-500x slower since ~Nov 2025 (x86 + aarch64; possible measurement artifact) | 115 | needs design decision | review early; this can change release or regression risk | 0 |
| issue #189281: [CPU] operator_benchmark: embedding / embeddingbag ~70-130x slower since ~Nov 2025 (x86 + aarch64; possible measurement issue) | 115 | needs design decision | review early; this can change release or regression risk | 0 |
| issue #189239: [Inductor]Unable to apply layout optimization on convolution backward causes training performance regression. | 115 | needs design decision | review early; this can change release or regression risk | 0 |
| issue #189144: [XPU][B580] `flex-attn-causal` performance drop with `last_level_cache_size` cache clear between runs | 115 | needs design decision | review early; this can change release or regression risk | 1 |
| issue #144965: RuntimeError "global alloc not supported yet" when using TorchScript optimization. | 115 | needs design decision | review early; this can change release or regression risk | 1 |
| issue #189126: Report a bug, as requested by the error message | 115 | needs design decision | review early; this can change release or regression risk | 1 |
| issue #189065: CI for distributed tests on B200 fully red since ~2026-07-01: rank 0 "CUDA-capable device(s) busy or unavailable" -> NCCL id timeout -> 1170-min job timeout | 115 | needs design decision | review early; this can change release or regression risk | 2 |
| issue #170003: batch_isend_irecv with nccl causes illegal memory access depending on P2P ordering | 115 | needs design decision | review early; this can change release or regression risk | 23 |
| issue #174288: [distributed] Batched isend/irecv with NCCL backend hangs on high load | 115 | needs design decision | review early; this can change release or regression risk | 26 |
| issue #162731: DTensor cached op propagation results can be mutated when propagating other ops | 115 | needs design decision | review early; this can change release or regression risk | 27 |
| issue #153517: [CI][CUDA][Distributed] test_non_blocking_with_eager_init timeout | 115 | needs design decision | review early; this can change release or regression risk | 28 |
| issue #167693: DISABLED test_side_stream_backward_overlap_cuda (__main__.TestAutogradStreamSynchronizationCUDA) | 115 | needs design decision | review early; this can change release or regression risk | 28 |
| issue #116423: PyTorch Distributed Elastic Launch Segmentation Fault with Python 3.12 | 115 | needs design decision | review early; this can change release or regression risk | 31 |
| issue #106164: distributed.batch_isend_irecv() crash when send/recv refers to itself | 115 | needs design decision | review early; this can change release or regression risk | 31 |
| issue #85088: reentrant torch.utils.checkpoint does not work with NamedTuple outputs | 115 | needs design decision | review early; this can change release or regression risk | 31 |
| issue #71303: [RFC] Cross-Process Performance Analysis: Straggler Detection | 115 | needs design decision | review early; this can change release or regression risk | 32 |
| issue #75725: Warning originating in C10 backend does not get translated to Python warning if run from subprocess | 115 | needs design decision | review early; this can change release or regression risk | 32 |
| issue #76449: Enhance _verify_param_shape_across_processes | 115 | needs design decision | review early; this can change release or regression risk | 32 |
| issue #81684: Message exchange failure when perform alltoallv (cpus) | 115 | needs design decision | review early; this can change release or regression risk | 33 |
| issue #119845: Segmentation fault in dataloader worker sub-process | 115 | needs design decision | review early; this can change release or regression risk | 33 |
| issue #183459: Regression in torch.distributed._functional_collectives.all_to_all_single in 2.9.1 -> 2.11 | 115 | needs design decision | review early; this can change release or regression risk | 33 |
| issue #180088: [DTensor] RNG tracker does not advance state for CPU tensors on CUDA mesh, causing trunc_normal_ infinite loop | 115 | needs design decision | review early; this can change release or regression risk | 41 |
| issue #173921: [Inductor] Significant numerical divergence (4.7% relative error) in Conv2d on CPU with torch.compile | 115 | needs design decision | review early; this can change release or regression risk | 48 |
| issue #179502: [DTensor] sharded view incorrectly passes when redistribution is needed | 115 | needs design decision | review early; this can change release or regression risk | 93 |
| issue #145610: mmap fails on 64k page aarch64 systems for AOTI model loading | 115 | needs design decision | review early; this can change release or regression risk | 141 |
| issue #66504: BatchNorm runtimeError: one of the variables needed for gradient computation has been modified by an inplace operation | 115 | needs design decision | review early; this can change release or regression risk | 144 |
| pr #188632: Fix Dynamo opaque object staticmethod guards | 112 | ready for maintainer decision | final maintainer merge/release decision | 0 |
| pr #186429: Fix ONNX export of pad_sequence with symbolic split lengths | 112 | ready for maintainer decision | final maintainer merge/release decision | 0 |
| pr #186461: Fix SymBool equality handling in symbolic shapes | 112 | ready for maintainer decision | final maintainer merge/release decision | 0 |
| pr #184716: Fix conv_transpose2d meta output padding validation | 112 | ready for maintainer decision | final maintainer merge/release decision | 0 |
| pr #185691: Avoid UB in float to signed integer casts | 112 | ready for maintainer decision | final maintainer merge/release decision | 1 |
| pr #184939: Support SymInt steps for linspace/logspace export | 112 | ready for maintainer decision | final maintainer merge/release decision | 20 |
| pr #189307: [CI/CD] Copy newer CUPTI headers into manywheel binary-build images | 108 | waiting on CI/check fix | final maintainer merge/release decision | 0 |
| pr #186927: [MPS] Gemv kernels | 108 | waiting on CI/check fix | final maintainer merge/release decision | 0 |
| pr #189312: [CUDA][cuBLAS] Change cuBLAS default workspace size for SM 11.0 to 32 MiB | 108 | waiting on CI/check fix | final maintainer merge/release decision | 0 |
| pr #188301: [inductor] Fix inductor dropping ordering dep between effectful ops with different kernel types | 108 | waiting on CI/check fix | final maintainer merge/release decision | 0 |
| pr #184648: [Test] Add has_sufficient_memory() hook to DeviceTypeTestBase | 108 | waiting on CI/check fix | final maintainer merge/release decision | 1 |
| pr #188237: Optimize CSR SpMM CPU grain size and inner accumulation | 108 | waiting on CI/check fix | final maintainer merge/release decision | 1 |
| pr #187501: [Stable C Shim] Use error message retrieval shim if available at runtime | 108 | waiting on CI/check fix | final maintainer merge/release decision | 13 |
| pr #185882: [PrivateUse1] Return `False` instead of `None` when PU1 backend is not available | 108 | waiting on CI/check fix | final maintainer merge/release decision | 15 |
| issue #153410: [Feature request] `torch.export` .save/.load could support `safetensors` and/or `weights_only=True` | 107 | needs design decision | review linked PR state before touching issue | 171 |
| issue #188892: nvidia-cudnn-cu13==9.20.0.48 bundled by torch is incomplete → CUDNN_STATUS_SUBLIBRARY_VERSION_MISMATCH | 106 | has linked PR | confirm whether existing evidence is enough; otherwise wait for a reproducer | 1 |
| issue #188970: [MPS] SIGSEGV in MPSStream::copy with triton.cudagraphs (reduce-overhead/max-autotune) on deep fp16 graphs | 106 | has linked PR | confirm whether existing evidence is enough; otherwise wait for a reproducer | 1 |
| pr #135631: [scan] Autograd | 100 | waiting on CI/check fix | final maintainer merge/release decision | 521 |
| pr #188115: [Inductor][TEST] Align `test_main_loop_scaling` with H100 support surface | 98 | waiting on CI/check fix | identify whether block is CI, merge conflict, or review gate | 6 |
| issue #149447: DTensor slicing on sharded dimension leads to replication | 97 | needs design decision | review early; this can change release or regression risk | 200 |
| issue #164929: get_optimizer_state_dict modifies optimizer state | 97 | needs design decision | review early; this can change release or regression risk | 211 |
| issue #156563: Segmentation fault (core dumped) in `torch.profiler.profile` | 97 | needs design decision | review early; this can change release or regression risk | 231 |
| issue #122019: `F.max_pool3d`: Segfault due to incorrect `dilation` check in kernel | 97 | needs design decision | review early; this can change release or regression risk | 285 |
| issue #159238: Performance regression: torch.jit.trace() significantly slower on RTX 5090 than RTX 4060 (cu128 nightly) | 97 | needs design decision | review early; this can change release or regression risk | 333 |
| issue #121219: Torch profiler corrupted names with Python 3.11 | 97 | needs design decision | review early; this can change release or regression risk | 335 |
| issue #91165: [FSDP] FSDP with CPU offload consumes `1.65X` more GPU memory when training models with most of the params frozen | 97 | needs design decision | review early; this can change release or regression risk | 339 |
| issue #47260: DDP doesn't work with retain_graph = True | 97 | needs design decision | review early; this can change release or regression risk | 394 |
| issue #100253: profiler.export_stacks doesn't return stack trace unless experimental_config is provided | 97 | needs design decision | review early; this can change release or regression risk | 395 |
| issue #155147: torch.profiler raises Aborted (core dumped) failurer related with GIL (gilstate_tss_set) | 97 | needs design decision | review early; this can change release or regression risk | 397 |
| issue #155146: Segmentation fault when using torch.profiler | 97 | needs design decision | review early; this can change release or regression risk | 398 |
| issue #153941: TorchScript Fails Type Inference When Using .item()-Derived Scalar in Tensor Construction | 97 | needs design decision | review early; this can change release or regression risk | 414 |
| issue #153260: Export shouldn't run TS under the hood. | 97 | needs design decision | review early; this can change release or regression risk | 421 |
| issue #153145: Segmentation faults with CPU FlexAttention | 97 | needs design decision | review early; this can change release or regression risk | 425 |
| issue #152108: [inductor][cpu] AMP static shape default wrapper AOTInductor performance regression in 2025_04_20 nightly release | 97 | needs design decision | review early; this can change release or regression risk | 440 |
| issue #145088: [torchbench] torch._dynamo.exc.Unsupported: Graph break due to unsupported builtin None.morphologyEx | 97 | needs design decision | review early; this can change release or regression risk | 490 |
| issue #146788: Segmentation Fault in `torch.choose_qparams_optimized` with Invalid Parameters | 97 | needs design decision | review early; this can change release or regression risk | 514 |
| issue #124950: [distributed] First NCCL barrier does not respect timeout | 97 | needs design decision | review early; this can change release or regression risk | 525 |
| issue #71855: Regression in multi-node training speed with Transformers + PyTorch | 97 | needs design decision | review early; this can change release or regression risk | 588 |
| issue #136712: Segmentation fault (core dumped) in `torch.ao.nn.quantized.dynamic.LSTMCell/GRUCell` | 97 | needs design decision | review early; this can change release or regression risk | 650 |
| issue #136728: Aborted (core dumped) in `torch.package.package_exporter.PackageExporter`/`torch.package.PackageExporter` | 97 | needs design decision | review early; this can change release or regression risk | 650 |
| issue #135835: A segmentation fault will be raised when using the `torch._C._jit_to_backend` | 97 | needs design decision | review early; this can change release or regression risk | 664 |
| issue #81085: RuntimeError: required keyword attribute 'value' is undefined | 97 | needs design decision | review early; this can change release or regression risk | 667 |
| issue #128693: Segmentation fault (core dumped) in `torch.fused_moving_avg_obs_fake_quant` | 97 | needs design decision | review early; this can change release or regression risk | 716 |