edge_kind,source_item_type,source_number,target_item_type,target_number,confidence,evidence_source,evidence_excerpt,evidence_url,evidence_hash references,issue,141287,issue,77764,medium,issue.body,lace to list and track work on adding support to new ops for the MPS backend. Most requested ops (extracted from comments to this issue and #77764) PyTorch MPS Ops Project : Project to track all the ops for MPS backend. There are a very large number of operators in pytorch and...,https://github.com/pytorch/pytorch/issues/141287,27287ca4e6e3b6b4ece6528bc77edeef25116a884123cb7ee2d31af4ad80d401 references,issue,189287,issue,188492,medium,issue.body,"ibution needs confirmation (it is by failing-since alignment, not root-caused). Same #188251 bump has produced other confirmed regressions: #188492 (swin inductor accuracy) and #188841 (compile + float8 training numerics). Filed by pt2-oncall triage from HUD signal. cc @ptrblc...",https://github.com/pytorch/pytorch/issues/189287,9698400f2da6521759a9575e45c7f88418e35787b03ed55b7631eb8e3bc3e359 references,issue,189287,issue,188841,medium,issue.body,"failing-since alignment, not root-caused). Same #188251 bump has produced other confirmed regressions: #188492 (swin inductor accuracy) and #188841 (compile + float8 training numerics). Filed by pt2-oncall triage from HUD signal. cc @ptrblck @msaroufim @eqy @tinglvv @nWEIdia @...",https://github.com/pytorch/pytorch/issues/189287,25eed4d5223772761e053cdbd2af3b7cce89b4147c534b321bd6700d9a93f1d2 references,issue,188756,issue,84520,medium,issue.body,"09 notes flip is the only remaining MPSGraph user in Indexing.mm, but other files still lower gather/scatter-like ops to MPSGraph. Related: #84520 (MPS 32-bit limits umbrella — but here nothing asserts, it corrupts silently), #144824 (OOB indexing returns 0), #154235 (missing...",https://github.com/pytorch/pytorch/issues/188756,5ddb5dfa430472e79585e537d3eb57af3316c88b726367c07226c054df366a82 references,issue,188756,issue,144824,medium,issue.body,"slicing — the device copy is intact; only the gather is broken. All indices involved are valid; this is not the OOB-index case of #154235 / #144824, though the failure surface is the same ""OOB reads return 0"" behavior. Root cause (as far as I can tell) In 2.12.0, index_select_...",https://github.com/pytorch/pytorch/issues/188756,f092cb9517d78cce4120db3a3afe47aaf2128377257db42ad51791c110aabff3 references,issue,189270,pr,187328,medium,issue.body,"Summary kernel_information.json (#160540, extended in #183952) and the profiler-timeline provenance (#186230, #187328) map each generated kernel to a set of pre/post-grad node names, stack traces, and now extern I/O shapes. What is still missing is the phys",https://github.com/pytorch/pytorch/issues/189270,cf1d27af3d74794a50a5e5b334efdcaf82727003595ee820cb256914508b80a0 references,issue,129131,issue,97894,medium,issue.body,"/_inductor/fx_passes/pad_mm.py#L636 https://github.com/pytorch/pytorch/blob/main/torch/_dynamo/convert_frame.py#L172 (Logic refers to issue #97894) Minified repro import torch @torch.compile def some_fn(x): return torch.pow(x, 0.1) / 2 with torch.device(""cpu""): x = torch.ones(...",https://github.com/pytorch/pytorch/issues/129131,0580bf147fcef88bd580d61c801894d53b3a463de97ba138037606fe24aed172 references,issue,187455,pr,182109,medium,issue.body,"s; OpInfo / common_methods_invocations.py remains the broad semantic owner. Revisit LayerNorm backward after the active reduce/loss stacks. #182109 already has feedback asking for clearer rationale, benchmark script, and before/after results. If revived, split it into smaller...",https://github.com/pytorch/pytorch/issues/187455,9f2d6dee8fd1f5c80583f3267a977d73119c9166fd00f0623328c3eafee36435 references,issue,187455,pr,182731,medium,issue.body,tern used by the leaves. No user-visible behavior change; this is the cross-op infrastructure reviewers should see first. Scalar read leaf: #182731. Rebase on the UMA base helper. Scope: _local_scalar_dense / .item() / scalar truthiness direct CPU read. MPS->CPU copy leaf: #18...,https://github.com/pytorch/pytorch/issues/187455,931ba6c553643a760ccb443d494823a0c03e5d2d705a5ab71b93f5a266fa0aa4 references,issue,187455,pr,182736,medium,issue.body,"leaf: #182731. Rebase on the UMA base helper. Scope: _local_scalar_dense / .item() / scalar truthiness direct CPU read. MPS->CPU copy leaf: #182736. Rebase on the UMA base helper. Scope: contiguous same-dtype copy_from_mps_ / .to(""cpu"") / .numpy() direct read. CPU->MPS copy le...",https://github.com/pytorch/pytorch/issues/187455,dd905b22f5c769d19280e2f72feb508eebe5b0c978eae6dcd77a6820a259b9d2 references,issue,187455,pr,182749,medium,issue.body,"182736. Rebase on the UMA base helper. Scope: contiguous same-dtype copy_from_mps_ / .to(""cpu"") / .numpy() direct read. CPU->MPS copy leaf: #182749. Rebase on the UMA base helper. Scope: contiguous same-dtype CPU->MPS direct write, with non-blocking/non-contiguous/dtype-cast f...",https://github.com/pytorch/pytorch/issues/187455,39ad423213d6c30d7edcd3619f8789b5914d4a0303b1c24a70bc34c5c6107928 references,issue,187455,pr,182791,medium,issue.body,"elper. Scope: contiguous same-dtype CPU->MPS direct write, with non-blocking/non-contiguous/dtype-cast fallbacks preserved. Fill/zero leaf: #182791. Rebase after the write-side copy leaf. Scope: dense zero_/fill_ via memset/std::fill_n, with negative-zero and unsupported dtype...",https://github.com/pytorch/pytorch/issues/187455,8ce2999c8c1c71211ab727a25ada447ebba73c20e3518765eae8f316d6715822 references,issue,187455,pr,183688,medium,issue.body,"182791 draft/pending rework Rebase onto the UMA base/write-side direction; do not review as a standalone final PR. MPS capture/replay stack #183688 is much larger than a reviewable PyTorch PR and should be treated as a reference branch to split, not a final PR. Order: Capture...",https://github.com/pytorch/pytorch/issues/187455,9df4bd150ddf31bc316ce476fd7f829ad13277ecb7bace954fc85b13c855a01c references,issue,187455,pr,187456,medium,issue.body,"lit at the last-dim vs non-last-dim boundary and stacked as cumulative-leaf forks on main (external contributor, no ghstack): softmax core (#187456, 52a3223): native last-dim _softmax/_softmax_backward (single-row + 8-wide-half + looped online-softmax); non-last-dim and huge-r...",https://github.com/pytorch/pytorch/issues/187455,345ffb2dce60cb2b34d124fed94231118ec9415c1281c040878c8ed45ac7f037 references,issue,187455,pr,187457,medium,issue.body,"-dim-only gate. Residual dim=0-backward tail is documented (MPSGraph's fused reduce still wins there; escape hatch covers it). log_softmax (#187457, 3d6f962): native log_softmax stacked on the softmax kernels. cross_entropy (#187458, ee60969): fused 2D cross-entropy Metal kern...",https://github.com/pytorch/pytorch/issues/187455,4693297ad47a1ed0ffaf041a3a3b91720068b11cbbadf04b36ed80bb5e106daa references,issue,187455,pr,187458,medium,issue.body,"antiations each), not chained on each other. Independent loss-style work: #187559 (nll_loss, review comments addressed at fc26ad2e55e3) and #187458 (cross_entropy) do not fit the same elementwise fused-loss template. Softmax / log_softmax / cross_entropy: refreshed and opened...",https://github.com/pytorch/pytorch/issues/187455,2d0d3aa86c56419e390e88ebb84b2aef1fe2f403d789ea9a8a900d66e98ac76a competes with,issue,187455,pr,187553,medium,issue.body,"into the existing TestMetalLibrary.test_reduction_utils owner surface instead of adding a standalone regression. Land the loss-stack base: #187553 (mse_loss). Scope: shared fused-loss infrastructure (FusedLossParams, fused_loss_pass1, fused_loss_bwd, fused_loss_reduce, TensorI...",https://github.com/pytorch/pytorch/issues/187455,5ead698bdfb6499f55eff3d9865a7aca19a56d5d71a0e0c3761c5d982d27dc30 references,issue,187455,pr,187556,medium,issue.body,"ases); OpInfo consistency (including the f16 grad test) remains the broad semantic owner. Stack: rebuild #187557 (binary_cross_entropy) and #187556 (smooth_l1_loss/huber_loss) as sibling leaves on the final #187553 (small op functor + instantiations each), not chained on each...",https://github.com/pytorch/pytorch/issues/187455,919f58fd26fa22c48388ef0ce0cc76ab80ab406ed90785dc78ba0b4a212ea983 references,issue,187455,pr,187557,medium,issue.body,"d, mixed dtypes, out-variant edge cases); OpInfo consistency (including the f16 grad test) remains the broad semantic owner. Stack: rebuild #187557 (binary_cross_entropy) and #187556 (smooth_l1_loss/huber_loss) as sibling leaves on the final #187553 (small op functor + instant...",https://github.com/pytorch/pytorch/issues/187455,4b4c4b9ff1073e29e28405a99aff70564283a96f27ef4b0c1d4d4b4699486118 references,issue,187455,pr,187559,medium,issue.body,"s) as sibling leaves on the final #187553 (small op functor + instantiations each), not chained on each other. Independent loss-style work: #187559 (nll_loss, review comments addressed at fc26ad2e55e3) and #187458 (cross_entropy) do not fit the same elementwise fused-loss temp...",https://github.com/pytorch/pytorch/issues/187455,c2e9079116af0b2664b8d9cf1704ea257b819b8f57dd55fd0c116be7e28a284d references,issue,187455,pr,187787,medium,issue.body,"guard; prod perf: outer/inner shape specializations, benchmark, perf table. Review var/std after the reduce base/prod direction is settled: #187787. Current state: rebuilt and pushed at cd01523561027; ready as the Welford follow-up, but should be read after #188156/#188333 bec...",https://github.com/pytorch/pytorch/issues/187455,952d13051c9d44a446a2be73e86248a8d47422d25eeff451dfe01383173f6430 references,issue,187455,pr,187976,medium,issue.body,"anup: complex enablement, integer rejection, duplicate-dim behavior, and large-output errors if those remain separate review concerns. Keep #187976 independent. #187976 fixes an int64 amax/amin partial-simdgroup bug in the existing value-reduction helper path. It is a sibling...",https://github.com/pytorch/pytorch/issues/187455,a6a44834bd4df619c0ecf08fbcf7fc9d812c5e0bf2cd91c7401bf348805177df references,issue,187455,pr,188156,medium,issue.body,"A standalone test_mps.py case needs a clear reason why no existing surface owns the route. Order of operations Land the reduce-family base: #188156. Scope: collapse generic sum/mean/nansum/count_nonzero onto the shared value_reduction path already used by min/max...",https://github.com/pytorch/pytorch/issues/187455,6d4f2eed50ae483231d5f1291cd1d08b298ba1ff098a4eed5c37fae987ed7a54 supersedes,issue,187455,pr,188333,medium,issue.body,"route for non-contiguous multi-dim sum-family reductions and load policies, not just a private helper. Land prod on top of the reduce base: #188333. Scope: migrate prod to value_reduction and remove its shape-dependent MPSGraph route. This supersedes the closed standal...",https://github.com/pytorch/pytorch/issues/187455,64907f0296db66e3f73e9c6406bcd1aa15b12816066489d2f98f9630b55a7190 references,issue,187455,pr,189146,medium,issue.body,"ned MPSGraph fallback, so it is correct standalone and strictly reduces the softmax graph-cache for the common last-dim case. softmax perf (#189146, 8fa0937): adds the huge-row two-pass and non-last-dim (coalesced/tiled/blocked) native kernels on top of core and removes the la...",https://github.com/pytorch/pytorch/issues/187455,ff04e385d39996a9012222e032a6741e7118b0823c5d3a9b59ce73aeaa1c8b26 references,issue,187175,issue,186241,medium,issue.body,to it unconditionally. There are several recently filed issues that highlight the TMA path for pointwise/reduction is not yet fully mature: #186241 #185513 #185222 #184714 #184563 Phase 2 — Once we are confident the TMA path is correct and mature: Enabling autotuning alone (wi...,https://github.com/pytorch/pytorch/issues/187175,447ca33d4dbd72febb21f81154fa1ee7db554dfabb534b423eff70ad1d01e8d1 references,issue,187810,issue,187811,medium,issue.comments[1].body,"Hi @chunhuanMeng , please help handle this issue and #187811",https://github.com/pytorch/pytorch/issues/187810,84c3b4bdab0afe48a014091b2f1016135305c1871a6dcf86c9eb8939aff8f862 references,issue,189138,issue,189135,medium,issue.body,"is replaced by _exposes_streams(device), which checks whether iface.Stream is not DeviceInterface.Stream. Please refer to the dedicated RFC #189135 for detailed design elaboration. Dynamo Registry Splicing Dynamo answers two questions for every value/call during symbolic execu...",https://github.com/pytorch/pytorch/issues/189138,e8b0f2becfc15424643138f062faee984ee45d6b1e5b8479bfcd04876886be85 references,issue,189138,issue,189136,medium,issue.body,"e-specialized tensor types — is handled by one optional slot dynamo_tensor_classes (default empty, safe). Please refer to the dedicated RFC #189136 for detailed design elaboration. Inductor Codegen Registration Interfaces Inductor's codegen path has four things that lack regis...",https://github.com/pytorch/pytorch/issues/189138,175dc64f062f304b809a5c88277b99d8dcf7740ecb1dafd45a75009868f8126c references,issue,189138,issue,189137,medium,issue.body,"E_TO_ATEN fallback. No-c-shim backends will still fail AOTI at link time. This needs a separate proposal. Please refer to the dedicated RFC #189137 for detailed design elaboration. For a more comprehensive and detailed understanding of each component's design, we highly encour...",https://github.com/pytorch/pytorch/issues/189138,75a60e2adbd3fd72c533e18e7d53249ad6bf4b74cb2ed1705175cb815516f709 references,issue,144965,issue,69078,medium,issue.body,"der some specific conditions, raise a RuntimeError with message: ""Global alloc not supported yet."". I think this is linked to an old issue: #69078, however, I managed to consistently reproduce this error. The code that reproduces the bug is quite long and needs some explanatio...",https://github.com/pytorch/pytorch/issues/144965,af0a8b92d21995f8ec4b0b5a5d44613aaaaab920e32715959df72818a2c6066d references,issue,188970,issue,180397,medium,issue.body,"not collect Expected behavior Compile succeeds, or the unsupported/unsafe option is ignored with a warning — not a native SIGSEGV. Related #180397 proposes a capture/replay API for MPS and notes MPS currently has no CUDA-Graphs equivalent — possibly useful context, since trito...",https://github.com/pytorch/pytorch/issues/188970,8dc4c11946ae99660774b1637ca1d88153f73a9a68462b4a70a8207182287fd5 references,issue,187093,pr,189127,medium,issue.comments[1].body,Opened a draft PR with the scoped post-grad rewrite: #189127 The patch extends the existing addmm bias-unfuse pattern to baddbmm: it only applies on GPU when the input is bias-like/broadcast and all b,https://github.com/pytorch/pytorch/issues/187093,56c41fc8de4001e0f38282047b6bfb8a47f1e098c860eb723fb2cd2a8c320d4c references,issue,189121,pr,189124,medium,issue.comments[1].body,"Opened a draft PR with the focused cache-gate patch: #189124 The patch keeps the ordinary single-config fast path unchanged, but lets DSR-eligible single-config reductions create/load an autotune cach",https://github.com/pytorch/pytorch/issues/189121,d2b79c1ccd9aea3c11d6213c11b639a5cc50550df8650ad729bda49e2271fd54 references,issue,188602,pr,186245,medium,issue.body,used to optimize the binary. The current plan is that this needs the following high level steps: Integrate LLVM-BOLT into the build system #186245 Add a profile staleness metric script which guides profile updates Enable profile collection in PyTorch CI cc @malfet @ptrblck @ms...,https://github.com/pytorch/pytorch/issues/188602,eafa00736571e47a2117cc23a189b6b089ae0d1f7ef2932593feb7b174ca8374 references,issue,189078,pr,186245,medium,issue.body,add this so we can keep track of how various changes impact build times. The main motivation for this comes from the desire to ensure that #186245 does not introduce major build time overhead but metrics should be helpful in general. Alternatives No response Additional context...,https://github.com/pytorch/pytorch/issues/189078,15a0ad69b44471a4d861bf8220eee35293248bd9f1f3227256c083a28432d825 references,issue,189098,pr,189097,medium,issue.body,o Kineto could utilize the export_using_protocol method in order to access them in Pytorch. A draft PR of the changes has been opened here: #189097 Alternatives No response Additional context No response cc @robieta @chaekit @guotuofeng @guyang3532 @dzhulgakov @davidberard98 @...,https://github.com/pytorch/pytorch/issues/189098,02871cf4cce200404cad179b35736868e013efe2530612dc49c5e2b89b1af758 references,issue,189077,issue,188557,medium,issue.body,"d-environment linux-jammy-cuda13.0-py3.12-gcc11-sm100, single unsharded (1, 1) job on one Blackwell GPU. This is the NVIDIA/b200 sibling of #188557 (which tracks the same failure mode on ROCm gfx950). Same root cause and fix levers; different infra/owners, so filing separately...",https://github.com/pytorch/pytorch/issues/189077,a245fc0a78e2348b1ce75282062a5e55c27c5db63737321c27f17e4ea8a3485b references,issue,189075,issue,140556,medium,issue.body,"never completes. Genuine per-test failures in the same job (e.g. the test_backward_prod_cuda_float32 mem-leak-check, tracked separately in #140556) are a distinct problem from this timeout. Fix levers (CI-side) Increase the shard count (currently 8) so each shard's slice fits...",https://github.com/pytorch/pytorch/issues/189075,682fa9067cfe772f7378d44a92d3c58c9d10c338662ca6b4b9b83579609f2bde references,issue,189065,issue,154297,medium,issue.body,"0 runner pool: investigate GPU state on linux.dgx.b200.8 (leaked processes, compute mode, or reimage). Nothing to fix in test code. Related #154297 - Hangs and timeouts on dist.reduce_scatter on B200 GPU (user-reported; possibly shared underlying B200 cause). #162178 - [CUDA]...",https://github.com/pytorch/pytorch/issues/189065,e7c80387a19e22fcfe21acebc2f3514911d5bab4d80513a29d95491c258a1062 references,issue,189065,issue,162178,medium,issue.body,test code. Related #154297 - Hangs and timeouts on dist.reduce_scatter on B200 GPU (user-reported; possibly shared underlying B200 cause). #162178 - [CUDA] Umbrella Issues/Failures on B200 Runner (stale umbrella). cc @ezyang @gchanan @kadeng @msaroufim @awgu @wanchaol @fegin @...,https://github.com/pytorch/pytorch/issues/189065,b36586005896bfd37626f9eb243032b2e5752e9be2914e7be637b68a084a6805 references,issue,188733,issue,45111,medium,issue.body,"al MLA / attention changes that alter KV-cache layout/strides and Sm90 backend selection for the MLA path -- most notably: vllm-project/vllm#45111 - ""Re-enable cross-layer KV cache layout for MLA via stride-aware kernels"" (changes MLA KV-cache strides/layout) -- primary suspec...",https://github.com/pytorch/pytorch/issues/188733,1d558cb9461e859db727514da4a760bfc7ca0df7fb0de4238ebc15a6b331d43a references,issue,189034,issue,173059,medium,issue.body,"iPy to a NumPy-2-compatible build (SciPy >=1.13 supports NumPy 2). Whoever owns the H100 image / bumped numpy 2.2.6 should confirm. Related #173059 -- same class (NumPy binary/ABI incompatibility breaking test imports), different instance (pandas + test_dataloader/test_datapip...",https://github.com/pytorch/pytorch/issues/189034,4ff8aa912ca76a90e33f07b086d70223c3964f06cf9565c1b1a97a7b8e85df01 references,issue,188900,issue,158861,medium,issue.comments[0].body,"s-is, so only the slice at index 0 along the broadcast dimension gets values and the rest stay zero. This looks like the same root cause as #158861, where `@nikitaved` noted that explicit broadcasting via `sparse_broadcast_to` followed by coalesce is missing. Proposed fix: bef...",https://github.com/pytorch/pytorch/issues/188900,ea7c404fac461685b6138c3782e0a7aa2a90ce8b26aa1c804b94bd8e9cf24563 references,issue,188792,pr,188791,medium,issue.body,"each worker do a serial 1-element torch.sqrt first; this drops the rate to 0 (5,200 first-calls), which is the basis of the fix proposed in #188791. mkl_vml_race_repro.py """"""Minimal repro: MKL VML first-call dispatch race in torch.sqrt. Spawns waves of worker processes on an o...",https://github.com/pytorch/pytorch/issues/188792,7a83aacf57b04dde3d1b2448cf0f4270ff7c363003b1b84fc081ebdb5c886f78 references,issue,188564,issue,152595,medium,issue.body,"Here the build is fine, the GPU is found, tests run and pass, and the job is killed purely by the 270-min wall-clock. (Other gfx1100 issues #152595 / #165141 / #169857 / #177834 are unrelated.) Suggested fix Unblock now (CI-side): increase shard count for this job (it's num_sh...",https://github.com/pytorch/pytorch/issues/188564,7f98c780b875fbc5917ea441da987ef5c78b69b8ed2330a6f9d3b6d21d088cde references,issue,188564,issue,165040,medium,issue.body,"s confirmed for the current runs but not individually proven as the culprit for all 388 red runs. This is NOT covered by the existing issue #165040 ""DISABLED rocm / linux-jammy-rocm-py3_10-gfx1100 / test (default)"" (OPEN; module: rocm, module: ci) is adjacent but distinct: it'...",https://github.com/pytorch/pytorch/issues/188564,1727021f32469529a36c72b5fc4f013ba7de4ed2f4b810df8c96ebecb1fc8be8 references,issue,188564,issue,165141,medium,issue.body,"uild is fine, the GPU is found, tests run and pass, and the job is killed purely by the 270-min wall-clock. (Other gfx1100 issues #152595 / #165141 / #169857 / #177834 are unrelated.) Suggested fix Unblock now (CI-side): increase shard count for this job (it's num_shards=2 tod...",https://github.com/pytorch/pytorch/issues/188564,60d9d4eb9c8753752450ee1167bd539202a190c38396b057e8ef0c3262a396e7 references,issue,188564,issue,169857,medium,issue.body,"ne, the GPU is found, tests run and pass, and the job is killed purely by the 270-min wall-clock. (Other gfx1100 issues #152595 / #165141 / #169857 / #177834 are unrelated.) Suggested fix Unblock now (CI-side): increase shard count for this job (it's num_shards=2 today) so eac...",https://github.com/pytorch/pytorch/issues/188564,1f4d60a6b9f14e9096dc8a717bf935fce550e4de05919055bdbbf7467eb50770 references,issue,188564,issue,177834,medium,issue.body,"U is found, tests run and pass, and the job is killed purely by the 270-min wall-clock. (Other gfx1100 issues #152595 / #165141 / #169857 / #177834 are unrelated.) Suggested fix Unblock now (CI-side): increase shard count for this job (it's num_shards=2 today) so each Test ste...",https://github.com/pytorch/pytorch/issues/188564,75777d9e62901abcd78bcf4fcf56ad11f56310d0079b192934bc543a14640ce5 references,issue,188545,issue,187332,medium,issue.body,"e returns [nan, nan, nan], but Inductor returns [nan, 0., nan]. This changes the isnan result for the +inf element. This appears related to #187332 , but it affects torch.special.bessel_y0. Repro import torch def test_bessel_y0_mismatch(): x = torch.tensor([float('nan'), float...",https://github.com/pytorch/pytorch/issues/188545,6ccedbc5c7e5104c094940d20c35e1caf868f0491c04b80f0ac66cfa51bc06db references,issue,188008,pr,187955,medium,issue.body,"ing unwinding → std::terminate. This is a pre-existing latent bug (the per-segment unmap loop had the same defect); it is not introduced by #187955, but that PR's unmapHandles rework touches the same code, so filing as follow-up. Mechanism map() / fromShared() set handles_[beg...",https://github.com/pytorch/pytorch/issues/188008,a0b4f7b49dff0e32705ba8c38ef4f342544be554b1648ec1a1fed54210ffcf10 references,issue,188008,pr,187955,medium,issue.comments[0].body,Being fixed in #187955: mapAndSetAccess is now strongly exception-safe — on a partial map/setAccess failure it unmaps the mapped prefix and releases the range's h,https://github.com/pytorch/pytorch/issues/188008,aaa82942a131c603561fa9378f65c45b3a3d5cc7cd2000202884d9ffee5f006f references,issue,140718,issue,77764,medium,issue.body,"y implemented for the MPS device. If you want this op to be added in priority during the prototype phase of this feature, please comment on #77764. As a temporary fix, you can set the environment variable PYTORCH_ENABLE_MPS_FALLBACK=1 to use the CPU as a fallback for this op....",https://github.com/pytorch/pytorch/issues/140718,913bb5cf94a50de8ea0c409f50d5e078de09feb8a75dfcd334449ef7e68ce925 references,issue,188953,issue,179891,medium,issue.body,"xpu: GPU detected, data transfer works, but any non-trivial computation (matmul) fails with could not create an engine. Related to existing #179891 (get_device_properties segfault on BMG). Intel compute-runtime issue cross-reference: intel/compute-runtime#947 Environment Compo...",https://github.com/pytorch/pytorch/issues/188953,edf3792aa9ff53d09b457d938437da2f12830a88948c595bbf819c7e2ba2b893 references,issue,188718,pr,188719,medium,issue.comments[1].body,Thanks @Sahil170595. The PR is #188719,https://github.com/pytorch/pytorch/issues/188718,d8de2cb15f5aab0c45ed4e1a918095e627d417e6a759e0bb722011f90c8d4967 references,issue,175891,issue,62530,medium,issue.body,"itional nodes, switch can do the same in CUDA >= 12.8. Alternatives No response Additional context This has previously been requested here: #62530 cc @mcarilli @ezyang @eellison @penguinwu @BoyuanFeng @chauhang @ydwu4 @bdhirsh @bobrenjc93 @aorenste",https://github.com/pytorch/pytorch/issues/175891,9a00af6f1c18741467c70dc5876e37327ef8e01a8789fa2f2fe567b1877b247f references,issue,182277,issue,177827,medium,issue.body,"aise a clean Python exception such as OverflowError, ValueError, or a normal CUDA OOM RuntimeError, not SystemError. This may be related to #177827, but it does not seem to be the same issue. #177827 is about negative allocation sizes, while this issue is triggered by a positi...",https://github.com/pytorch/pytorch/issues/182277,0c280b4f4ac95b5cd7b9eb9d77e20e1830114ba4cb6350cfd1897c45adbc30d9 references,issue,188812,pr,188801,medium,issue.body,MetalShaderLibrary& MetalShaderLibrary::getBundledLibrary() { static BundledShaderLibrary* l = new BundledShaderLibrary(); return *l; } PR: #188801 (will be scoped down to just this leak). cc @malfet @aditvenk @kulinseth @DenisVieriu97 @jhavukainen @Isalia20 Versions PyTorch v...,https://github.com/pytorch/pytorch/issues/188812,c2134ecc4dba4b1da73c0d76b5a01693ba1c84a1c535b16b58927411e651d4df references,issue,188319,pr,188951,medium,issue.comments[1].body,"Implemented the core pipeline changes in #188951, covers matrix generation, Makefile, Dockerfile, and workflow env vars. Would appreciate any feedback on the approach.",https://github.com/pytorch/pytorch/issues/188319,2efa83a6f79f702c8ae605facf76fdd636076faba758ac5e93db3089c80368d1 references,issue,177116,issue,117826,medium,issue.body,"shaders, fixed for torch.mm/torch.bmm via tiling in PR #117549 but not all code paths), #122045 (F.linear wrong results for large inputs), #117826 (incorrect einsum gradient on MPS). Minimal reproduction The model is Embedding → 2 residual Linear blocks → output Linear, traine...",https://github.com/pytorch/pytorch/issues/177116,8ed85e7ca715f1fab7ccca1a190b0fd1593e47912a1109682dbc672e09b66fad references,issue,184575,pr,188066,medium,issue.comments[1].body,~1e12 in backward (not NaN). This is a silent gradient corruption — users see finite gradients when they should see NaN. Submitted test PR: #188066 Detected via NeuralDBG — causal diagnostic engine for PyTorch training.,https://github.com/pytorch/pytorch/issues/184575,9d5bc7624ee970ffb047679d26c36ed9cf9a8dab9b283dc84652743252a3ca4a references,issue,188891,issue,187988,medium,issue.body,"1-D weight correctly, which is why eager rank >= 3 works. Context Found while adding 1-D-weight OpInfo samples to sample_inputs_linear for #187988 / PR #187989 (an unrelated MPS backward crash with 1-D weights): a sample with a scalar bias turns 14 functorch OpInfo tests red o...",https://github.com/pytorch/pytorch/issues/188891,5c0571484e233f01f4babb49eac24cc12f29f0aa634da0332f085e0c7df35515 references,issue,188891,pr,187989,medium,issue.body,"orrectly, which is why eager rank >= 3 works. Context Found while adding 1-D-weight OpInfo samples to sample_inputs_linear for #187988 / PR #187989 (an unrelated MPS backward crash with 1-D weights): a sample with a scalar bias turns 14 functorch OpInfo tests red on all backen...",https://github.com/pytorch/pytorch/issues/188891,c614743a0c8262a4303764eb0e730a4fdd43b87def99992865382c70940be5f6 references,issue,188670,pr,169694,medium,issue.body,nce / ROCm follow-up #182028 ROCm MI200 — flex vs SDPA assert_close tolerances #186889 XPU — relaxed assertEqual vs eager (atomic ordering) #169694 Transformer tests — golden precision / reference choice #183691 Test infra — tolerance overrides (helps symptoms; root cause is s...,https://github.com/pytorch/pytorch/issues/188670,198e53966f626e02bbdf17b398887cd138b4318c2c9a462aa2ae6ccc1ae15874 references,issue,188670,pr,183691,medium,issue.body,ose tolerances #186889 XPU — relaxed assertEqual vs eager (atomic ordering) #169694 Transformer tests — golden precision / reference choice #183691 Test infra — tolerance overrides (helps symptoms; root cause is still design in individual tests) 4. Conclusion & proposed next s...,https://github.com/pytorch/pytorch/issues/188670,e77c5c394c4db350dad1287de44cbd34ca3016b7a29375880c9e86389e7b7a03 references,issue,188670,pr,184531,medium,issue.body,"oesn't solve the root of the issue. Tight same() / eager-fp16 equality is stricter than the bug needs and varies by hardware (e.g. #179958, #184531) — false failures even when the grid fix is sound. 3. Other examples (same pattern family) Link Note #183630 test_bmm_large_batch...",https://github.com/pytorch/pytorch/issues/188670,bd51bf5985bf862772fa908fe0af098cbe2480688a50b0a83f66cafb860b7ba2 references,issue,188670,pr,187050,medium,issue.comments[0].body,"Here is a PR that implements tests for bmm issue in a way that I think is better than current implementation: #187050 It shouldn't be used to close the issue, just serve as an example of the direction that I propose to move towards to.",https://github.com/pytorch/pytorch/issues/188670,a6427297580fb8d1214da154523bc28d92d379f1130f2178df72aaef3e3c91e8 references,issue,188727,pr,188500,medium,issue.comments[1].body,ike GraphPP and Eager Torchtitan PP. We are changing PP to no longer handle microbatch splitting and let the user supply microbatches here: #188500 Torchtitan will now produce microbatches right at source i.e. dataloader: pytorch/torchtitan#3856 So no changes are required from...,https://github.com/pytorch/pytorch/issues/188727,fa1d6ca45bc0918ba4f584c1be38f60155ad873dd4d1789d3473e40dc5f8e10a references,issue,188772,pr,188698,medium,issue.body,tem.py (which is brittle) and implement their own workarounds for asynchronous saving over fsspec. Additional context This is drafted in PR #188698 and discussed offline with @LucasLLC and @meetv18. Opening this issue to take it further with the process cc @awgu @wanchaol @feg...,https://github.com/pytorch/pytorch/issues/188772,f672141928f220635f80b6689a8a4bf16d8e328f2cb1f56d4c4965a7c26729e1 references,issue,188711,issue,188675,medium,issue.body,"fast_accum_True_scaling_block_sizes2_cuda One of ~14 failing test_main_loop_scaling parametrizations (w = BlockWise1x128 recipes). See also #188675, #188704. cc @ptrblck @msaroufim @eqy @tinglvv @nWEIdia",https://github.com/pytorch/pytorch/issues/188711,380e8b3b0e3d23c78bc388f0c39b430e7091ba46c1b318dee418bc47871250aa references,issue,188711,issue,188704,medium,issue.body,"m_True_scaling_block_sizes2_cuda One of ~14 failing test_main_loop_scaling parametrizations (w = BlockWise1x128 recipes). See also #188675, #188704. cc @ptrblck @msaroufim @eqy @tinglvv @nWEIdia",https://github.com/pytorch/pytorch/issues/188711,5d870351811498879b4882276cdb3b8ad227981535b4d9babb5dbe37ee5349d0 references,issue,188707,issue,188675,medium,issue.body,"ast_accum_False_scaling_block_sizes1_cuda One of ~14 failing test_main_loop_scaling parametrizations (w = BlockWise1x128 recipes). See also #188675, #188704. cc @ptrblck @msaroufim @eqy @tinglvv @nWEIdia",https://github.com/pytorch/pytorch/issues/188707,086d0df5d173b95613cf8cc09a713ec31938bbe00b72919802796fb04158cce8 references,issue,188707,issue,188704,medium,issue.body,"_False_scaling_block_sizes1_cuda One of ~14 failing test_main_loop_scaling parametrizations (w = BlockWise1x128 recipes). See also #188675, #188704. cc @ptrblck @msaroufim @eqy @tinglvv @nWEIdia",https://github.com/pytorch/pytorch/issues/188707,8aad4c9a3346ff569d46f2e73ceaea770185441378d8365818c617c77dc759ca references,issue,188706,issue,188675,medium,issue.body,"ast_accum_False_scaling_block_sizes0_cuda One of ~14 failing test_main_loop_scaling parametrizations (w = BlockWise1x128 recipes). See also #188675, #188704. cc @ptrblck @msaroufim @eqy @tinglvv @nWEIdia",https://github.com/pytorch/pytorch/issues/188706,800422d2dbb726ce7dc6d5d2b41483d112e6bfa3de8f3cbb186fd73a351fbc2f references,issue,188706,issue,188704,medium,issue.body,"_False_scaling_block_sizes0_cuda One of ~14 failing test_main_loop_scaling parametrizations (w = BlockWise1x128 recipes). See also #188675, #188704. cc @ptrblck @msaroufim @eqy @tinglvv @nWEIdia",https://github.com/pytorch/pytorch/issues/188706,e945e5c0dec8f603f8a1226536e8ebb9ceec1f90de8152f9072435d0f78e09f7 references,issue,102207,issue,108798,medium,issue.body,"_.from needs to be renamed"" 'test_advancedindex_mixed_cpu_devices' - ""FIXME"" 'test_advancedindex_mixed_devices_error' - ""FIXME"" #108181 and #108798 'test_typed_storage_deprecation_warning' - ""FIXME"" 'test_pin_memory' - ""pin_memory isn't yet supported in TorchInductor"" 'test_me...",https://github.com/pytorch/pytorch/issues/102207,9eac0f3787e9f7be4bae833fe135aa372e815033066e68e5669d452d2ff375ce references,issue,188721,issue,188477,medium,issue.body,: module 'cutlass.cute.arch' has no attribute 'ProxyKind' (unpinned cutlass_api vs pinned nvidia-cutlass-dsl==4.5.2). Root cause tracked in #188477. Unstable until the cutlass_api / cutlass-dsl version skew is resolved. cc @malfet @pytorch/pytorch-dev-infra,https://github.com/pytorch/pytorch/issues/188721,10df704ed063fc92983779ef79ab6ae6c3f17cb633156e78fca63c99cc6dbec4 references,issue,188714,issue,188675,medium,issue.body,"fast_accum_True_scaling_block_sizes1_cuda One of ~14 failing test_main_loop_scaling parametrizations (w = BlockWise1x128 recipes). See also #188675, #188704. cc @ptrblck @msaroufim @eqy @tinglvv @nWEIdia @mruberry @chauhang @penguinwu @voznesenskym @EikanWang @jgong5 @Guobing-...",https://github.com/pytorch/pytorch/issues/188714,c0ffd4ffb8e61edc1a7c08df3df4169354695b25febc034bde3c4b3e5565bf21 references,issue,188714,issue,188704,medium,issue.body,"m_True_scaling_block_sizes1_cuda One of ~14 failing test_main_loop_scaling parametrizations (w = BlockWise1x128 recipes). See also #188675, #188704. cc @ptrblck @msaroufim @eqy @tinglvv @nWEIdia @mruberry @chauhang @penguinwu @voznesenskym @EikanWang @jgong5 @Guobing-Chen @Xia...",https://github.com/pytorch/pytorch/issues/188714,f6642f1800f107c17de69c746216551c020e432c2aafb41e60b78c092f123237 references,issue,188715,issue,188675,medium,issue.body,"fast_accum_True_scaling_block_sizes2_cuda One of ~14 failing test_main_loop_scaling parametrizations (w = BlockWise1x128 recipes). See also #188675, #188704. cc @mruberry @chauhang @penguinwu @voznesenskym @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @we...",https://github.com/pytorch/pytorch/issues/188715,5a37ccad30a9eecfe83cb39ab9baab004caf4c3a94d0e77bb4de32f34f0a85b7 references,issue,188715,issue,188704,medium,issue.body,"m_True_scaling_block_sizes2_cuda One of ~14 failing test_main_loop_scaling parametrizations (w = BlockWise1x128 recipes). See also #188675, #188704. cc @mruberry @chauhang @penguinwu @voznesenskym @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...",https://github.com/pytorch/pytorch/issues/188715,eabbe0bea09182cafeb8da6c348ee12821462fde0eea5a628c7d38bee1b51061 references,issue,188712,issue,188675,medium,issue.body,"ast_accum_False_scaling_block_sizes1_cuda One of ~14 failing test_main_loop_scaling parametrizations (w = BlockWise1x128 recipes). See also #188675, #188704. cc @ptrblck @msaroufim @eqy @tinglvv @nWEIdia @mruberry @chauhang @penguinwu @voznesenskym @EikanWang @jgong5 @Guobing-...",https://github.com/pytorch/pytorch/issues/188712,f290676f1ab2d12c4f511352083abff8c8d2dfde47772cd0e89040927553fa49 references,issue,188712,issue,188704,medium,issue.body,"_False_scaling_block_sizes1_cuda One of ~14 failing test_main_loop_scaling parametrizations (w = BlockWise1x128 recipes). See also #188675, #188704. cc @ptrblck @msaroufim @eqy @tinglvv @nWEIdia @mruberry @chauhang @penguinwu @voznesenskym @EikanWang @jgong5 @Guobing-Chen @Xia...",https://github.com/pytorch/pytorch/issues/188712,4c5f171e6f5735599314786e8f0bbcedce80c4d199357c5c956644a673e21452 references,issue,188709,issue,188675,medium,issue.body,"fast_accum_True_scaling_block_sizes0_cuda One of ~14 failing test_main_loop_scaling parametrizations (w = BlockWise1x128 recipes). See also #188675, #188704. cc @mruberry @chauhang @penguinwu @voznesenskym @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @we...",https://github.com/pytorch/pytorch/issues/188709,970a255e54311b85311aeac6f6fb42b2e9aea2c6e665b9f70ddeefe26ce11610 references,issue,188709,issue,188704,medium,issue.body,"m_True_scaling_block_sizes0_cuda One of ~14 failing test_main_loop_scaling parametrizations (w = BlockWise1x128 recipes). See also #188675, #188704. cc @mruberry @chauhang @penguinwu @voznesenskym @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...",https://github.com/pytorch/pytorch/issues/188709,7f22e90fe7a4f84c421309be4f27ee8ed5f0f2fdaddfe5504b01287f7d528fd6 references,issue,188713,issue,188675,medium,issue.body,"ast_accum_False_scaling_block_sizes2_cuda One of ~14 failing test_main_loop_scaling parametrizations (w = BlockWise1x128 recipes). See also #188675, #188704. cc @ptrblck @msaroufim @eqy @tinglvv @nWEIdia @mruberry @jianyuh @nikitaved @walterddr @lezcano @chauhang @penguinwu @v...",https://github.com/pytorch/pytorch/issues/188713,506ac453723407ac03c04d17c5f9ed7f2effbe2e149b13e3d8cd1224c3b808c5 references,issue,188713,issue,188704,medium,issue.body,"_False_scaling_block_sizes2_cuda One of ~14 failing test_main_loop_scaling parametrizations (w = BlockWise1x128 recipes). See also #188675, #188704. cc @ptrblck @msaroufim @eqy @tinglvv @nWEIdia @mruberry @jianyuh @nikitaved @walterddr @lezcano @chauhang @penguinwu @voznesensk...",https://github.com/pytorch/pytorch/issues/188713,faa883fd8e576c81e29b73be484507092dc97967ef43815ea59cb7f1df56deb7 references,issue,188717,issue,188675,medium,issue.body,"ast_accum_False_scaling_block_sizes2_cuda One of ~14 failing test_main_loop_scaling parametrizations (w = BlockWise1x128 recipes). See also #188675, #188704. cc @mruberry @chauhang @penguinwu @voznesenskym @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @we...",https://github.com/pytorch/pytorch/issues/188717,9d6aa963f9e1b3659683551432bb0e30abe511d2cb5a6e2f058abbd3e5d11022 references,issue,188717,issue,188704,medium,issue.body,"_False_scaling_block_sizes2_cuda One of ~14 failing test_main_loop_scaling parametrizations (w = BlockWise1x128 recipes). See also #188675, #188704. cc @mruberry @chauhang @penguinwu @voznesenskym @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...",https://github.com/pytorch/pytorch/issues/188717,64eee4d83aca66f94e0f6b4b11c4304546de08d3f2a7916b64e4a00d94bb2bc6 references,issue,188708,issue,188675,medium,issue.body,"ast_accum_False_scaling_block_sizes2_cuda One of ~14 failing test_main_loop_scaling parametrizations (w = BlockWise1x128 recipes). See also #188675, #188704. cc @mruberry @chauhang @penguinwu @voznesenskym @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @we...",https://github.com/pytorch/pytorch/issues/188708,1e7aa4b708515aed1abe502d66f11ba976d81326e4cff82fddbbe69d393fb8f0 references,issue,188708,issue,188704,medium,issue.body,"_False_scaling_block_sizes2_cuda One of ~14 failing test_main_loop_scaling parametrizations (w = BlockWise1x128 recipes). See also #188675, #188704. cc @mruberry @chauhang @penguinwu @voznesenskym @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...",https://github.com/pytorch/pytorch/issues/188708,a63ac04f3aa92924d143b5f3c76de537074f8941a76015bd963bb28801d1a9b8 references,issue,188716,issue,188675,medium,issue.body,"ast_accum_False_scaling_block_sizes1_cuda One of ~14 failing test_main_loop_scaling parametrizations (w = BlockWise1x128 recipes). See also #188675, #188704. cc @mruberry @chauhang @penguinwu @voznesenskym @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @we...",https://github.com/pytorch/pytorch/issues/188716,0ea007c8d36e2ffeef844c1311a78327436c618f2324070d1743f416d8c89c4e references,issue,188716,issue,188704,medium,issue.body,"_False_scaling_block_sizes1_cuda One of ~14 failing test_main_loop_scaling parametrizations (w = BlockWise1x128 recipes). See also #188675, #188704. cc @mruberry @chauhang @penguinwu @voznesenskym @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...",https://github.com/pytorch/pytorch/issues/188716,3533f21e3e7b72b453b8187347400979a070d26555cfd388a18ddb2c8e08d22a references,issue,188710,issue,188675,medium,issue.body,"fast_accum_True_scaling_block_sizes1_cuda One of ~14 failing test_main_loop_scaling parametrizations (w = BlockWise1x128 recipes). See also #188675, #188704. cc @mruberry @chauhang @penguinwu @voznesenskym @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @we...",https://github.com/pytorch/pytorch/issues/188710,4ae6d492ad2b3c3935086523dc68b6b689c360a25dac91b1ca3984b599f5c4d2 references,issue,188710,issue,188704,medium,issue.body,"m_True_scaling_block_sizes1_cuda One of ~14 failing test_main_loop_scaling parametrizations (w = BlockWise1x128 recipes). See also #188675, #188704. cc @mruberry @chauhang @penguinwu @voznesenskym @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...",https://github.com/pytorch/pytorch/issues/188710,fdc73e68154a961c66e385be08098f7660819a75001ad588ab6a8e9a4480183e references,issue,188652,pr,188263,medium,issue.body,"he system, it downloads its own python-six sources files from the web (pypi to be exact) This makes pytorch not buildable from sources only #188263 shows the changes needed for allowing air-gapped builds To reproduce the bug simply try to build pytorch on a machine without int...",https://github.com/pytorch/pytorch/issues/188652,c64925d08a1e2b495fad1e9153b7c9abe99e2fda39eaab3f5468c1ed5467b161 competes with,issue,188425,issue,188420,medium,issue.comments[1].body,"@Blooming-Tree I noticed that #188420, #188430, and #188432 appear to have the same methodology issue. They compare eager FP64 against compiled FP32 rather than eager FP32 again",https://github.com/pytorch/pytorch/issues/188425,6030c43457e684c489aa1189093263292c95b210873f3bdec9848039f82a9e41 competes with,issue,188425,issue,188430,medium,issue.comments[1].body,"@Blooming-Tree I noticed that #188420, #188430, and #188432 appear to have the same methodology issue. They compare eager FP64 against compiled FP32 rather than eager FP32 against compil",https://github.com/pytorch/pytorch/issues/188425,dfc3ea4565e9d75e9bd4e2d1a354936446792c6c243acbf699c6d0b7b85e6376 competes with,issue,188425,issue,188432,medium,issue.comments[1].body,"@Blooming-Tree I noticed that #188420, #188430, and #188432 appear to have the same methodology issue. They compare eager FP64 against compiled FP32 rather than eager FP32 against compiled FP32, so t",https://github.com/pytorch/pytorch/issues/188425,25952773a33260bddb92f3d86fd71d0c8614fcce65ca99bf69a67ad8950ae3e5 references,issue,132845,issue,15963,medium,issue.body,"ue, please check the box on this page and assign the issue to yourself. Thanks! C10D NCCL #131203 #112815 #94003 #58856 GLOO #115740 #16295 #15963 Tcpstore #61671 #50701 MultiProcessing #65559 MultiThreadedTestCase #107425 #107424 #100459 #100448 RPC #125365 #74180 #70753 #705...",https://github.com/pytorch/pytorch/issues/132845,6c9111a02c5a0197606cd4c4228cdde6e8327e9346c654d4895eb3d1e5c8c6f5 references,issue,132845,issue,16295,medium,issue.body,"an issue, please check the box on this page and assign the issue to yourself. Thanks! C10D NCCL #131203 #112815 #94003 #58856 GLOO #115740 #16295 #15963 Tcpstore #61671 #50701 MultiProcessing #65559 MultiThreadedTestCase #107425 #107424 #100459 #100448 RPC #125365 #74180 #7075...",https://github.com/pytorch/pytorch/issues/132845,3f5de47a55f7828e5b27e907165fc3be2fdd9844956ed32d2d4924eb21942134 references,issue,132845,issue,51062,medium,issue.body,1 #50701 MultiProcessing #65559 MultiThreadedTestCase #107425 #107424 #100459 #100448 RPC #125365 #74180 #70753 #70546 #69179 #69178 #64093 #51062 DeviceMesh and DTensor DeviceMesh #132833 #123906 #122525 #122664 #121951 #121776 #121372 #120706 #120692 #110026 #98875 #98059 #9...,https://github.com/pytorch/pytorch/issues/132845,ee4a99abde9564c99b6d193c26838b234f9ec91787e275ee2a2591f21c900c0e references,issue,132845,issue,58856,medium,issue.body,"u would like to take an issue, please check the box on this page and assign the issue to yourself. Thanks! C10D NCCL #131203 #112815 #94003 #58856 GLOO #115740 #16295 #15963 Tcpstore #61671 #50701 MultiProcessing #65559 MultiThreadedTestCase #107425 #107424 #100459 #100448 RPC...",https://github.com/pytorch/pytorch/issues/132845,2fcc1403a4a723e7d1a8c76cddfc97094f497485ee4cc71d2ba7a0b7ff58fc5c references,issue,132845,issue,64093,medium,issue.body,e #61671 #50701 MultiProcessing #65559 MultiThreadedTestCase #107425 #107424 #100459 #100448 RPC #125365 #74180 #70753 #70546 #69179 #69178 #64093 #51062 DeviceMesh and DTensor DeviceMesh #132833 #123906 #122525 #122664 #121951 #121776 #121372 #120706 #120692 #110026 #98875 #9...,https://github.com/pytorch/pytorch/issues/132845,6de948c3fa1a17c69f726908fad72c292a277d6af77e747ec94e4c1149179992 references,issue,132845,issue,69178,medium,issue.body,Tcpstore #61671 #50701 MultiProcessing #65559 MultiThreadedTestCase #107425 #107424 #100459 #100448 RPC #125365 #74180 #70753 #70546 #69179 #69178 #64093 #51062 DeviceMesh and DTensor DeviceMesh #132833 #123906 #122525 #122664 #121951 #121776 #121372 #120706 #120692 #110026 #9...,https://github.com/pytorch/pytorch/issues/132845,ce44fb6f9ea09942a5f6cff4f68ee347d11719250e1d08e63c138b3c854a64a6 references,issue,132845,issue,69179,medium,issue.body,#15963 Tcpstore #61671 #50701 MultiProcessing #65559 MultiThreadedTestCase #107425 #107424 #100459 #100448 RPC #125365 #74180 #70753 #70546 #69179 #69178 #64093 #51062 DeviceMesh and DTensor DeviceMesh #132833 #123906 #122525 #122664 #121951 #121776 #121372 #120706 #120692 #11...,https://github.com/pytorch/pytorch/issues/132845,fcc954b1e21d7dfdc0b1f09f4a8123017402c67814372fb459ebc2b642715795 references,issue,132845,issue,94003,medium,issue.body,". If you would like to take an issue, please check the box on this page and assign the issue to yourself. Thanks! C10D NCCL #131203 #112815 #94003 #58856 GLOO #115740 #16295 #15963 Tcpstore #61671 #50701 MultiProcessing #65559 MultiThreadedTestCase #107425 #107424 #100459 #100...",https://github.com/pytorch/pytorch/issues/132845,04a20f55ce13ee93eb9c8f836de1361967940dc10cc7073e4b90f29a0233183a references,issue,132845,issue,103625,medium,issue.body,"14 #132447 #131446 #127373 #126868 #126852 #126493 #123294 DDP, FSDP, PiPPy FSDP1 #123726 #105024 FSDP2 #120961 DDP #106361 #104328 #104011 #103625 #77342 #77317 pipeline #121234 Other DCP #113937 #113936 Functional Collectives #107278 ShardedTensor #78068 cc @XilunWu @H-Huang...",https://github.com/pytorch/pytorch/issues/132845,7104f447a89af157a51d4b595d5cde1609425c7884a153dfe0b77f4d707f13a3 references,issue,132845,issue,104011,medium,issue.body,"in #132114 #132447 #131446 #127373 #126868 #126852 #126493 #123294 DDP, FSDP, PiPPy FSDP1 #123726 #105024 FSDP2 #120961 DDP #106361 #104328 #104011 #103625 #77342 #77317 pipeline #121234 Other DCP #113937 #113936 Functional Collectives #107278 ShardedTensor #78068 cc @XilunWu...",https://github.com/pytorch/pytorch/issues/132845,6a719ae638ac282e95ba9b87e203085b4f5bc6c23cb2aeadeffcde444daac0de references,issue,187473,issue,73764,medium,issue.body,"ributed 500/shm_broadcast hang) and 4 (CPU-MM 3 timeout) are silent on build #72925 (2026-06-18) — both passed; never filed upstream. Build #73764 status (2026-06-23) Still reproducing (fixes in pytorch main, not yet in test-channel wheel): #187484 (Inductor FxGraphCache guard...",https://github.com/pytorch/pytorch/issues/187473,f2ab20b1cd7894d89eca1ab09af87d149a93a12aedd167abdac443781319eb50 references,issue,187473,pr,184193,medium,issue.body,"eam. #187484 — [Inductor] w8a8 block-fp8 BackendCompilerFailed: inductor (DeepGEMM H100, FusedMoE H100/B200, Kernels MoE Test 1–4) · Cause: #184193 · Fix: revert #187666 (release/2.13, bisected — all 75 failures pass reverted) #187735 — [CPU multimodal] qwen2_vl multi-image to...",https://github.com/pytorch/pytorch/issues/187473,f90706cfb2736e37955d14ed1b6cce7c4c26e149d7614ae98ce25afef6ad29e3 references,issue,158968,issue,158371,medium,issue.body,"ture, motivation and pitch These enhancements are to have a better UX when using foreach_map, suggested in a few places, but most recently, #158371 (comment) These enhancements should allow easier compiler-first custom optimizer implementations. Alternatives No response Additi...",https://github.com/pytorch/pytorch/issues/158968,a06136a5d639c01df4166d0ecc0fe59446cd3972cf8b65cf24fae0d7baf8aa69 references,issue,77764,issue,141287,medium,issue.body,lace to list and track work on adding support to new ops for the MPS backend. Most requested ops (extracted from comments to this issue and #141287) PyTorch MPS Ops Project : Project to track all the ops for MPS backend. There are a very large number of operators in pytorch an...,https://github.com/pytorch/pytorch/issues/77764,43524886f1e21123a7453a8d5ab0713bb84880c7928fbbaa614e5646b57bf808 references,issue,139792,issue,101699,medium,issue.comments[0].body,"speeding up pickling itself, it would be great to have core support for arrays of strings represented as tensors (and simple data frames): #101699 If there is a million-sized list of python strings, pickle probably would be slower compared to a single tensor of concatenated ut...",https://github.com/pytorch/pytorch/issues/139792,00c944a126a20273d4aac658a2c66174e32c8e948480ea861b96d06cad8000e8 references,issue,188113,pr,188114,medium,issue.comments[1].body,"ses manual attention. It works on both AOTriton v0.11.2 and CK: a, w = attn(x, x, x, need_weights=True) # OK on both backends Updated PR at #188114 with all four files.",https://github.com/pytorch/pytorch/issues/188113,91dd135026cbb2814acedea56df2a1908c59fe3c7d459e78130daf7898a0c19e references,issue,184229,issue,175467,medium,issue.body,"uted.device_mesh import init_device_mesh from torch.distributed.tensor.parallel import ColwiseParallel, parallelize_module # Workaround for #175467 (DTensorSpec not registered as pytree constant). torch.utils._pytree.register_constant( torch.distributed.tensor._dtensor_spec.DT...",https://github.com/pytorch/pytorch/issues/184229,5551942bc14ca08353529e9362d5cbd688dbe26d12bff87952fa9e12d33b43f3 references,issue,184229,pr,184508,medium,issue.comments[1].body,There is an AI generated fix here: #184508 Which is waiting for review.,https://github.com/pytorch/pytorch/issues/184229,92976a5224c4d6aa9a3b36bac0bede9cf2ba2600b8f0bf6608f71e57d8370b41 references,issue,187953,issue,135859,medium,issue.body,"recompile under torch.compile on every new input shape, while the functional variants specialize only once. Reported originally as part of #135859 for linalg.norm and cholesky; the max / min / topk cases (structured ops) are fixed in #187952, but the composite ops need a diffe...",https://github.com/pytorch/pytorch/issues/187953,ecfef6ec584732aa9ed14763fda420b2313435bb050dd5ee5a3a838487499aeb references,issue,187953,pr,187952,medium,issue.body,"ze only once. Reported originally as part of #135859 for linalg.norm and cholesky; the max / min / topk cases (structured ops) are fixed in #187952, but the composite ops need a different fix. Minimal repro: import torch torch._logging.set_logs(recompiles=True) def f(x): out =...",https://github.com/pytorch/pytorch/issues/187953,03517d2896bd196f1b2189c6ae7890ac461dccc3b686339326c7962bc719747c competes with,issue,162588,issue,142228,medium,issue.comments[0].body,"ient-attention (cutlass FMHA) kernel that wraps when the sequence index crosses 2**16. Note a separate failure mode exists for large batch (#142228) that DOES raise (grid-dim launch error); the seq-length case here is silent, consistent with an index wrap rather than a launch-...",https://github.com/pytorch/pytorch/issues/162588,5e23dbe8436616af1f2b30e33e2db089548d1a234069bd386f39ded83cc9d68e references,issue,186350,pr,186209,medium,issue.comments[0].body,@jeffdaily Can you take a look at the draft PR? #186209,https://github.com/pytorch/pytorch/issues/186350,d6943f4901f816ddbed4ed47f96f1b3cabfa128cc53335d9ab43996a86e645b0 references,issue,188153,issue,167007,medium,issue.comments[0].body,"These two issues are similar but, as far as I can tell, not quite the same: #167007, #170672",https://github.com/pytorch/pytorch/issues/188153,16eae6e3ace15ee774e0752ce5fe9aee2048977036c39cc3769120b6261db45c references,issue,188153,issue,170672,medium,issue.comments[0].body,"These two issues are similar but, as far as I can tell, not quite the same: #167007, #170672",https://github.com/pytorch/pytorch/issues/188153,c4fe7610c5d6e5b89d6f7b674a01fd31c877f044a34779ce2beadbf4e6b72df1 references,issue,188133,pr,184481,medium,issue.body,"ic-base AOTAutograd/Inductor pattern with mutated aliased inputs. The same pattern was observed to fail on main before the guard changes in #184481, so tests added there skip only the Pallas-generated variant. Minimal shape of the failing pattern: import torch def fn(x, y): x....",https://github.com/pytorch/pytorch/issues/188133,13bbd20e0a3d6d5896155a070b6287664c2239452175fe2142a4f674980c0f78 competes with,issue,187615,issue,114299,medium,issue.body,"exactly as much copy bandwidth as they must to fit, and no more. This is the natural continuation of the per-parameter FSDP direction (RFC #114299) and is adjacent to the activation-offload work in #174960, on the parameter axis rather than the activation axis. Proposed API @d...",https://github.com/pytorch/pytorch/issues/187615,f93d1b661da8d8eb5e77fbb563c0f1cbfc6778743ac60caa5f07f553e31900a0 competes with,issue,187615,issue,174960,medium,issue.body,"more. This is the natural continuation of the per-parameter FSDP direction (RFC #114299) and is adjacent to the activation-offload work in #174960, on the parameter axis rather than the activation axis. Proposed API @dataclass(frozen=True) class PartialOffloadPolicy(OffloadPol...",https://github.com/pytorch/pytorch/issues/187615,aa57377386fd0707b4048413a9d02ee29240b8e01cf8e4ebc7b8fc9e985f0fd1 references,issue,186825,issue,72525,medium,issue.body,"log_prob (mirroring Geometric), so the forward value is exact while gradients keep the clamped path. Possibly related (same clamp family): #72525. For reference, scipy.stats.bernoulli(0).logpmf(1) returns -inf. Versions torch 2.12.0, CPU (arm64, macOS); Bernoulli.log_prob unch...",https://github.com/pytorch/pytorch/issues/186825,3c4e324629965fd44d5a12c17a8839f73367d907efcd2a1f1ad32b2fc333f137 references,issue,188020,pr,187696,medium,issue.body,"extual text to these fatal worker crash messages while preserving the existing PyTorch error message. I prototyped one possible approach in #187696 using configurable prefix/suffix text, but the exact API shape is open for discussion. I am opening this issue first, as requeste...",https://github.com/pytorch/pytorch/issues/188020,c8e6f5a432bff8d212ed86cd97a5871a7e5b2b7edf6d9bd16ce99348f6999326 references,issue,185484,issue,185470,medium,issue.body,"65625, indicating a discrepancy in how the scalar type casting or precision is handled during the kernel execution. This appears related to #185470, but it is not the same reproducer: #185470 reports F.threshold and says softshrink is fixed, while this reproducer shows F.softs...",https://github.com/pytorch/pytorch/issues/185484,8b59d6e8e6fe2d85e78edda234e55502f3f04a09a4b4f98defc1fa06fc7b2e7e references,issue,187183,pr,180208,medium,issue.body,Commit: 3aa1217 (3aa1217) PR: #171448 (#171448) Both tests removed from slow_tests.json. Open PR: Added Again Commit: 71ad4ce (71ad4ce) PR: #180208 (#180208) test_sort_stable_cpu: 1046.259033203125s test_split_cumsum_cpu: 64.4380009969s Other cases test_ddp_uneven_inputs (main...,https://github.com/pytorch/pytorch/issues/187183,ea7ba1f438e9cbc0f436236e2c5dd2a566cd537d028313797f28565a3178d19a references,issue,170839,issue,137574,medium,issue.body,e/fullgraph/CUDAGraph) Will varlen support be added to FlexAttention as well? Maybe FlexAttention could also be added as a backend in SDPA? #137574 and thus power as Triton attention impl Alternatives No response Additional context No response @drisspg cc @albanD @mruberry @jb...,https://github.com/pytorch/pytorch/issues/170839,b94f3a13809a035227001186c9fc127f944c10f0033e8cc24c6874130ca14c66 competes with,issue,187721,issue,187336,medium,issue.body,"lying torch.copysign to the output yields 1.0 in Eager mode but -1.0 in Compiled mode. Note: This issue involves a signbit discrepancy like #187336, but the manifestation is opposite (Eager is 0.0, Compiled is -0.0) and it occurs during a reduction operation rather than a poin...",https://github.com/pytorch/pytorch/issues/187721,68ff308037145e7544661ad143fb079721431c9f092253165904a9c898aaec17 references,issue,186465,issue,172568,medium,issue.body,": Importantly, this RFC does not address implementing any missing features across the HOPs. There is a separate issue tracker for this, see #172568 Any feedback is welcome. If this proposal makes sense, I am happy to embark on it. cc @chauhang @penguinwu @voznesenskym @EikanWa...",https://github.com/pytorch/pytorch/issues/186465,7b1ec46e00bd5699117975ff0f5a0a95bbd19a2b23348436c591bc98eac15104 references,issue,186465,issue,186107,medium,issue.comments[0].body,"+1, Maybe not in scope, but I also want ""consistency with jax"", see #186107 for some things jax.scan can do but ours can't",https://github.com/pytorch/pytorch/issues/186465,caf54fab54fa66105eb87bc892e3576c95b4637eacada7be7be37a1d4903f566 references,issue,146938,issue,80821,medium,issue.body,"ot callable. Also, ModuleDict does not explicitly implement the MutableMapping[str, Module] interface. But this seems to be a duplicate of: #80821 Possible/Suggested solutions Add an override of __getattr__ in ModuleDict, provided that parameters and buffers are not allowed in...",https://github.com/pytorch/pytorch/issues/146938,094aadf9f81e9c24dbcb1f8e18d9a56da77f7578c5e1cfd2bdfae8965c6232ce references,issue,186826,issue,186824,medium,issue.body,"s (or a shared helper used by subclasses) would cover all distributions. Alternatively, document that icdf performs no validation. Related: #186824 (even valid boundary quantiles q=0/1 produce wrong finite values for Cauchy/Gumbel). Versions torch 2.12.0, CPU (arm64, macOS); b...",https://github.com/pytorch/pytorch/issues/186826,e8e34d29eb25e6858b62c56b22579dac8cc6608dd35fd5b538b5b063026f4c7a competes with,issue,187178,issue,184036,medium,issue.body,TripletMarginLoss(p=1e-6) is affected identically. This is the small-p companion of the large-p +inf overflow in torch.nn.PairwiseDistance (#184036): both come from evaluating the p-norm power directly instead of in log space. import torch import torch.nn.functional as F torch...,https://github.com/pytorch/pytorch/issues/187178,f26e4fafaaffbf49586a1dac3c650907a447963be88658ca04ce71c0524a215a references,issue,187178,pr,187243,medium,issue.comments[0].body,"Raised a fix in #187243. Adds a ValueError for p < 1 in both the functional API and the module constructor, since the p-norm overflows fp32 for small p and silentl",https://github.com/pytorch/pytorch/issues/187178,dada15f9a7b6f739494135fec75a2aaa8bf12e61cc2577c9c96c80166552a7dd references,issue,187178,pr,187243,medium,issue.comments[1].body,"Raised #187243 for this. The fix rejects p < 1 in both the functional API and the module constructor, matching the documented p (int) contract and the exi",https://github.com/pytorch/pytorch/issues/187178,737f51c92c9ca0c45d9026965ffe59f7f0cd3c3ce61b0f00ef815c99814b4dd7 references,issue,187177,pr,187242,medium,issue.comments[0].body,"Raised a fix in #187242. The approach validates that k and alpha are not both non-positive at construction time (and in the functional API), since that combination",https://github.com/pytorch/pytorch/issues/187177,92a6bc63f63e71a2cbded9b7941f5ff5843bcbbf70f9f6ff4ba010f07ebe277f references,issue,187177,pr,187242,medium,issue.comments[1].body,"Raised #187242 for this. The fix validates that k and alpha aren't both non-positive when beta is non-zero, since that combination collapses the denominat",https://github.com/pytorch/pytorch/issues/187177,0df2d5e12b9f99dc2be3094eca60fc155d6061b96c3414319f7da4207c821736 references,issue,187698,issue,160230,medium,issue.body,"s and backend status, and/or Track work items for enabling Vulkan linear autograd/training support. Related broader Vulkan feature request: #160230 cc @peterjc123 @mszhanyi @skyline75489 @nbcsm @iremyux @Blackhex @nkhasbag-nv @ezyang @albanD @gqchen @nikitaved @soulitzer @Vara...",https://github.com/pytorch/pytorch/issues/187698,d08fe8a4a8c43726ba56926a5cdca94820ca2959799653e1fa34bdea041a504a references,issue,164666,issue,144040,medium,issue.body,"ed"" kernel, resulting in NotImplementedError: ""sampled_addmm_out_sparse_csr"" not implemented for 'BFloat16'. Here's a related issue I found #144040. I could try my hand at fixing this issue, but I need guidance on which approach to take. import torch # @ device = torch.device(...",https://github.com/pytorch/pytorch/issues/164666,dbb2809e8562477a4bf9719e262f919b946a23dcb8a927c2f81d4086726b50cf references,issue,187239,issue,186601,medium,issue.body,"sor For param i, rank r writes: output_offset = split_offsets[i] * world_size + r * split_sizes[i] this is inspired from proposal from rocm #186601 #!/usr/bin/env python3 """"""Minimal MORI param-contiguous all-gather layout probe. This intentionally avoids FSDP. It shows the bac...",https://github.com/pytorch/pytorch/issues/187239,a0dc8594ae4aaa6d07ac335e42e10c5cc83199048690678da75137307870627d references,issue,187239,issue,177427,medium,issue.comments[0].body,"cc @kwen2501 on param-major AG, this might share similar primitive to support non-contiguous input #177427 cc @syed-ahmed @fduwjj if you have interests as well",https://github.com/pytorch/pytorch/issues/187239,10879768b2bc4bfa688e1f9a7101e2316ec5b7d9e0891dbed308cacc94570b6e references,issue,177010,issue,158827,medium,issue.body,"our categories in a backend-agnostic way. And generalize the advanced runtime functionality, Graph capturing and replay, for accelerator on #158827. Today, generator APIs are still mostly backend-specific (torch.cuda.*, torch.xpu.*). This creates API fragmentation and requires...",https://github.com/pytorch/pytorch/issues/177010,550a7fedfc53986935fb618761d311ff61251102fdf1938ed042e5ec5780c1af references,issue,187179,issue,184036,medium,issue.body,"rm in log space via exp((1/p) * logsumexp(p * log(|diff| + eps))), finite for any p > 0. Same overflow family as torch.nn.PairwiseDistance (#184036, large p) and torch.nn.functional.triplet_margin_loss (small p). Versions PyTorch version: 2.5.1+cu121 Is debug build: False CUDA...",https://github.com/pytorch/pytorch/issues/187179,1659c054a6c53cbdaeba7b105a81c1ec2142d4592717fe3694de4e8c38b982da references,issue,136264,issue,58743,medium,issue.comments[1].body,Is this feature identical to the one being tracked in #58743?,https://github.com/pytorch/pytorch/issues/136264,f6953d26388fb4b7b162a98d2dec2d5b6359e891d1cf14dd9061412b056e5c3d references,issue,187591,issue,187590,medium,issue.body,nd recompilation would not be triggered whenever they change. Alternatives No response Additional context This would be a good follow up to #187590 as it could be implemented by a string specialized LazyConstant which materializes and guards in the appropriate situations. cc @...,https://github.com/pytorch/pytorch/issues/187591,856b5d9f055d06243c8ba3fde11a918159e242e44bbc340466d1cc55a39cabaf competes with,issue,187590,pr,170092,medium,issue.body,aph computation. Constants that are not involved in computation can be reconstructed rather than being burned in and guarded on. This stack #170092 which is partially merged contains some prior work to support this pattern. It has become difficult to merge and is probably wort...,https://github.com/pytorch/pytorch/issues/187590,53defeb3ec47c301721fc8da7b6a2d434bcf168dd351609402f4ab47e8886709 references,issue,118107,issue,88137,medium,issue.body,"3843 Loss functions Cross-entropy, log_softmax, NLL loss): #99142 Binary cross-entropy: #89125 eq, to, masked_select, index_select, narrow: #88137 LayerNorm on non-contiguous (transposed) NJT: #130538 cc @cpuhrsch @bhosmer @drisspg @soulitzer",https://github.com/pytorch/pytorch/issues/118107,35a21a6150eac21da0ba0040083731db4248caa776481345adbcac5d026222ac references,issue,118107,issue,89125,medium,issue.body,"across constant dimension: #108567 EmbeddingBag: #93843 Loss functions Cross-entropy, log_softmax, NLL loss): #99142 Binary cross-entropy: #89125 eq, to, masked_select, index_select, narrow: #88137 LayerNorm on non-contiguous (transposed) NJT: #130538 cc @cpuhrsch @bhosmer @dr...",https://github.com/pytorch/pytorch/issues/118107,96051e98ae09237efa5b9900f0afbc3d841f5f7418b9a981285fc7bee77e2247 references,issue,118107,issue,99142,medium,issue.body,"utate in-place Slicing of NTs across constant dimension: #108567 EmbeddingBag: #93843 Loss functions Cross-entropy, log_softmax, NLL loss): #99142 Binary cross-entropy: #89125 eq, to, masked_select, index_select, narrow: #88137 LayerNorm on non-contiguous (transposed) NJT: #13...",https://github.com/pytorch/pytorch/issues/118107,546fa9df1eb287366103054a7d4347eeb9543382761b91e1e5f34835024aebea references,issue,118107,issue,108567,medium,issue.body,". Prior requests: Conv2d: #114090 KV-cache related ops: #115978 Includes ops that mutate in-place Slicing of NTs across constant dimension: #108567 EmbeddingBag: #93843 Loss functions Cross-entropy, log_softmax, NLL loss): #99142 Binary cross-entropy: #89125 eq, to, masked_sel...",https://github.com/pytorch/pytorch/issues/118107,f227e9e2c3b4b56121ad44b2584041dd64dd924c4a27493f31c3298f88cb98d8 references,issue,118107,issue,114090,medium,issue.body,"ike a specific op to be implemented for nested tensors, please add a comment here so we can prioritize effectively. Prior requests: Conv2d: #114090 KV-cache related ops: #115978 Includes ops that mutate in-place Slicing of NTs across constant dimension: #108567 EmbeddingBag: #...",https://github.com/pytorch/pytorch/issues/118107,faa2bf481b35b0b41bc51811c6d7da1466707ae030f013052eeb1dd4b497fe31 references,issue,118107,issue,115978,medium,issue.body,"ented for nested tensors, please add a comment here so we can prioritize effectively. Prior requests: Conv2d: #114090 KV-cache related ops: #115978 Includes ops that mutate in-place Slicing of NTs across constant dimension: #108567 EmbeddingBag: #93843 Loss functions Cross-ent...",https://github.com/pytorch/pytorch/issues/118107,d7e4f67ab5f1553de319b3388eaf026a0448fce8061173d0dec6115d4c27a16d references,issue,118107,issue,130538,medium,issue.body,"oss): #99142 Binary cross-entropy: #89125 eq, to, masked_select, index_select, narrow: #88137 LayerNorm on non-contiguous (transposed) NJT: #130538 cc @cpuhrsch @bhosmer @drisspg @soulitzer",https://github.com/pytorch/pytorch/issues/118107,29b736daf49abc4970907434a9623fb17f80e2248bce2194b63ac2ad47e1dcdf references,issue,163429,issue,162178,medium,issue.body,🐛 Describe the bug Tracked in umbrella bug: #162178 https://github.com/pytorch/pytorch/actions/runs/17882055506/job/50851525565 this job has many failures Failure 1: 2025-09-20T17:55:32.73910,https://github.com/pytorch/pytorch/issues/163429,e4ad3720da45b776f2f1cfffe756c67e8bf831903a8cdaac8dc3940622702ed9 references,issue,163429,issue,187158,medium,issue.comments[0].body,"Failures 1 and 3 are now passing, however failure 2 is still happening and it is reported as #187158 with open PR",https://github.com/pytorch/pytorch/issues/163429,6e69a7b4a5ce9cf6f1394492df47794bea0ce49595e901c9e3034ab517fa7951 supersedes,issue,163429,issue,187158,medium,issue.comments[1].body,Failure 2 will be fixed by #187371 which supersedes the previously flagged #187158,https://github.com/pytorch/pytorch/issues/163429,f862d24467e16c4013bf6583f23addc5f5d558ae46b14a670a9baec29e16342b references,issue,187293,issue,185770,medium,issue.comments[0].body,This is a duplicate of #185770. Fix submitted in #185790.,https://github.com/pytorch/pytorch/issues/187293,531d327546f3057c627a9eea4a8817bf04591418f1362d1a8f73a0e0f412fdf5 references,issue,187293,pr,185790,medium,issue.comments[0].body,This is a duplicate of #185770. Fix submitted in #185790.,https://github.com/pytorch/pytorch/issues/187293,2a94ed17375a993d53ff6dab0a0b47f2a81ae99eb3b36433e9b3a8d4597c5e32 references,issue,187338,issue,178593,medium,issue.comments[1].body,"I remember seeing an issue about it already @malfet Thanks for taking a look! I wonder if the issue you remember might be #178593, which I reported previously. That issue was about incorrect type promotion when mixing signed (int8) and unsigned (uint8) tensors in torch",https://github.com/pytorch/pytorch/issues/187338,fb26c9d970b019a62f8169ee325368b3b14e209aef567b31b70133d621c00860 references,issue,187389,issue,60858,medium,issue.body,"vec = torch.randn((4,), dtype=dtype, device=device).unsqueeze(1) ovec = torch.triangular_solve(vec,mat) print(ovec) Related issues: #87358 #60858 Versions Collecting environment information... PyTorch version: 2.12.0 Is debug build: False CUDA used to build PyTorch: None ROCM...",https://github.com/pytorch/pytorch/issues/187389,e1b3b9b7cbd14c0b5d398628831d12ee6d1a83f68f8b3b72fee446ca2074a774 references,issue,187389,issue,87358,medium,issue.body,"e_csr() vec = torch.randn((4,), dtype=dtype, device=device).unsqueeze(1) ovec = torch.triangular_solve(vec,mat) print(ovec) Related issues: #87358 #60858 Versions Collecting environment information... PyTorch version: 2.12.0 Is debug build: False CUDA used to build PyTorch: No...",https://github.com/pytorch/pytorch/issues/187389,4bc4dd16252905cd4315de076678b7674a2306a791265fbf8d2254a779ac571c references,issue,176298,issue,176296,medium,issue.body,"ted for 'UInt16' print(a * b) # NotImplementedError: ""mul_stub"" not implemented for 'UInt16' Same for uint32 and uint64. Context Related to #176296 (same gap on MPS). PR #159094 added foundational support for these types but binary op kernels were not registered on either back...",https://github.com/pytorch/pytorch/issues/176298,9788cabba50d613e8ab29efad560ec45347f73746db9335b7e73619bd8c57d7e references,issue,186537,pr,186583,medium,issue.comments[0].body,cc @rtimpe can you see if #186583 fixes the issue?,https://github.com/pytorch/pytorch/issues/186537,bf93f1da58517dde42903ea259c2d86e3ac07328cd2085a0ead75465368c6ea3 references,issue,162358,pr,181516,medium,issue.comments[1].body,"ourse, it's a fun bug. Have a draft MR up, with some notes on the justification and considerations of various alternate possible approaches #181516.",https://github.com/pytorch/pytorch/issues/162358,46d1fe7c89d3c7439851ea51f01e060ad78b79eb94a452c399cb208f524b7d31 references,issue,187094,issue,165230,medium,issue.body,"on GB300 — we measured an INT8-quantized MLP at 0.17x the speed of its fp32 baseline, dominated by the _int_mm calls. Related but distinct: #165230 reports the row-major-rhs layout penalty (reproduced here as the 2.08x col-vs-row gap). This issue is about the sm_103 dispatch i...",https://github.com/pytorch/pytorch/issues/187094,2a49f7d981f51b1092b39b6b0da70e78ee586302a24daacbdb5c8ca85419b87d competes with,issue,187266,pr,176265,medium,issue.body,"tays 3D, so every attention call takes the slow math path instead of flash/efficient/cuDNN. This is the same per-example-vmap motivation as #176265 (backward batching rule), but for the forward eligibility check rather than the batching rule itself. Alternatives Add the batch...",https://github.com/pytorch/pytorch/issues/187266,a13dfb81042fa223b2258cda30ffaf55740a237d15f9cef981ca6f68cba75090 references,issue,179374,issue,160875,medium,issue.comments[0].body,Also related on first-class support for uv in pytorch: #160875 (comment),https://github.com/pytorch/pytorch/issues/179374,1c2a44e01a2da7837be5ade64b5eb6b46730c57cf74531082f7d6887c85887a4 references,issue,186837,pr,186835,medium,issue.body,es. A fixed set of representative benchmarks for these paths would catch broad regressions. Comments and suggestions welcome. Companion PR: #186835 Alternatives No response Additional context No response cc @malfet @pytorch/pytorch-dev-infra,https://github.com/pytorch/pytorch/issues/186837,72f088282a925ac5c30ec388dbb1c08398980de549df1fa4adc272eb310c8532 references,issue,187008,issue,145529,medium,issue.body,"he EMA weakref lives in the EMA object, not in any fake-tensor cache. #186796 — weakref comes from Dynamo guards on a live compiled module. #145529 — requires set_swap_module_params_on_conversion(True) and is inside @torch.compile. Here the flag is off and there is no compile....",https://github.com/pytorch/pytorch/issues/187008,feadd59ed5fc3db58059f74153cbce875e67b439a446443cfbe6029db27ca0f4 references,issue,187008,issue,170137,medium,issue.body,. #145529 — requires set_swap_module_params_on_conversion(True) and is inside @torch.compile. Here the flag is off and there is no compile. #170137 — same spirit (a tool's weakref blocks swap) but for the memory profiler on regular tensors. So an EMA (or any user-held weakref)...,https://github.com/pytorch/pytorch/issues/187008,61e7b9c9f26262fa1583469b6f0e99caa303f52d0cf3d9b160d240bf2a209b94 references,issue,187008,issue,170770,medium,issue.body,"forces-swap / swap_tensors-rejects-weakref limitation as a few other issues, but the weakref source — and therefore the fix — is different: #170770 / #141548 — weakref comes from FakeTensorMode / MetaConverter after torch.compile. The fix discussed there (clear the FakeTensorM...",https://github.com/pytorch/pytorch/issues/187008,cb98e57674ba2de6d74b5a77303d3479e62a52f0bdf8a5e767a1c68ae3cce0b5 references,issue,187008,issue,186796,medium,issue.body,"Mode memos at end of compile, e.g. #171209) does not help this case: the EMA weakref lives in the EMA object, not in any fake-tensor cache. #186796 — weakref comes from Dynamo guards on a live compiled module. #145529 — requires set_swap_module_params_on_conversion(True) and i...",https://github.com/pytorch/pytorch/issues/187008,c947d9b058f2a22873c8366289944d0061f8b0d43ab55ae6da2d6a1112fe7f9b references,issue,187008,pr,187090,medium,issue.comments[0].body,"Hello, I have submitted a pull request resolving this issue here: #187090 Technical Summary of the Fix: Root Cause: The RuntimeError originates from the Python-level checks (weakref.getweakrefs) inside torch.utils",https://github.com/pytorch/pytorch/issues/187008,569e9fb74c894665106040865a203f1f653a2496f33390669c85ab4bf8eb45eb references,issue,166638,issue,152406,medium,issue.comments[0].body,Seems duplicate to #152406.,https://github.com/pytorch/pytorch/issues/166638,668d57f8cd07dcf522f3f0a5dd80494e6339f433558368ca6583a37cfa4bf4ad references,issue,156217,issue,157495,medium,issue.body,"ops only included that for a certain target shardings. To achieve this, we require OpSpec always include tensor_meta for input and output. #157495 OpStrategy should return all valid OpSpec's. OpStrategy should never raise error when finding unsupported input strategy in OpStra...",https://github.com/pytorch/pytorch/issues/156217,02f7d6fe4a4763f430d6846c4c40c0acbfe0043ec69ed5334a38113484e23e5a references,issue,184195,pr,184308,medium,issue.comments[0].body,n change the generated code to call aoti_torch_call_dispatcher for dispatcher look up when the implementation doesn't exist. something like #184308? cc @desertfire,https://github.com/pytorch/pytorch/issues/184195,1fd1968a309093fa42d5bc3340b11f1f657ddceced096e885681b5b4eb8b083a references,issue,50688,issue,46168,medium,issue.body,"ained scan/reduce/map ops could give users extra confidence. Also memory could be economized at backprop if shapes are guaranteed to match: #46168 This may also be interesting for differntiating through optimization loops, since training loop may be represented as such ops as...",https://github.com/pytorch/pytorch/issues/50688,c528edd05b769d3d8830fc9cce36c5800c5ac93d2d7e19b6db5d4a37ec1c73a4 references,issue,50688,issue,95408,medium,issue.body,an ONNX op: https://github.com/onnx/onnx/blob/master/docs/Changelog.md#Scan-11 Related feature request on parallel fast associative scans: #95408 Could possibly be a way for coding of custom RNN that is guaranteed/confident for correct tracing/export. + the code becomes more f...,https://github.com/pytorch/pytorch/issues/50688,02d04314daaa00cce7b4aadcee3f06dfc715a02ad7134e387a05521575af5084 references,issue,153484,pr,153557,medium,issue.comments[0].body,Potential fix here (needs more design): #153557,https://github.com/pytorch/pytorch/issues/153484,94f0487b7cf2c8d0ede360d4682da70b2ff430ef281fa7eea37573ce52e2e56c competes with,issue,186256,pr,176234,medium,issue.body,"sition), or — as a strict improvement until then — it should raise rather than return silently-wrong values (cf. the layer_norm proposal in #176234). Related #175754 — same bug for native_layer_norm (closed; root cause documented there) #176234 — proposed layer_norm guard; rev...",https://github.com/pytorch/pytorch/issues/186256,95635c02f84b1b4ba79d2243297f7eedf8fd70efe93baec18a4b98b54deaf147 references,issue,186256,pr,176234,medium,issue.comments[1].body,"sing DelayedError, with a focused regression test and an explicit compiled-autograd skip so it does not repeat the CI failure that reverted #176234.",https://github.com/pytorch/pytorch/issues/186256,357a0be63558f62c96566adddecec5085f63002f2a6ab564ac786b08c2389415 references,issue,178153,pr,183276,medium,issue.body,N Introduce shared heuristics/registry.py with both template and codegen registration APIs; template/registry.py becomes re-export shim PR: #183276 3/N Extract pointwise heuristics from triton_heuristics.py into triton_codegen/pointwise.py with PointwiseHeuristic + ROCmPointwi...,https://github.com/pytorch/pytorch/issues/178153,c2cecaaee019ba0325f5062cff087b38458dce3d5062e5c23186a7b9437f83c3 references,issue,178153,pr,183277,medium,issue.body,ics from triton_heuristics.py into triton_codegen/pointwise.py with PointwiseHeuristic + ROCmPointwiseHeuristic + XPUPointwiseHeuristic PR: #183277 4/N Extract reduction heuristics from triton_heuristics.py into triton_codegen/reduction.py with ReductionHeuristic + ROCmReducti...,https://github.com/pytorch/pytorch/issues/178153,81801edad4bf0832d1087dd1212ac12b104488eab0d2a9b7b567b908bae4d01f references,issue,178153,pr,183278,medium,issue.body,ics from triton_heuristics.py into triton_codegen/reduction.py with ReductionHeuristic + ROCmReductionHeuristic + XPUReductionHeuristic PR: #183278 Alternatives No response Additional context cc @jerryzh168 @chauhang @penguinwu @voznesenskym @EikanWang @jgong5 @Guobing-Chen @X...,https://github.com/pytorch/pytorch/issues/178153,749173131de385d70401cf098844421633105781d1a999f6ecc5371760d1d8b5 references,issue,140069,issue,88838,medium,issue.comments[0].body,Depends on DTensor Strided Sharding: #88838 (comment) #129627,https://github.com/pytorch/pytorch/issues/140069,7a59f71d515055391da9608ae2812df0224ab94d729f9b047d34f888ed84cccc references,issue,140069,issue,129627,medium,issue.comments[0].body,Depends on DTensor Strided Sharding: #88838 (comment) #129627,https://github.com/pytorch/pytorch/issues/140069,87efc6369fe2e30964565c328e4c3b546ee8e68191ef7ccc99e8a3e72a881bc3 competes with,issue,137014,issue,136312,medium,issue.body,"on and pitch Advanced users want to control ProcessGroup abort and recovery, instead of being controlled by watchdog (such as SIGABRT). See #136312 for such ask. In this situation, an abort group API is necessary because that's the only way to stop a spinning NCCL kernel. Ther...",https://github.com/pytorch/pytorch/issues/137014,f18d56df822bec0ccad5f843e8e2c2832d1cfed080901c26dc5f528b88f33799 references,issue,186548,issue,171938,medium,issue.body,"4N both train cleanly under it on Sunspot. But it's a band-aid we'd like to delete once the upstream override + split() impl land. Related #171938 — a sibling issue (""Device id in distributed process group with multiple backends"") that touches the same device_id + subgroup-cre...",https://github.com/pytorch/pytorch/issues/186548,9dd20c5028540d349b75f6225d7f303d87a2a84c59fe37bc8f2c2299862f0865 references,issue,168044,issue,162591,medium,issue.body,"🚀 The feature, motivation and pitch RFC: Improving the Internal Consistency of PyTorch LR Schedulers (Inspired by Issue #162591) Summary While trying to understand and fix the behavior described in issue #162591, where SequentialLR resumes training with completely un",https://github.com/pytorch/pytorch/issues/168044,3dcc61d30fc541bf9bbbb8a09406e0fd113492ef4e1b7272b5e967a7786acbba references,issue,168044,issue,29697,medium,issue.comments[0].body,About closed-form / purely-functional scheduler formulas: #68332 #29697,https://github.com/pytorch/pytorch/issues/168044,babca9b559d17c9d4bf9b7dd31ec2b6be546a71235f956e46c7a778355837977 references,issue,168044,issue,68332,medium,issue.comments[0].body,About closed-form / purely-functional scheduler formulas: #68332 #29697,https://github.com/pytorch/pytorch/issues/168044,1a2d749961ab8496799ff95373e66933c79616c86f29ba6ea133f9d649d33d10 competes with,issue,171938,issue,186548,medium,issue.comments[1].body,"Cross-linking a sibling issue I just filed: #186548. Same theme (xpu + device_id binding + subgroup creation), but the xccl-only init path instead of the multi-backend xpu:xccl,cpu:gloo path",https://github.com/pytorch/pytorch/issues/171938,e57e50beeac568eee6e0509dfe9d6871fa5203b8c927754c6b4a33d326ddabf3 references,issue,73332,issue,66504,medium,issue.comments[0].body,Related: #66504,https://github.com/pytorch/pytorch/issues/73332,88d4b2c77d9758aa7dcd431bed1fe50a4d63b0fe64c571ae0e99cce4e314c5cb references,issue,73332,issue,66504,medium,issue.comments[1].body,"cc @zhaojuanmao , we should investigate: Why converting to syncBatchNorm doesn't work, but it is recommended as a workaround in #66504 Whether custom buffer communication hooks implemented in https://github.com/pytorch/pytorch/blob/master/torch/nn/parallel/distributed.py#L1",https://github.com/pytorch/pytorch/issues/73332,1c8cd07107d917c7234fa5c45e99f45c2645bec4fa8994cc9f0254a0d0761648 references,issue,73960,issue,59552,medium,issue.comments[0].body,"Thanks for reporting this issue! I believe isend/irecv APIs are known to have a couple issues as of now, see: #68866, #64688, #63486, #59552, NVIDIA/nccl#501 To confirm that this is an issue with isend/irecv APIs, could you switch the code to doing a simple dist.barrier() after p",https://github.com/pytorch/pytorch/issues/73960,01be3dec8dc4a01aa29638b0c14747452bae54cac9ac5ff99a41f5c69340c1d9 references,issue,73960,issue,63486,medium,issue.comments[0].body,"Thanks for reporting this issue! I believe isend/irecv APIs are known to have a couple issues as of now, see: #68866, #64688, #63486, #59552, NVIDIA/nccl#501 To confirm that this is an issue with isend/irecv APIs, could you switch the code to doing a simple dist.barrier()",https://github.com/pytorch/pytorch/issues/73960,bfabcc9255577ee80b7555c450ad6700c277465332e8898bd90345f1fa0cfb89 references,issue,73960,issue,64688,medium,issue.comments[0].body,"Thanks for reporting this issue! I believe isend/irecv APIs are known to have a couple issues as of now, see: #68866, #64688, #63486, #59552, NVIDIA/nccl#501 To confirm that this is an issue with isend/irecv APIs, could you switch the code to doing a simple dist.b",https://github.com/pytorch/pytorch/issues/73960,e5ac7be7c3b7a88ebec88283ad948028d1edd14dbcd72dda92559a236770aad0 references,issue,73960,issue,68866,medium,issue.comments[0].body,"Thanks for reporting this issue! I believe isend/irecv APIs are known to have a couple issues as of now, see: #68866, #64688, #63486, #59552, NVIDIA/nccl#501 To confirm that this is an issue with isend/irecv APIs, could you switch the code to doing a simpl",https://github.com/pytorch/pytorch/issues/73960,8f0e9884d86aa2e833657d2dbcf3861018d32c348d900fe963bbbebba8313c30 references,issue,74528,issue,74734,medium,issue.comments[0].body,The original issue: #74734. Maybe we can use this as the feature request.,https://github.com/pytorch/pytorch/issues/74528,3d7076992f73038653a2cdfe8058fc7decffe5279a102fac146975346c98387a references,issue,74734,issue,74528,medium,issue.comments[1].body,New feature request is here: #74528,https://github.com/pytorch/pytorch/issues/74734,2c8bcb6e3d769904029f21552dfee3e7ba589fa18b85ebe34c1984fd36b9a0d5 references,issue,75725,issue,72948,medium,issue.body,"varma @gqchen @aazzolini @osalpekar @jiayisuse @SciPioneer @H-Huang @albanD , @ezyang , @mruberry , it is most likely linked to this issue: #72948 Thanks a lot! Versions Collecting environment information... PyTorch version: 1.11.0+cu113 Is debug build: False CUDA used to buil...",https://github.com/pytorch/pytorch/issues/75725,c3e691e9670327a28e19f89ad2cad3a14bc31b2ebf632dc70bd4008d04f44706 references,issue,76103,issue,70160,medium,issue.body,"DEBUG: UNSET I0420 13:19:03.251551 238663 ProcessGroupNCCL.cpp:728] [Rank 0] NCCL watchdog thread started! op_registration_test is probably #70160 nnpack_test occasionally failed. ProcessGroupNCCLTest consistently failed, but running the executable directly is fine. Sequential...",https://github.com/pytorch/pytorch/issues/76103,a68db8bc1cb5435ec9c29e64f7bbfdda3e4777bfd77ee1781b1497999bb30c15 references,issue,76282,issue,67590,medium,issue.body,"sure if any other optimizer wrapper may need to run certain communications even if no local optimizer step. One relevant feature request is #67590, which proposes to signal that the scaler has been stepped and optimizer.step() has been actually called. Note that this issue may...",https://github.com/pytorch/pytorch/issues/76282,742e8757c57ec2a72a26c2d50edfc5d0f33a677f89b25fe2c2c395cbac4e4881 references,issue,77154,issue,23430,medium,issue.comments[0].body,"Related: #23430, Lightning-AI/pytorch-lightning#12866 if generic distributed sampler wrapper gets created, then probably existing multinomial sampler can b",https://github.com/pytorch/pytorch/issues/77154,4ef1d22d086dc4757ce17b42c364d793513276262a5bd803171f4c60b65f6140 references,issue,77154,issue,23430,medium,issue.comments[1].body,") I see #23430 is related, if I can supply a WeightedSampler object into a DistributedSampler class to get a DistributedWeightedSampler. I will wait for a",https://github.com/pytorch/pytorch/issues/77154,523f0f12557b763361e5e223be90d14977d8bd4791103ed55297ab01e4a74fae references,issue,80832,issue,47260,medium,issue.body,"C._EngineBase object at 0x7fb37a2ec200> returned NULL without setting an error But if I call the prepare_for_backward before backward (like #47260), ddp_model.reducer.prepare_for_backward(loss) loss.backward(retain_graph=True) ddp_model.reducer.prepare_for_backward(output) out...",https://github.com/pytorch/pytorch/issues/80832,81bf60ffce180b424f5c1ee18caea4cdc9bb9aece9519798e91cba7f581d4847 references,issue,112164,issue,105499,medium,issue.body,"ariable_builder, tx.fake_mode.from_tensor( -> r = self.meta_tensor DDP: DDP, autocast (elastic DDP) #111794 Previous related issues #105348 #105499 #103862 #102731 #96372 Possibly related: FSDP, grad_enabled (not torch.compile) #111958 FSDP, grad_enabled (not torch.compile, In...",https://github.com/pytorch/pytorch/issues/112164,38ae2c329932176d1e54891a0b79534b7848c18490da505b294d1f07a522504d references,issue,112164,issue,111317,medium,issue.body,"ug FSDP FSDP, autocast (MosaicML Diffusers) #110797 Error raised: aot_autograd, r.grad = self.meta_tensor FSDP, autocast (LlamaForCausalLM) #111317 Error raised: fsdp/flat_param.py, flat_param.grad = flat_param_grad FSDP, autocast (Llama2, LlamaRotaryEmbedding) #108211 Error r...",https://github.com/pytorch/pytorch/issues/112164,b95b57a377a1a583e571bf4476f6241e0d780da675d5eff1a2ee162118b7d780 references,issue,112164,issue,111552,medium,issue.body,"99 #103862 #102731 #96372 Possibly related: FSDP, grad_enabled (not torch.compile) #111958 FSDP, grad_enabled (not torch.compile, InternLM) #111552 Likely the same root cause - improper global state / side-effects handling when splitting the graph Solution May not be solvable...",https://github.com/pytorch/pytorch/issues/112164,448d78b021cb0043372c3ac1822d7e0efc8a62f9afa8ab7b27889d35a6818a59 references,issue,112164,issue,105499,medium,issue.comments[0].body,"these global states should only affect compute and backward graph building, not communication, I have a suspicion that they do in fact. See #105499 @awgu would you be able to comment?",https://github.com/pytorch/pytorch/issues/112164,a2167200248ce2131d08e85310e75b2bdc7ef50ad88b3c8860c8988f3b636650 references,issue,146326,issue,146328,medium,issue.body,"requires communication APIs that accept GPU metadata (they don't exist today), and ragged gemm API that reads offsets from the GPU memory. #146328, #146329 Block quantization for fp8 #146368 Hierarchical implementation for the a2a communication to reduce cross-node traffic #14...",https://github.com/pytorch/pytorch/issues/146326,558d7d722d4156ef517966cf96ec7df1c12e95188449ec4786ea9d21d83c084f references,issue,146326,issue,146329,medium,issue.body,"communication APIs that accept GPU metadata (they don't exist today), and ragged gemm API that reads offsets from the GPU memory. #146328, #146329 Block quantization for fp8 #146368 Hierarchical implementation for the a2a communication to reduce cross-node traffic #146331 MLA...",https://github.com/pytorch/pytorch/issues/146326,8cd0ad9a27167d06a7f44a9cb917a88876c6e08ab97f91012a752081dd17f58f references,issue,146326,issue,146330,medium,issue.body,"implementation for the a2a communication to reduce cross-node traffic #146331 MLA attention, currently not implemented in an efficient way #146330 Overlapping communication and computation - DeepSeek uses fine-grain overlapping strategy, where forward of one microbatch is over...",https://github.com/pytorch/pytorch/issues/146326,643519f1ee4bb402608a775f67d0021039ecbe5c18e7f0222142a45ab0715dd7 references,issue,146326,issue,146331,medium,issue.body,"ory. #146328, #146329 Block quantization for fp8 #146368 Hierarchical implementation for the a2a communication to reduce cross-node traffic #146331 MLA attention, currently not implemented in an efficient way #146330 Overlapping communication and computation - DeepSeek uses fi...",https://github.com/pytorch/pytorch/issues/146326,c1ae6bf4324614c35e6a8d14df76e1f19843f3343b055b207873125d33601d46 references,issue,146326,issue,146332,medium,issue.body,"ly don't have a way to conveniently express that, possibly torch.compile could help but there is a wide open space for design options here. #146332 More flexibility for mixed precision optimizers (e.g. DeepSeek mentions that they keep optimizer states in bf16) #146542 cc @awgu...",https://github.com/pytorch/pytorch/issues/146326,9f03e5230de37de8382a7214d1b4592d5839233900c8bbb62bca57c300882534 references,issue,146326,issue,146368,medium,issue.body,"metadata (they don't exist today), and ragged gemm API that reads offsets from the GPU memory. #146328, #146329 Block quantization for fp8 #146368 Hierarchical implementation for the a2a communication to reduce cross-node traffic #146331 MLA attention, currently not implemente...",https://github.com/pytorch/pytorch/issues/146326,dc5310b1aab375cef8baee65fbc226c3352ef6f7140d85225b299f7d0fa178c9 references,issue,146326,issue,146542,medium,issue.body,sign options here. #146332 More flexibility for mixed precision optimizers (e.g. DeepSeek mentions that they keep optimizer states in bf16) #146542 cc @awgu @wanchaol @fegin @fduwjj @wz337 @wconstab @d4l3k @pragupta @msaroufim @dcci @aditvenk @weifengpy @drisspg @liangel-02 @h...,https://github.com/pytorch/pytorch/issues/146326,077d971a86bdf34a9d0af1b9f34447cfbc3bc798a96a807a5f4ce78cd0b20891 references,issue,146326,issue,146328,medium,issue.comments[1].body,"@eqy let's take this discussion to #146328, the API you linked is for fp32/fp64 only, and has sizes/pointers on the CPU, not GPU.",https://github.com/pytorch/pytorch/issues/146326,e8a465aab2de8153159bc50a86b2516ea0de4f71a64e19ad292901caf6308c5e references,issue,152822,issue,115305,medium,issue.comments[0].body,Is it possible to consider this option in conjunction with #115305,https://github.com/pytorch/pytorch/issues/152822,82e1db582359bb6eaa57ae725b4683a96417af00585dc88b89323432172a768c references,issue,162179,pr,162706,medium,issue.body,---------------------------------------------------------- Ran 0 tests in 0.000s NO TESTS RAN The issue with the decorator will be fixed by #162706 The test itself fails however. I assume it got broken which went unnoticed as it wasn't run Versions At least PyTorch 2.7.1 up to...,https://github.com/pytorch/pytorch/issues/162179,bd20fa9211fecc20bb12f7c5d387f4c86a7fbdb4c77840f14dcc3c965c3d2072 references,issue,162179,pr,162706,medium,issue.comments[1].body,I already worked on this in #162706 after the previous one caused failures due to a typo and had to be reverted. The change is simply redefining the function: def skip_if_lt_x,https://github.com/pytorch/pytorch/issues/162179,7e5af40a053fed85ad2c547eb5f2e7f036b8918377e38279774d9debc4d12ef0 references,issue,162745,issue,162178,medium,issue.body,"🐛 Describe the bug Tracking this in Umbrella Bug: #162178 Job link: https://github.com/pytorch/pytorch/actions/runs/17470577491/job/49628024468 The two tensors are ""close"" but not close enough it s",https://github.com/pytorch/pytorch/issues/162745,ec61ec6bf61d80c7d98cd3448fe66fee44854650b2a5bcb686e849e18c732d40 references,issue,162748,issue,162178,medium,issue.body,🐛 Describe the bug Tracked in #162178 Job link: https://github.com/pytorch/pytorch/actions/runs/16764647752/job/47468897559 On a 8-GPU runner (e.g. 8xB200): the following passes,https://github.com/pytorch/pytorch/issues/162748,2fbc305c0a20b9698454ed8930018b9a526d103191134f55781fde9dfbe0d5d0 references,issue,162871,issue,162178,medium,issue.body,to decorate test_gather_object with require_exact_world_size_of_4. Or simply skip it for now. Versions TOT Tracking this in umbrella issue: #162178 cc @awgu @wanchaol @fegin @fduwjj @wz337 @wconstab @d4l3k @pragupta @msaroufim @dcci @aditvenk @weifengpy @malfet @pytorch/pytorc...,https://github.com/pytorch/pytorch/issues/162871,90bb1a6764b54708eafce3a349ec514cf3451e5fb45256d030a999a814ac1c38 references,issue,162897,issue,162178,medium,issue.body,🐛 Describe the bug Tracking via umbrella bug: #162178 Failure Job link: https://github.com/pytorch/pytorch/actions/runs/17680564307/job/50281580313 Failure message: ` 2025-09-13T08:32:02.378563,https://github.com/pytorch/pytorch/issues/162897,e1b0d05d53a73185129a8a681798020ec40d9d739e41a1996f4a14f05fc206c1 references,issue,162917,issue,162178,medium,issue.body,"🐛 Describe the bug Tracked in B200 Umbrella Bug: #162178 Job link: https://github.com/pytorch/pytorch/actions/runs/17706537397/job/50336308097 Error message: First, excessive amount of messages (E",https://github.com/pytorch/pytorch/issues/162917,39cd3abda3b38defefcf593e0ba9a46a6fa853334cd6df10036c7073980bba9d references,issue,162940,issue,162178,medium,issue.body,🐛 Describe the bug Tracking this via: #162178 Platform: B200 Job link: NA - manual run with the following: python test/distributed/test_symmetric_memory.py AsyncTPTest.test_fused_scaled,https://github.com/pytorch/pytorch/issues/162940,adfbba877772695f3db88c80830cc44f1c619a7b59d1a542415523b2b4063855 references,issue,162940,issue,162917,medium,issue.body,"8, 0.0747, ..., 0.0752, 0.0732, 0.0732]]], device='cuda:0', dtype=torch.bfloat16)) This error would have been encountered if it was not for #162917 2025-09-14T16:15:45.6935514Z distributed/test_symmetric_memory.py::AsyncTPTest::test_fused_scaled_matmul_reduce_scatter_scatter_d...",https://github.com/pytorch/pytorch/issues/162940,cf366f8eb52226426e0c6b84a61ce487ce8f0abfd855d8dce3972869344c2b3e references,issue,165727,issue,145949,medium,issue.body,may be related to how DDP handles tied/shared parameters on Blackwell architecture Novel bug - No documented cases found in PyTorch issues #145949 (Blackwell tracking) or #157549 (Blackwell compatibility) Frameworks tested (ALL exhibit same hang): torchrun with native DDP Hugg...,https://github.com/pytorch/pytorch/issues/165727,13cd16af92bd54c4fd6f220b13c34f52f1fbbdf38ebc85bc5c5e29b89159a3d2 references,issue,171412,issue,155992,medium,issue.comments[0].body,"This new mode should also be more practical for loading from S3: #155992 And I also wonder, if HF model AutoLoader can do the same - to avoid hitting the HF servers more times than needed, before the weights are",https://github.com/pytorch/pytorch/issues/171412,feb9ff497ff4c50c253d472672c5506f7ed006b9cabe30d5b0f56f340a81b7dc references,issue,171412,issue,155992,medium,issue.comments[1].body,"ode should also be more practical for loading from S3: [feature request] Native checkpointing to/from s3:// (DCP / torch.load / torch.save) #155992 And I also wonder, if HF model AutoLoader can do the same - to avoid hitting the HF servers more times than needed, before the we...",https://github.com/pytorch/pytorch/issues/171412,f9c347a3be5eb82c4ec2b4dc1d12d02e44d0dfc9a5526baaf207ead0081a6c37 references,issue,171604,issue,138422,medium,issue.body,"to unwrap functorch's TensorWrapper during JVP execution, but wait_tensor cannot access the storage of wrapped tensors. Similar to #161943 #138422 Code to Reproduce import torch from torch.distributed._functional_collectives import _maybe_wrap_tensor inp = torch.randn(1) class...",https://github.com/pytorch/pytorch/issues/171604,b0543acd0aadcc404e1172bd10a6c3059b83a848027ddb03582c451dc0939370 references,issue,173362,issue,173164,medium,issue.body,"e() or transpose()), even though the implementation handles this by calling .contiguous() internally. This is similar to the issue fixed in #173164 for all_gather, but affects reduce_scatter instead. import os import torch import torch.distributed as dist import torch.multipro...",https://github.com/pytorch/pytorch/issues/173362,1fe222f62af75c875323c1b9191eef430692d1353551abb3def757e584349806 references,issue,185888,pr,186205,medium,issue.comments[0].body,Opened a fix: #186205,https://github.com/pytorch/pytorch/issues/185888,b2a5075b9e1a94ec4b005a77d89c9514b176b4d8bf429993e7e6cd5e8fac7786 references,issue,186099,issue,184352,medium,issue.body,"Background Related to #184352 (Python 3.15 support tracking) Description The triton, triton-rocm, and triton-xpu nightly wheels published to download.pytorch.org cannot",https://github.com/pytorch/pytorch/issues/186099,efcc1348a699fbd7136570bc85deb68725476025a2419561e8d895c444c8c70f references,issue,168652,issue,168651,medium,issue.body,file: test/inductor/test_select_algorithm.py Test class: TestSelectAlgorithm Test name: test_mm_dropout Defined at line: 329 Parent issue: #168651 Code reference: https://github.com/pytorch/pytorch/blob/main/test/inductor/test_select_algorithm.py cc @sunway513 @jithunnair-amd...,https://github.com/pytorch/pytorch/issues/168652,9246e65978635b54d2522306c2d5a434b3fa5dd21aac654151e07b3c8b40c75d references,issue,185780,issue,184361,medium,issue.body,dering if something like that at a smaller scale could be implemented for the pipelining module Alternatives No response Additional context #184361 (comment) cc @awgu @wanchaol @fegin @fduwjj @wz337 @wconstab @d4l3k @pragupta @msaroufim @dcci @aditvenk @weifengpy,https://github.com/pytorch/pytorch/issues/185780,f927aae409bdc2615b5205cba04e43dc05ea607e8d6780e379f6fd70f58e3425 references,issue,185240,issue,16706,medium,issue.body,"ailures — including both the bnb cold-start race (bitsandbytes#1936) and the L4T 36.4.7 NvMap allocation-tracking cap (llama.cpp discussion #16706, NVIDIA forum 349752, CVE-2025-33177). Both populations hit the byte-identical crash signature, so a single PyTorch fix addresses...",https://github.com/pytorch/pytorch/issues/185240,8cbc9d756aa29e50082d386cf8b4be429b86821957fd278fd4df83d68d398f50 references,issue,183434,pr,185813,medium,issue.comments[1].body,"blematic to include either too few or too many license files (and especially if they're incompatible licenses of course). @zklaus @rgommers #185813 implements #183434 (comment) (explicit license-files + audit test, Option 2). This also fixes pip install torch failures on Windo...",https://github.com/pytorch/pytorch/issues/183434,3c9d3d0525f6be49a9c63c898855521fc7513dac086effef1fcd0537da4a5ef5 references,issue,170487,issue,167729,medium,issue.body,"🚀 The feature, motivation and pitch Currently it fails, no matter how I try to hack around: #167729 (comment) This is often needed to do torch.autograd.grad calls inside torch.compile, in order to satisfy its requirements (it doesn't like",https://github.com/pytorch/pytorch/issues/170487,736eea028f6c03e19f89c2ebd7c8206c29f40b64771d4f6a021b2b294ba2413d references,issue,170487,issue,167729,medium,issue.comments[1].body,at these tensors should not be visible from outside of the compiled region because we don't know how to recreate autograd state. . Related: #167729,https://github.com/pytorch/pytorch/issues/170487,b09bd3722d6e274dbcddd8f6d7dbf39b64f2265c222aa5948a4c078d06db02fb references,issue,149534,issue,24826,medium,issue.comments[0].body,"Relevant older issue: #24826 Note that we generally maintain a high bar for inclusion of new modules into torch.nn, as each addition comes with a steep maintenance cost",https://github.com/pytorch/pytorch/issues/149534,337205566f70eb0975178098bda1ebe237f64329ea350a8cf369672fb685d32b references,issue,173225,issue,86076,medium,issue.body,PS would be taking over the namespace which would make it clearer where the developer who wants to add a Metal kernel should head over to. *#86076 cc @kulinseth @malfet @DenisVieriu97 @aditvenk,https://github.com/pytorch/pytorch/issues/173225,9a00a3041fcfd958c140ce432b37a34bf2f4966caaaf5a2e9b9a12c268e1f2af references,issue,185590,issue,174469,medium,issue.body,"act same goal as this refactoring initiative and provide crucial underlying technical context, serving as perfect complements to this plan: #174469 #185142 Looking forward to great discussions and contributions in the Slack channel and PR reviews! cc @mruberry",https://github.com/pytorch/pytorch/issues/185590,f98deb180257fa79c2f06ab448b09fb2322c81df7f7179711299407d2cc2a641 references,issue,185590,issue,185142,medium,issue.body,"goal as this refactoring initiative and provide crucial underlying technical context, serving as perfect complements to this plan: #174469 #185142 Looking forward to great discussions and contributions in the Slack channel and PR reviews! cc @mruberry",https://github.com/pytorch/pytorch/issues/185590,2350bad93e32743279db554428000b78da1d106829fff866b7fa41bdfcd6e18a references,issue,183913,issue,157807,medium,issue.body,25 regex-parses the same Python dict to extract expected cudnn versions. Prerequisites This RFC depends on the scikit-build-core migration (#157807) landing first. The current setuptools-based build already supports dynamic dependencies: setup.py:main() composes install_requir...,https://github.com/pytorch/pytorch/issues/183913,f81002f8ff44092b4c291ccfeda0dc73f2d7e01844afd46a59f008aaa6ffb483 references,issue,183913,pr,185208,medium,issue.body,"and the lock's variant fork derive from the same group, so they can't silently diverge. Proof of concept A working PoC is up as a draft PR (#185208), stacked on the migration: e1fb24c — the rt-* groups, [tool.uv] conflicts, and the group-reading provider; uv lock produces a si...",https://github.com/pytorch/pytorch/issues/183913,c4ea7522d91f10fa08868d0bd8ae37d5765ba2de63ae4ae5d0c47a6e233f4c45 references,issue,183913,issue,160875,medium,issue.comments[0].body,"Related on better supporting uv flow (on the consumer side): #160875 vllm-project/vllm#24218 - currently the index files don't play very well with uv, so having variants specified in square brackets (but with",https://github.com/pytorch/pytorch/issues/183913,a7c4bc7812cfcabb13c06584bab48ede0d868ae61fd415f3b395309a12e4405f references,issue,185135,issue,141184,medium,issue.body,"00+ lowered buffers would likely hit similar costs Related #158007 — torch.compile cannot compile LSTM (blocked by allow_rnn=False default) #141184 — (separate dynamo bug, 0 graphs compiled) #163207 Environment PyTorch: 2.13.0a0+git795be92 CUDA: 13.0 GPU: NVIDIA H200 OS: RHEL...",https://github.com/pytorch/pytorch/issues/185135,5eff16f51fab970db35fe8cb7602d4aa8c9fb8def20a68f6d41a8421a0f023fd references,issue,185135,issue,158007,medium,issue.body,"ion slow for common RNN configurations. This is a one-time compilation cost (cached runs are fast), but it's high enough that users hitting #158007 who try allow_rnn=True as suggested by the error message will likely give up thinking it's hung. Measured data All measurements f...",https://github.com/pytorch/pytorch/issues/185135,833c590f7ec2fe7372f0cf19aa8afd15978023817788e2539a3b68f12b77e797 references,issue,185200,pr,185231,medium,issue.comments[1].body,"apper), so before every access to the data-ptr it needs to be unwrapped. Strangely e.g. a ...to(...).contiguous() wraps it again. I drafted #185231 which handles wrapped tensors What I'm still missing is a way to set the precision of the output as I'm debugging an accuracy issue",https://github.com/pytorch/pytorch/issues/185200,30911eb9201947394ee5c59ac25a33d0a52c192dabe9852a3e44bad82d470bef references,issue,183758,issue,181946,medium,issue.body,"contiguity Parametrize the test_noncontiguous_samples to run across all the expanded test cases above Additional context Motivated by Issue #181946 that modifies handling of non-contiguous tensors for a number of MPS tensor ops, and may introduce bugs that will not be caught w...",https://github.com/pytorch/pytorch/issues/183758,a39fa0097a2b47f06c799f86bbd6483f2bd90fe3505c6f22424446784a56d0c3 references,issue,183060,issue,114392,medium,issue.comments[0].body,This is a historical issue that was reported: #114392 But I want to point out that both JAX PyTree and PyTorch Python PyTree have their limitation: JAX PyTree does not preserve the key order of,https://github.com/pytorch/pytorch/issues/183060,92d2d8a9570ee76f0f1a82dc8782f7683d53af68e0849d0c299848689c3c86b3 references,issue,184996,issue,162700,medium,issue.body,"ve(). This makes the exported artifact state inconsistent and harder to diagnose. Related issue check The closest related report I found is #162700, which also mentions tensors with non-zero elements but unallocated/zero storage. I did not find an existing minimal report for t...",https://github.com/pytorch/pytorch/issues/184996,155fe275a5dc8aca3b24dfd10c76493a873a70e0a9f7012d849f75ac21c1ffc5 references,issue,170636,issue,151549,medium,issue.comments[1].body,Duplicated with #151549,https://github.com/pytorch/pytorch/issues/170636,148207fe2ab8e24fe0933e24e66adf126d271da9504ba5a8debfb1b8b978084b references,issue,177815,issue,158793,medium,issue.body,"bmit a PR. This benefits any torch.compile backend that lowers collectives to static IR through the more efficient use of caching. Related: #158793 (DeviceMesh iterations RFC), #159017 (PG bookkeeping) Alternatives DeviceMesh.get_all_replica_groups(mesh_dim) public method — co...",https://github.com/pytorch/pytorch/issues/177815,09e5c903f03f0a26fca7db2d26b3456afa9ae52c612cb8e5b6ea6318eb91fc9d references,issue,177815,issue,159017,medium,issue.body,"mpile backend that lowers collectives to static IR through the more efficient use of caching. Related: #158793 (DeviceMesh iterations RFC), #159017 (PG bookkeeping) Alternatives DeviceMesh.get_all_replica_groups(mesh_dim) public method — computes the right answer from _layout...",https://github.com/pytorch/pytorch/issues/177815,20b87cd4df468d707c343b0d9c045da6581bf27d7d5c22881dea4d5f62dac685 references,issue,87961,issue,85329,medium,issue.body,"ta, './test_package.zip') tensor([[-1.3200, -1.0457, -0.0773], [ 0.6770, 1.8295, -0.8031]]) abort (core dumped) This is probably related to #85329, except that in the code example above, no Error is thrown and it directly crashes. Versions pytorch 1.12.1",https://github.com/pytorch/pytorch/issues/87961,737b615612acd5d3fee63a3eeac4c989762ff33c9418c8f0dc8047fe80a5db1c references,issue,176877,issue,174469,medium,issue.body,"es Decorator bypass extending DistTestCases.backend_feature so PrivateUse1 devices count as accelerators for GPU-count checks. Related RFC: #174469 5. Documentation A distributed.md page added to docs/source/accelerator/, covering: c10d backend, backend registration, transport...",https://github.com/pytorch/pytorch/issues/176877,ea614921ea951f161d6b5d503a8efde9c93ab31266ab974cb99575449ac31bcf references,issue,174960,issue,158657,medium,issue.body,"i OffloadManager: Batch-level offload/onload with ping-pong GPU buffers, pre-allocated pinned memory, and separate D2H/H2D streams. Related #158657 — proposes extending torch.utils.checkpoint with offloading (different design direction; couples offloading with recomputation) c...",https://github.com/pytorch/pytorch/issues/174960,ed27cfc4891e7698b905595ce61a766b8a7742631d34f853241e51c2441a03a1 references,issue,161113,issue,161111,medium,issue.body,"🐛 Describe the bug Following #161111, I tried _dynamo.disable on the causal_conv1d triton kernel, and ran into this runtime failure with selective_scan instead: INFO 08-20 16:5",https://github.com/pytorch/pytorch/issues/161113,24629bd55bb60024ce99aa5d02f5e611304343968137fce396eb8fb4a0725efb references,issue,179368,pr,179372,medium,issue.comments[0].body,"Hi @ad8e, thanks for the detailed reproducer! I’ve investigated this and submitted a fix in PR #179372. The root cause is that Inductor’s layout planner optimizes the tensors saved for backward into Channels Last (NHWC) format to improve Conv",https://github.com/pytorch/pytorch/issues/179368,bc88c33f42006da02d2cd45e939c133fbbba41b47d63da9760ef5777a7cbc4b4 references,issue,168042,issue,126024,medium,issue.comments[0].body,"ile is not thread safe overall (neither dynamo, nor inductor). For example, I guess this issue is tracking for dynamo: #118260. I also see: #126024 and #118387. AFAIK, this is a large project not currently on the roadmap.",https://github.com/pytorch/pytorch/issues/168042,0779ff56445f6e14cc1cc3ebf8ea7c93886c1880a8393ce27d62c94e618711bc references,issue,144362,issue,144247,medium,issue.body,"ompile Mode: Yields regular outputs (I guess implicit data type casting happens under torch.compile) Some related issues: #144314, #144310, #144247. Although this dtype-check-missing issue may not be severe, in case you are interested, I cherrypick a few operators where dtype...",https://github.com/pytorch/pytorch/issues/144362,63bbe261a26e3962cf089d81a3a8e38c58aadd12a8c38b3c713e540c54bbbee5 references,issue,180156,pr,181895,medium,issue.comments[1].body,"and adds an explicit float32 → double dispatch branch, following the same approach the Intel XPU team used in intel/torch-xpu-ops#3337. PR: #181895",https://github.com/pytorch/pytorch/issues/180156,4e3b755080cf0f5f7312e4971156b7d19ae5fc0185bb51266c6edb44b7e08f59 references,issue,182403,issue,167568,medium,issue.body,"Trying to unsafely apply AC to a non-functional graph with the default partitioner. Falling back to min-cut partitioner. This is related to #167568, but this repro does not require a Triton mutation. It uses an ordinary ATen indexed write. Minimal repro from collections import...",https://github.com/pytorch/pytorch/issues/182403,60af9e2ff2527ad1d0e2fa101c96db317ef96fab379f7f6a65b20a8e64ab97ce references,issue,184084,issue,173382,medium,issue.body,"ction is documented in docs/source/notes/cuda.rst but only as a private API Related issues about incomplete memory release: #17157, #46602, #173382, #20837 cc @ptrblck @msaroufim @eqy @jerryzh168 @tinglvv @nWEIdia @csarofeen",https://github.com/pytorch/pytorch/issues/184084,2e253f7e9bd8abcca378afc432671a9d2fa4e5dfe8787cd6f20ae486e85b8366 references,issue,37410,issue,32021,medium,issue.comments[0].body,(Related reset concept request but for lr: #32021),https://github.com/pytorch/pytorch/issues/37410,e1220c256dab5b9d5d9d496f8d78d6ee7fdd533350a8d464476d4e1ee230e97f references,issue,183795,pr,183487,medium,issue.body,"As part of the XNNPACK submodule update in #183487, I'm adding a temporary preprocessor define XNNPACK_NO_CODE_CACHE to gate the feature in OSS. Once fully landed, we can remove this flag an",https://github.com/pytorch/pytorch/issues/183795,b78bb6825394779881729b333ee9b2380b584849d0fca84a46d4b58612a42e3c references,issue,146941,issue,147115,medium,issue.comments[0].body,another example: #147115,https://github.com/pytorch/pytorch/issues/146941,a17c9d957ba08aec4e7a51417d29a6c39f54b5c4d206c09c19062c910362b977 references,issue,146111,issue,146108,medium,issue.comments[0].body,@ydwu4 assigned to you because of similar #146108 Feel free to unassign.,https://github.com/pytorch/pytorch/issues/146111,cab82591bd436ed40c53d91b79a5c405641f32c802d0a6a196efa8cadbde19c2 references,issue,183621,issue,182817,medium,issue.body,🐛 Describe the bug Context Related to the discussion at issue #182817 this is an orthogonal issue with the current testing approach and a suggestion on how we could address it: When testing op correctness for,https://github.com/pytorch/pytorch/issues/183621,c371df6d57f00efcc8941e545bcbdd4ec1cf6e05741b844029e61cc4e7b81137 references,issue,9867,issue,128959,medium,issue.comments[1].body,"removed from this repository. Evidence: The ""Remove Caffe2"" tracking issue (#72536) was closed in Oct 2024, and the umbrella removal issue (#128959) tracks the cleanup work. The caffe2/python/, caffe2/operators/, and caffe2/proto/ directories no longer exist in main — only caf...",https://github.com/pytorch/pytorch/issues/9867,c93131ed0337524cdbc877a9b78f491eaeb503171a972b7d0b1ce65205571f99 references,issue,11869,issue,128959,medium,issue.comments[0].body,"removed from this repository. Evidence: The ""Remove Caffe2"" tracking issue (#72536) was closed in Oct 2024, and the umbrella removal issue (#128959) tracks the cleanup work. The caffe2/python/, caffe2/operators/, and caffe2/proto/ directories no longer exist in main — only caf...",https://github.com/pytorch/pytorch/issues/11869,1d0a16497b4e50828a28fcc1f951b1b4ab5e35d5f708059c453092c3457dd06c references,issue,10818,issue,128959,medium,issue.comments[0].body,"removed from this repository. Evidence: The ""Remove Caffe2"" tracking issue (#72536) was closed in Oct 2024, and the umbrella removal issue (#128959) tracks the cleanup work. The caffe2/python/, caffe2/operators/, and caffe2/proto/ directories no longer exist in main — only caf...",https://github.com/pytorch/pytorch/issues/10818,6c9ee9b6644a8c16094061bacd747008cc3458be893519d446835be171d338ca references,issue,10464,issue,128959,medium,issue.comments[0].body,"removed from this repository. Evidence: The ""Remove Caffe2"" tracking issue (#72536) was closed in Oct 2024, and the umbrella removal issue (#128959) tracks the cleanup work. The caffe2/python/, caffe2/operators/, and caffe2/proto/ directories no longer exist in main — only caf...",https://github.com/pytorch/pytorch/issues/10464,705b4551910d9d6199589e876adc459686114fc122dcef2e56de5c2bc05459c4 references,issue,10440,issue,128959,medium,issue.comments[0].body,"removed from this repository. Evidence: The ""Remove Caffe2"" tracking issue (#72536) was closed in Oct 2024, and the umbrella removal issue (#128959) tracks the cleanup work. The caffe2/python/, caffe2/operators/, and caffe2/proto/ directories no longer exist in main — only caf...",https://github.com/pytorch/pytorch/issues/10440,7b0141d217fc5ab4e1ff369a70369c2071a8e8d4030fad8e96c72e32f63e2344 references,issue,9707,issue,128959,medium,issue.comments[0].body,"removed from this repository. Evidence: The ""Remove Caffe2"" tracking issue (#72536) was closed in Oct 2024, and the umbrella removal issue (#128959) tracks the cleanup work. The caffe2/python/, caffe2/operators/, and caffe2/proto/ directories no longer exist in main — only caf...",https://github.com/pytorch/pytorch/issues/9707,d65a55694e22912286b12334dcf125c6aa47cd7048181f8486c8bfaaa0c15031 references,issue,9349,issue,128959,medium,issue.comments[0].body,"removed from this repository. Evidence: The ""Remove Caffe2"" tracking issue (#72536) was closed in Oct 2024, and the umbrella removal issue (#128959) tracks the cleanup work. The caffe2/python/, caffe2/operators/, and caffe2/proto/ directories no longer exist in main — only caf...",https://github.com/pytorch/pytorch/issues/9349,a09002eeb6f5703a85df12799255113206f5d1a00710cd075b4503b210501d6a references,issue,8595,issue,128959,medium,issue.comments[0].body,"removed from this repository. Evidence: The ""Remove Caffe2"" tracking issue (#72536) was closed in Oct 2024, and the umbrella removal issue (#128959) tracks the cleanup work. The caffe2/python/, caffe2/operators/, and caffe2/proto/ directories no longer exist in main — only caf...",https://github.com/pytorch/pytorch/issues/8595,4b337bf6fb20c0795945bebbf675deae7f4e1f9c3986f15d780d9c6a730d5eca references,issue,8533,issue,128959,medium,issue.comments[0].body,"removed from this repository. Evidence: The ""Remove Caffe2"" tracking issue (#72536) was closed in Oct 2024, and the umbrella removal issue (#128959) tracks the cleanup work. The caffe2/python/, caffe2/operators/, and caffe2/proto/ directories no longer exist in main — only caf...",https://github.com/pytorch/pytorch/issues/8533,fdc99d3328fa381a2a3bcd32a860c351e3920ba56cafbc5fda5e8621b601fb38 references,issue,8442,issue,128959,medium,issue.comments[0].body,"removed from this repository. Evidence: The ""Remove Caffe2"" tracking issue (#72536) was closed in Oct 2024, and the umbrella removal issue (#128959) tracks the cleanup work. The caffe2/python/, caffe2/operators/, and caffe2/proto/ directories no longer exist in main — only caf...",https://github.com/pytorch/pytorch/issues/8442,cea1bbefad3e1b822578bd16efd6038fa5157b20dac6419496d378432848ff36 references,issue,7835,issue,128959,medium,issue.comments[0].body,"removed from this repository. Evidence: The ""Remove Caffe2"" tracking issue (#72536) was closed in Oct 2024, and the umbrella removal issue (#128959) tracks the cleanup work. The caffe2/python/, caffe2/operators/, and caffe2/proto/ directories no longer exist in main — only caf...",https://github.com/pytorch/pytorch/issues/7835,bc53a987fd21ac41d776bc8de01acd8e8ed98fadef5902a0217a5b45c405dfac references,issue,7614,issue,128959,medium,issue.comments[0].body,"removed from this repository. Evidence: The ""Remove Caffe2"" tracking issue (#72536) was closed in Oct 2024, and the umbrella removal issue (#128959) tracks the cleanup work. The caffe2/python/, caffe2/operators/, and caffe2/proto/ directories no longer exist in main — only caf...",https://github.com/pytorch/pytorch/issues/7614,f6c4b81a2dced0735eba2b84802ce891e4d03021f1292ac88d5fff22badd5b5a references,issue,7569,issue,128959,medium,issue.comments[0].body,"removed from this repository. Evidence: The ""Remove Caffe2"" tracking issue (#72536) was closed in Oct 2024, and the umbrella removal issue (#128959) tracks the cleanup work. The caffe2/python/, caffe2/operators/, and caffe2/proto/ directories no longer exist in main — only caf...",https://github.com/pytorch/pytorch/issues/7569,45ad3013299d59508cc5e392cd9660b35e82e8f9a3df3f7025a9952c85b7a311 references,issue,7491,issue,128959,medium,issue.comments[0].body,"removed from this repository. Evidence: The ""Remove Caffe2"" tracking issue (#72536) was closed in Oct 2024, and the umbrella removal issue (#128959) tracks the cleanup work. The caffe2/python/, caffe2/operators/, and caffe2/proto/ directories no longer exist in main — only caf...",https://github.com/pytorch/pytorch/issues/7491,6f12cbf084090cdc306de0eadddad10db6402a771003334ef2549ad41a3a90a4 references,issue,7490,issue,128959,medium,issue.comments[0].body,"removed from this repository. Evidence: The ""Remove Caffe2"" tracking issue (#72536) was closed in Oct 2024, and the umbrella removal issue (#128959) tracks the cleanup work. The caffe2/python/, caffe2/operators/, and caffe2/proto/ directories no longer exist in main — only caf...",https://github.com/pytorch/pytorch/issues/7490,1d1e2ae3a3d8a2368005598aa5df8ca80808d5c8e21656c160a975499e092b89 references,issue,7374,issue,128959,medium,issue.comments[0].body,"removed from this repository. Evidence: The ""Remove Caffe2"" tracking issue (#72536) was closed in Oct 2024, and the umbrella removal issue (#128959) tracks the cleanup work. The caffe2/python/, caffe2/operators/, and caffe2/proto/ directories no longer exist in main — only caf...",https://github.com/pytorch/pytorch/issues/7374,a95a0af309a315f4c8c00f495ea104c41263899ed07acbdde933af773ef36d6c references,issue,7362,issue,128959,medium,issue.comments[0].body,"removed from this repository. Evidence: The ""Remove Caffe2"" tracking issue (#72536) was closed in Oct 2024, and the umbrella removal issue (#128959) tracks the cleanup work. The caffe2/python/, caffe2/operators/, and caffe2/proto/ directories no longer exist in main — only caf...",https://github.com/pytorch/pytorch/issues/7362,1ab5e45a7930732348da80942cc1df53b2ca96648402e202a78a863304363b21 references,issue,7060,issue,128959,medium,issue.comments[0].body,"removed from this repository. Evidence: The ""Remove Caffe2"" tracking issue (#72536) was closed in Oct 2024, and the umbrella removal issue (#128959) tracks the cleanup work. The caffe2/python/, caffe2/operators/, and caffe2/proto/ directories no longer exist in main — only caf...",https://github.com/pytorch/pytorch/issues/7060,96e8816bae8a9c6e093543a0afcd6c1144b904d15169c9eb636480535f3dd9be references,issue,6979,issue,128959,medium,issue.comments[0].body,"removed from this repository. Evidence: The ""Remove Caffe2"" tracking issue (#72536) was closed in Oct 2024, and the umbrella removal issue (#128959) tracks the cleanup work. The caffe2/python/, caffe2/operators/, and caffe2/proto/ directories no longer exist in main — only caf...",https://github.com/pytorch/pytorch/issues/6979,db184340c799c99377bfb3ad65f2a0ae000070b00523bddbf0adf50b43587282 references,issue,6785,issue,128959,medium,issue.comments[0].body,"removed from this repository. Evidence: The ""Remove Caffe2"" tracking issue (#72536) was closed in Oct 2024, and the umbrella removal issue (#128959) tracks the cleanup work. The caffe2/python/, caffe2/operators/, and caffe2/proto/ directories no longer exist in main — only caf...",https://github.com/pytorch/pytorch/issues/6785,625725c5c4d77109a39bcd348bda56c3759a15e8d66aac9dd8adb3918ef5c91c references,issue,6549,issue,128959,medium,issue.comments[0].body,"removed from this repository. Evidence: The ""Remove Caffe2"" tracking issue (#72536) was closed in Oct 2024, and the umbrella removal issue (#128959) tracks the cleanup work. The caffe2/python/, caffe2/operators/, and caffe2/proto/ directories no longer exist in main — only caf...",https://github.com/pytorch/pytorch/issues/6549,2ff8c265d12e164c12d1d00535d3b93d3fec5e88f9511811b286e50e3398315c references,issue,181946,pr,182712,medium,issue.comments[0].body,181949 3.1x HistogramKernel PR #181951 1.08x Im2Col PR #182709 1.25x Indexing N/A 1.28x Regresses contiguous performance to 0.77x Linear PR #182712 1.77x LossOps PR #182714 1.44x Pooling PR #182715 1.11x RangeFactories N/A 1.00x Does not improve contiguous or non-contiguous pe...,https://github.com/pytorch/pytorch/issues/181946,54186da4281f404547f5e2d619a6b0459f712498d9229b0f786633fecdc9d587 references,issue,178492,issue,154052,medium,issue.body,"ehavior) - OS: macOS 26.1 / 26.2 - Hardware: MacBook Pro M1 Pro (32 GB) and Mac mini M4 Pro (64 GB) — same issue on both - Python: 3.12 cc: #154052 (context from the ""most requested MPS ops"" discussion)",https://github.com/pytorch/pytorch/issues/178492,d651a41b192944d0a2e1f1a39f7312174a564cd72321895ccc8a2f547e7f8cab references,issue,182382,pr,182377,medium,issue.body,"a one-sided algorithm (or any new algorithm we wish to register) will perform better. I have separated this out into a set of discrete PRs: #182377 - Update DTensor to store local tensor shards using SymmetricMemory, if enabled. #182378 - Add a Python-level get primitive to Sy...",https://github.com/pytorch/pytorch/issues/182382,2ae9c34aee2d0f6a26b9624236093233967fd5191867a93d3351ceff8c47930a references,issue,182382,pr,182379,medium,issue.body,"82378 - Add a Python-level get primitive to SymmetricMemory, useful for implementing one-sided communication algorithms. (Orthogonal to 1.) #182379 - Update DTensor's matrix multiply dispatch logic to support additional types of algorithms. This is a draft and where I would li...",https://github.com/pytorch/pytorch/issues/182382,42ac083074b05efa9c00b422ff59d63d299a34ca5963a92613cb6eeb5ed8badc references,issue,182661,issue,182641,medium,issue.body,"oducer for setting up a DDP model and jvp'ing through it. I also cannot try the same via FSDP2's replicate mode, because that's broken too (#182641). the consequence of these two errors is that there is no distributed way to train MeanFlow models. which is a shame because the...",https://github.com/pytorch/pytorch/issues/182661,a8b55386c9cc1bf079ad61e64eb5ab0fe7729634860fbbec723931182837b334 references,issue,182661,issue,180284,medium,issue.comments[0].body,"ils, please see https://pytorch.org/docs/main/notes/extending.func.html so _DDPSink doesn't use the setup_context idiom. this is similar to #180284.",https://github.com/pytorch/pytorch/issues/182661,c399b1a79437f7476fdad03b07a4a347ae22f3a11dddfab84e458272fb2dd7e1 references,issue,182842,issue,170487,medium,issue.comments[0].body,Maybe related: #170487 #171158,https://github.com/pytorch/pytorch/issues/182842,10d97ac151d3892470cad53cbaac8c02cfedd77ac35307c72763410acd1b8429 references,issue,182842,issue,171158,medium,issue.comments[0].body,Maybe related: #170487 #171158,https://github.com/pytorch/pytorch/issues/182842,9dc129098d942e3355590f70f7ae775cfaa4c10b01a9e60ed7f1550d731d6335 references,issue,148077,issue,64359,medium,issue.comments[0].body,A bit related: #64359,https://github.com/pytorch/pytorch/issues/148077,b3e2047a1c8e6380a081055320859afde451845ae18e06b3abc05a2d420d11de references,issue,182263,issue,82098,medium,issue.comments[0].body,t to have some dynamic diagnostics on the actually blas loaded into the process (can be dynamically fetched from /proc/self/maps). Related: #82098 It would be good that PyTorch at least printed a warning about this or even threw Python exception if the loaded blas (or openmp)...,https://github.com/pytorch/pytorch/issues/182263,4384bbb60113a11bc591e0dfe0c02343a2806fb53971510b42d83274f4e4b678 references,issue,182036,issue,181891,medium,issue.body,"🐛 Describe the bug Similar to #181891 , I tested torch.while_loop() and confirmed it has a similar limitation: if the body_fn returns nothing, it is not supported and leads to f",https://github.com/pytorch/pytorch/issues/182036,d3e3e23ad9f5e19eb8ac8c3aaeb59cfdb408b83be6eff59c1f32279a84c6b483 references,issue,145555,issue,123800,medium,issue.comments[0].body,"oh, I have just discovered I already filed it #123800 but it got us nowhere - perhaps this can revive the old ticket?",https://github.com/pytorch/pytorch/issues/145555,1d275b53a28a14ad5b1f6ae90d0c7bce01919e7ddb329e7fd172b2a677190e7f references,issue,164469,issue,157807,medium,issue.body,"setup_helpers and friends as we complete the transition to a more modern build system; in this latter sense, this issue is complementary to #157807. Projects such as Numpy, Scipy, Scikit-image, and others have adopted Spin, benefiting from the common, shared commands and addin...",https://github.com/pytorch/pytorch/issues/164469,4a140bbf0b88c4b99c2374e6bece92b4f6dd9d8ffa2b974ddd39c23dda474206 supersedes,issue,164469,issue,169479,medium,issue.body,"owed: ARM (#172965), Darwin (#175185), and clearer failure messaging (#173694). Ongoing community ideation for additional commands lives in #169479. Tracking Build build develop will be implemented as a straight-up replacement for setup.py develop, possibly improved after the...",https://github.com/pytorch/pytorch/issues/164469,840101224f372eb45183bb71c0baeb9c104013ccada4a4c38a353e05df30932f references,issue,177258,issue,177254,medium,issue.body,ing from the base repo dir: python test/jit/test_freezing.py TestFreezing.test_module_with_shared_type_instances Possibly the same issue as #177254 Affects Neoverse N1 and V1 Versions Commit - 08b6f48 cc @EikanWang @jgong5 @wenzhe-nrv @sanchitintel @mruberry @snadampal @milpuz...,https://github.com/pytorch/pytorch/issues/177258,f0b43a26e0fc0a03f599e8b5d3596a290eaacdbdb2fd6c8e08b89600458b9604 references,issue,181347,pr,181823,medium,issue.comments[1].body,"Fixed it, PR is here #181823",https://github.com/pytorch/pytorch/issues/181347,9a3906e6ac797957cb7d3224b0beb164c1b64408157e513c3fab78cacaca0b75 references,issue,180833,pr,180353,medium,issue.comments[0].body,sorry assigned myself wrongly. I think it's best to link your PR here: #180353,https://github.com/pytorch/pytorch/issues/180833,6fd63231494b0a72f85c57cdcbabdc6c05bf0a5324c141a137f7d30855333788 references,issue,180833,pr,180353,medium,issue.comments[1].body,Tracking status: Part 1: PR open at #180353 Part 2: folded into #180353 to avoid landing a regression between parts,https://github.com/pytorch/pytorch/issues/180833,0007f6c500f9f52b32e31c387aeb3aeed5acb2432edfc0288ad4d5ab09df6b03 references,issue,181725,issue,100347,medium,issue.body,"in the wrapper / functional layer, not a difference in the math. The same comparison on CPU shows ~1× (no gap). This is not a duplicate of #100347. That issue is about need_weights=True falling off the SDPA path entirely. This issue is about the residual ~9× gap that remains o...",https://github.com/pytorch/pytorch/issues/181725,a8f4e5d07e3b7221a88daac6bc1d94a143c2b6e34272a31b28ad2466dc44cf25 references,issue,179008,issue,179010,medium,issue.comments[0].body,"Hi, I'm working on #179010 and plan to address this issue in the same tutorial update. The new section will: Document torch._C._acc.register_python_privateuseone_hook",https://github.com/pytorch/pytorch/issues/179008,c0026a88e99c02c0371b2b86ecd6b10aacdcba57ecf310cc6c1144e63c98447c references,issue,179010,issue,179008,medium,issue.comments[0].body,ch.library op registration patterns Complete end-to-end NumPy-backed example with autograd Hooks/device guard customization (also addresses #179008) C++ vs Python comparison table and limitations I'll also cross-reference torch_openreg and the test file at test/test_privateuse...,https://github.com/pytorch/pytorch/issues/179010,2c1af482faedcc2f970999ae50b11854aae1652e91c515ca3b91c4e64cc21ec3 references,issue,154318,issue,106704,medium,issue.body,ass Time: 4.772 seconds I have to reduce to batch size since DataLoader fails to complete due to unexpected bus error Versions 2.7 Relevant #106704 cc @msaroufim @jerryzh168 @andrewkho @divyanshk @ssnl @VitalyFedyunin @dzhulgakov,https://github.com/pytorch/pytorch/issues/154318,9145e0512c5cb0b7d48a0ea129b86d9095a896560c18ec18ab5e48e6c2d3b67b references,issue,180377,issue,154356,medium,issue.body,") or torch could throw an error here too. If anyone has any thoughts on this, please let me know and I can add a this check. (related issue #154356) Versions PyTorch version: 2.12.0a0+gitd7d0482 Is debug build: True CUDA used to build PyTorch: 12.8 ROCM used to build PyTorch:...",https://github.com/pytorch/pytorch/issues/180377,e9d8c2065b442866f67f13d104bfafe1aeae4249becdda09b57eec3e5882fe76 references,issue,179294,issue,175873,medium,issue.body,"ice-agnostic impl. So a dedicated MPS backward pass needs to be implemented. Problem - Unused second output tensor As some issues (#176730, #175873) have mentioned, _scaled_dot_product_attention_math_for_mps returns a second output tensor that is currently not used by anything...",https://github.com/pytorch/pytorch/issues/179294,cf8d018b821685289079e9d15cf74b5fedc1e4a22e6f82ec6f36a168ee55de2f references,issue,179294,issue,176730,medium,issue.body,"ing a device-agnostic impl. So a dedicated MPS backward pass needs to be implemented. Problem - Unused second output tensor As some issues (#176730, #175873) have mentioned, _scaled_dot_product_attention_math_for_mps returns a second output tensor that is currently not used by...",https://github.com/pytorch/pytorch/issues/179294,076e569670e3e679fa33cdc516064371eccb83dd292bdfee3214e87d98fc352b references,issue,166209,issue,148819,medium,issue.body,"🚀 The feature, motivation and pitch #2 from here #148819 (comment) We chose to be explicit about accepting only matrices for Muon. We'd like to standardize the ways people smoosh particular types",https://github.com/pytorch/pytorch/issues/166209,c9e48148a6add08936a1c289c62e42a94271e25cbbbeff5565fba7cda281c429 references,issue,180975,issue,171659,medium,issue.body,"Background Following the prior RFC (#171659), we have concluded multiple rounds of discussions and initial experiments. This effort has been jointly driven by XuanTie Team(from Alibab",https://github.com/pytorch/pytorch/issues/180975,c477860ed4c22b8beb06ca4329affeedcf69ed94b6d1538c69a191836f9d8830 references,issue,158917,pr,165837,medium,issue.comments[0].body,80 #165897 Approved but waiting to be merged: None Needs review(Dependency from top to bottom): #165728 #165631 #166115 #166395 Developing: #165837 #166288 #166128,https://github.com/pytorch/pytorch/issues/158917,6b8522160f553c18103d0d8400b92dd59e9c94692a2840d1257297bdab674c3e references,issue,141896,issue,112583,medium,issue.comments[0].body,"The root cause should be that autocast cache does not key on requires grad. #112583 What happens is that when checkpoint recomputes, we are in a fresh autocast context that did not first run the same module under no-grad, e",https://github.com/pytorch/pytorch/issues/141896,93ea614fd407fda32e07918aeb2a8c047a3439993db0a53364825ed83dd12a1b references,issue,119613,issue,119574,medium,issue.comments[0].body,Is this similar to #119574,https://github.com/pytorch/pytorch/issues/119613,b59bbbe237d315f1bc0fdf6877c7a1896a71af6c7032e0bcd370cf8125d616ba references,issue,180762,issue,44279,medium,issue.body,"allelism / device linking support for custom PyTorch CUDA extensions #78225 added relocatable device code linking support for CUDAExtension #44279 shows user confusion around how to achieve this in cpp extensions, especially for JIT-like workflows Expected outcome: JIT cpp ext...",https://github.com/pytorch/pytorch/issues/180762,352a2fbc582c575aea71d7f606bbc9120b90f1aca44bb42e38ed0431657db392 references,issue,180712,issue,180343,medium,issue.body,"🚀 The feature, motivation and pitch This probably lands somewhere between #180343 and #180345, not sure if it is justified in having its own issue, but in TorchTPU we have set up autoloading of our tpu module, but we only",https://github.com/pytorch/pytorch/issues/180712,ef5868b1ee30af084844d0fab54beeeb1451ce7326f13bf8eefba8e7b8427f9d references,issue,180712,issue,180345,medium,issue.body,"🚀 The feature, motivation and pitch This probably lands somewhere between #180343 and #180345, not sure if it is justified in having its own issue, but in TorchTPU we have set up autoloading of our tpu module, but we only load it if",https://github.com/pytorch/pytorch/issues/180712,987a2fd953cc2d0dce1c2f6d9722794de30f2c05edfb4eee9ed5a7bee53b4f7d references,issue,157807,pr,180247,medium,issue.body,CMAKE_BUILD_TYPE=Release in the Windows CI setup script. Phase 2: The changeover Migrate build system from setuptools to scikit-build-core (#180247) — switch build-backend in pyproject.toml from setuptools.build_meta to scikit_build_core.build. Add scikit-build-core configurat...,https://github.com/pytorch/pytorch/issues/157807,73fc09ba2d57fd6a94c4da18e8fbbd0897513f2478801058123e860f3e21b6c9 references,issue,157807,pr,180248,medium,issue.body,"setuptools infrastructure and address remaining rough edges. Included in the initial landing: Remove setup.py and setuptools build helpers (#180248) — delete setup.py, tools/setup_helpers/, tools/build_pytorch_libs.py, and related test files. Retain a minimal tools/build_libto...",https://github.com/pytorch/pytorch/issues/157807,8069808fb27709c65017fd388379a98c93c22b63c3b6d9589ad16350859f3170 references,issue,157807,pr,180250,medium,issue.body,cikit-build-core's redirect-mode editable installs can locate cmake-built artifacts alongside source-tree Python files. Add setup_.py shim (#180250) — provide a compatibility shim (named with underscore to avoid build tool detection) that translates legacy setup.py invocations...,https://github.com/pytorch/pytorch/issues/157807,294c80d23ccb62a5db26488867c164c16e3c74a07e3347aa8098e413b4e32683 references,issue,180624,issue,157807,medium,issue.body,"Summary With PyTorch's own build backend migration to scikit-build-core (see #157807), this RFC proposes to refactor torch.utils.cpp_extension so that PyTorch no longer carries a runtime dependency on setuptools, and to clar",https://github.com/pytorch/pytorch/issues/180624,d04925e879c3bf8efb18f5dfa0815c1102c94034ee76e760b86c94b671c871bc references,issue,180596,issue,180595,medium,issue.body,"compiled 3 times using a lightweight counting backend: Baseline: no flags (default config) +capture_scalar_outputs: (tested separately, see #180595) +capture_dynamic_output_shape_ops: torch._dynamo.config.capture_dynamic_output_shape_ops = True We measured subgraph count per r...",https://github.com/pytorch/pytorch/issues/180596,7b0410abacd74db2dfafef4fe913e1c506e417e4256ad28be9e3b942e0d5a196 references,issue,148402,issue,122840,medium,issue.body,There is a triton bug that such reduction will have un-coalesced memory access due to the non-const mask for a potential vectorized load. ( #122840 ) The load is not vectorized and can be less efficient Here is an optimization to fix it. We can split the reduction loop to 2 lo...,https://github.com/pytorch/pytorch/issues/148402,08750e4d18dd818656c1c00f327d6158e464354184ec2a67696df8cd440e6223 references,issue,180362,pr,173519,medium,issue.body,nly Intended to validate upstream CI integration and detect regressions Current Implementation: Initial upstream enablement implemented in: #173519 Purpose: Validate Power CI behavior on real upstream PRs Collect reliability and signal‑quality data before introducing scheduled...,https://github.com/pytorch/pytorch/issues/180362,89a57d0b97c11fc3292778b9471f04f87fcad3de1be7da23ebdcf74c21450082 references,issue,169277,issue,143112,medium,issue.body,"rator.get_device_capability for users to get some device capabilities (currently focus on data types) to solve some issues like #165038 and #143112. The supported data types in PyTorch are: float32, float64, float16, bfloat16, float8_e5m2, float8_e4m3fn, float8_e5m2fnuz, float...",https://github.com/pytorch/pytorch/issues/169277,f27c723d6895cb7318157a0a1782c607ec1596014b6951412f9104cf4e21e0e5 references,issue,169277,issue,165038,medium,issue.body,"torch.accelerator.get_device_capability for users to get some device capabilities (currently focus on data types) to solve some issues like #165038 and #143112. The supported data types in PyTorch are: float32, float64, float16, bfloat16, float8_e5m2, float8_e4m3fn, float8_e5m...",https://github.com/pytorch/pytorch/issues/169277,8ba947aadf35fb64edd9fdf7aa53690ebcdd371999508c2880f9d947535ba53a references,issue,179510,issue,172026,medium,issue.body,"piled_autograd = True which yields a different error message: NotImplementedError: Cannot access storage of TensorWrapper Possibly related: #172026 Full MWE with eager / compile import torch from torch import Tensor, nn @torch.no_grad() def _fixpoint_iteration(fn, x: Tensor) -...",https://github.com/pytorch/pytorch/issues/179510,65d571c8ef08d0a9f03dc52e0327c8a7e7d46f8520f75ead4b1ef1ded059e59b references,issue,167631,issue,167645,medium,issue.comments[0].body,"As another important application, I want to highlight #167645, that is export + parametrizations. Caching parametrizations is one of those additional state-modifying functions I was talking about, whic",https://github.com/pytorch/pytorch/issues/167631,3c876ef77ed9176e957a390bec9335ae96f9e74f0d5d83ef3d9e05052e8a29ba references,issue,127078,issue,173881,medium,issue.comments[0].body,"I run POWER8 S824 hardware and was looking to submit the fix, but the code is gone. Happy to help with any other ppc64le build issues — see #173881.)",https://github.com/pytorch/pytorch/issues/127078,d8e90f989cc69cd47adb3c66888ec880216d72b2ac33eb502a8867b98b171929 references,issue,178379,pr,178298,medium,issue.body,"🐛 Describe the bug As per title. Will be addressed in #178298. Here is the repro: In [1]: import torch In [2]: A=torch.tensor([[ 1.0000, 0.0000, 0.0000, 0.0000, 0.0000], ...: ...: [-0.1264, 1.0000, 0.0",https://github.com/pytorch/pytorch/issues/178379,245e597d31189b1839ed0b83d018874498f448eebb0355e18a05300c2302c6fe references,issue,178379,pr,178298,medium,issue.comments[1].body,Will be addressed in #178298,https://github.com/pytorch/pytorch/issues/178379,017e969905a2163ee65428fcd7752147876913446f5e96c06d25270e8eda7378 references,issue,179128,issue,158698,medium,issue.comments[0].body,"Could this be related? #158698 Maybe not related to HSDP, but this also increases peak memory (as reduction is first done in fp32 buffer, and then a separate bf16 buffer",https://github.com/pytorch/pytorch/issues/179128,67cf1b73c31f15d046e8cba08cd4d7848e5dafc6c99bc60111e525e82f0b681a references,issue,174913,issue,168756,medium,issue.body,Context Source file: test/test_linalg.py Test class: TestLinalg Test name: test_tensorinv Defined at line: 3771 Parent issue: #168756 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_linalg.py cc @sunway513 @jithunnair-amd @pruthvistony @ROCmSupport @,https://github.com/pytorch/pytorch/issues/174913,fa173df9a63bfba31e2fc80bb99dc0b8090507250f7ec225f8d22bd90fceb850 references,issue,179250,pr,179273,medium,issue.comments[1].body,"I've tried a design in #179273 , I would be glad to see your thoughts about it.",https://github.com/pytorch/pytorch/issues/179250,7b53b302cd2322746f5337f4ee7c42f8ef0cf959581afc621212c412d06484c0 references,issue,171158,issue,170487,medium,issue.body,hose under @torch.no_grad context manager Currently only torch.func.grad allows to fullgraph-compile computing grads wrt inputs because of: #170487 so it's an important usecase torch.autograd.grad is fine with some leaf nodes being inplace updated if this happens under @torch....,https://github.com/pytorch/pytorch/issues/171158,5fbc381d13330507090816180d418c60aa1233a79ecd056367553ab3e07def08 references,issue,158710,issue,69431,medium,issue.body,"itemsize, a.nbytes) b = to_(a, torch.bfloat16) print(b, b.dtype, b.dtype.itemsize, b.nbytes) Or might be better to model this as torch.to_: #69431",https://github.com/pytorch/pytorch/issues/158710,42d9e4c880c9f894095093714623927f888af483bb37bb15dc9764e58c23593b references,issue,158710,issue,158698,medium,issue.body,"🚀 The feature, motivation and pitch Motivated by: #158698 It's sometimes useful in eager non-differentiable code to be able to downcast inplace (float32->bfloat16 for reduce_dtype->param_dtype, flo",https://github.com/pytorch/pytorch/issues/158710,2423479054c87b270b9b66f6bb9e5023ee973d1b813bf21720172c38e7aa1ee8 references,issue,175695,issue,174496,medium,issue.body,I first noticed this here #174496 (comment) I am making a separate issue becuase I plan to solve #174496 in a stacked PR on top of a PR to solve this specific issue. Essenti,https://github.com/pytorch/pytorch/issues/175695,5bceb54bf81e9bbae2f853ab890c4e17be802ee09950cc36d7f8d2e2b21bf305 references,issue,154052,issue,77764,medium,issue.body,"This issue will list the most requested ops for the MPS backend, that haven't been implemented yet, taken from the comments to #77764 and #141287. The script that produced this data is provided in this gist. The number of votes is computed as the number of unique users req",https://github.com/pytorch/pytorch/issues/154052,8c6a1e88bb7567aa756f7a4b7021c58507dd2da20c008b017ab3f3d8963aca83 references,issue,154052,issue,141287,medium,issue.body,"This issue will list the most requested ops for the MPS backend, that haven't been implemented yet, taken from the comments to #77764 and #141287. The script that produced this data is provided in this gist. The number of votes is computed as the number of unique users request...",https://github.com/pytorch/pytorch/issues/154052,c26f87b7192c13a1885bf42b35326c63d689bfb37ad4bded528014608f828b40 references,issue,44027,issue,30458,medium,issue.body,"Previously: #30458 An immutable tensor is a tensor which cannot be mutated, e.g., via inplace operations or out= operations. Due to the reduced API surface of",https://github.com/pytorch/pytorch/issues/44027,39e6bb1392943aee7ba069b2a315ef4ab26b15b5ff262af1608c5273120f4098 references,issue,166291,issue,177427,medium,issue.comments[1].body,"@kwen2501 has discussed supporting non-contig allgathers in his RFC (#177427), and @weifengpy has been working on FlexShard which would be a possible place to experiment with this (pytorch/torchtitan#2378)",https://github.com/pytorch/pytorch/issues/166291,563fbba98d4c32e65a04b03d48939e0c9af7d44907d028d7fca44109fa5e1a04 references,issue,178251,issue,126654,medium,issue.comments[1].body,ytorch.org/t/function-scaled-dot-product-efficient-attention-backward0-returned-nan-values-in-its-0th-output/191752 #119320 #125674 #138649 #126654 Demidov-N/LiquidSearcher@96dcd8c ORNL/MATEY#13 I haven't found a suitable solution under these issues (and the Math path is also...,https://github.com/pytorch/pytorch/issues/178251,16ac961145ed0542efd85d6d2a28403b9a1e44cb99efc644712e408bdf39bd1c references,issue,178251,issue,138649,medium,issue.comments[1].body,iscuss.pytorch.org/t/function-scaled-dot-product-efficient-attention-backward0-returned-nan-values-in-its-0th-output/191752 #119320 #125674 #138649 #126654 Demidov-N/LiquidSearcher@96dcd8c ORNL/MATEY#13 I haven't found a suitable solution under these issues (and the Math path...,https://github.com/pytorch/pytorch/issues/178251,935e9b4075ee8cf93dc4f20b864a2f9299ed113a4b8f9eec19aa7dde11987810 references,issue,177969,issue,175211,medium,issue.comments[0].body,"I agreed, and a similar idea from #175211",https://github.com/pytorch/pytorch/issues/177969,b8b82041d3e4a3c02fcf6e61e05c1b8eb78a44e879741b3c7b7e64fba20ec454 references,issue,152954,issue,152963,medium,issue.comments[0].body,This is a child issue of #152963 cc @IvanKobzarev,https://github.com/pytorch/pytorch/issues/152954,0cb122849cdf9efa977f30f02a0f69384da17ed8de6b2bec917895653df3c2e5 references,issue,155720,issue,65462,medium,issue.body,"allocator > > > >&) [clone .localalias.823] () from /root/miniforge3/lib/python3.12/site-packages/torch/lib/libamdhip64.so #65462 0x00007fffa62dd518 in hip::Graph::GetRunList(std::vector >, std::allo...",https://github.com/pytorch/pytorch/issues/155720,0b55e438b13f83e77eba8d5604ff6f0f59a5101572cd85d0d83230339ed46090 references,issue,155720,issue,65464,medium,issue.body,"aphInstantiate(hip::GraphExec**, hip::Graph*, unsigned long) () from /root/miniforge3/lib/python3.12/site-packages/torch/lib/libamdhip64.so #65464 0x00007fffa632a20a in hip::hipGraphInstantiateWithFlags(hipGraphExec**, ihipGraph*, unsigned long long) () from /root/miniforge3/l...",https://github.com/pytorch/pytorch/issues/155720,7c8b47ae6e28959f278a328198ff4530353f97a2c7f98cb5e713113f282eaa39 references,issue,155720,issue,65465,medium,issue.body,"ateWithFlags(hipGraphExec**, ihipGraph*, unsigned long long) () from /root/miniforge3/lib/python3.12/site-packages/torch/lib/libamdhip64.so #65465 0x00007fffd71b5fe7 in at::cuda::CUDAGraph::capture_end() () from /root/miniforge3/lib/python3.12/site-packages/torch/lib/libtorch_...",https://github.com/pytorch/pytorch/issues/155720,b8bb0334af32a8eaf4ae19bd164d643f947d44c140c982a03a672ac24cb69c4b references,issue,177833,pr,168223,medium,issue.body,_size=0) opt.step() sched.step() # ZeroDivisionError: integer modulo by zero Negative step_size values are also silently accepted. Related: #168223 #169107 #174025 (all stale) cc @vincentqb @jbschlosser @albanD @janeyx99 @crcrpar @malfet,https://github.com/pytorch/pytorch/issues/177833,d1d179910fa3822eda9ec8373d6a45b26c490cc3e2f2f517a1859cc3ec5aab2e references,issue,177833,pr,169107,medium,issue.body,opt.step() sched.step() # ZeroDivisionError: integer modulo by zero Negative step_size values are also silently accepted. Related: #168223 #169107 #174025 (all stale) cc @vincentqb @jbschlosser @albanD @janeyx99 @crcrpar @malfet,https://github.com/pytorch/pytorch/issues/177833,99942bdc1f21269cdfa06c763a86e3f6f81297a77b83915724d39b8d141048c6 references,issue,177833,pr,174025,medium,issue.body,p() sched.step() # ZeroDivisionError: integer modulo by zero Negative step_size values are also silently accepted. Related: #168223 #169107 #174025 (all stale) cc @vincentqb @jbschlosser @albanD @janeyx99 @crcrpar @malfet,https://github.com/pytorch/pytorch/issues/177833,0fe668a5c0321c19c78ed57257acba59b2a000e781cccf3f650eb1543603fbf9 references,issue,164429,issue,164057,medium,issue.comments[1].body,"Hi @AlstonTang , thanks for the report! This should be the same issue as #164057 (comment) Could you try with the pytorch 2.9 and latest driver?",https://github.com/pytorch/pytorch/issues/164429,cdc57f12c082c185bb9e3166a6cb20526418533aa6fd912d22a722f88eac7465 references,issue,164966,issue,161381,medium,issue.comments[1].body,"oblems regarding your use case: There is a known issue that the SYSMAN/L0 driver provides inconsistent free memory result. The issue is in: #161381 . For internal track, see JIRA: GSD-11670 The provided reproducer trys to allocate big memory at once. The Driver team said the U...",https://github.com/pytorch/pytorch/issues/164966,909f13f02a4e25a57819a35cbd3df0e80b643908d08ce0dde060799ed16c18a8 references,issue,76324,issue,78065,medium,issue.body,ssel and related functions as PyTorch operators. Enjoy! One of a five-part series of special functions issues: Gamma and Related Functions (#78065) Bessel and Related Functions (#76324) Orthogonal Polynomials (#80152) Elliptic Functions and Integrals (#80157) API Bessel Functi...,https://github.com/pytorch/pytorch/issues/76324,b90710f54f20ac9efb713ebcd82d95b616cf9b1725db185c689ff36fe3f6bb3d references,issue,76324,issue,80152,medium,issue.body,"part series of special functions issues: Gamma and Related Functions (#78065) Bessel and Related Functions (#76324) Orthogonal Polynomials (#80152) Elliptic Functions and Integrals (#80157) API Bessel Functions bessel_j(input: Tensor, n: int = 0) → Tensor Bessel function of th...",https://github.com/pytorch/pytorch/issues/76324,07802612b9942eaa97939c2ac3f98b2ca5a3d96d4c84c5c01e838231ac44ad66 references,issue,76324,issue,80157,medium,issue.body,"amma and Related Functions (#78065) Bessel and Related Functions (#76324) Orthogonal Polynomials (#80152) Elliptic Functions and Integrals (#80157) API Bessel Functions bessel_j(input: Tensor, n: int = 0) → Tensor Bessel function of the first kind, $J_{n}\left(\text{input}\rig...",https://github.com/pytorch/pytorch/issues/76324,e0893af3d11c1bcd15f472eabf422edd1611a7361f8b29c3ab047b34f898d0fb references,issue,58743,issue,58734,medium,issue.body,"ility issues with Python Array API specification: __array_namespace__ (add at/towards the end, it's the attribute that declares compliance) #58734 #58736 #58739 #58740 #58741 #58742 #55090 #58745 torch.nonzero diverges from the specification (it returns a tensor, not a tuple,...",https://github.com/pytorch/pytorch/issues/58743,ca1155816401eba330f2827c12ad6eba36f3c9d2840aec6e04d6f43ff9c5d1dd references,issue,58743,issue,58736,medium,issue.body,"ssues with Python Array API specification: __array_namespace__ (add at/towards the end, it's the attribute that declares compliance) #58734 #58736 #58739 #58740 #58741 #58742 #55090 #58745 torch.nonzero diverges from the specification (it returns a tensor, not a tuple, by defa...",https://github.com/pytorch/pytorch/issues/58743,ff1c2d90230841f81eea917c924886b8b17f2b211595ce98c2fe84693b3d1974 references,issue,58743,issue,58741,medium,issue.body,"ay API specification: __array_namespace__ (add at/towards the end, it's the attribute that declares compliance) #58734 #58736 #58739 #58740 #58741 #58742 #55090 #58745 torch.nonzero diverges from the specification (it returns a tensor, not a tuple, by default) - see gh-64502 #...",https://github.com/pytorch/pytorch/issues/58743,e8351730d07ba7352011b88a3894f0726d27f5f32a647f6de2501d05b47af4c9 references,issue,58743,issue,58742,medium,issue.body,"specification: __array_namespace__ (add at/towards the end, it's the attribute that declares compliance) #58734 #58736 #58739 #58740 #58741 #58742 #55090 #58745 torch.nonzero diverges from the specification (it returns a tensor, not a tuple, by default) - see gh-64502 #59786 #...",https://github.com/pytorch/pytorch/issues/58743,e240dc913642936621be835793c5488daaadbd3bb1e4b9559db5065e822bdac5 references,issue,58743,issue,58745,medium,issue.body,"__array_namespace__ (add at/towards the end, it's the attribute that declares compliance) #58734 #58736 #58739 #58740 #58741 #58742 #55090 #58745 torch.nonzero diverges from the specification (it returns a tensor, not a tuple, by default) - see gh-64502 #59786 #59787 #59867 #5...",https://github.com/pytorch/pytorch/issues/58743,72e77af9d3e6fb51a7103e8d912431d744e2214812a93467bd4fdd63205dfd6c references,issue,58743,issue,59786,medium,issue.body,"740 #58741 #58742 #55090 #58745 torch.nonzero diverges from the specification (it returns a tensor, not a tuple, by default) - see gh-64502 #59786 #59787 #59867 #59868 #70591 #70906 #70909 #70910 #70914 #70915 #70916 #70918 #70919 #70920 #70921 #70922 #9190 #70924 #70925 cc @e...",https://github.com/pytorch/pytorch/issues/58743,548d8c754ff961b92366417b50d0e6d8c47ca51cdeceeb5f9db763c87985eed4 references,issue,58743,issue,59787,medium,issue.body,"741 #58742 #55090 #58745 torch.nonzero diverges from the specification (it returns a tensor, not a tuple, by default) - see gh-64502 #59786 #59787 #59867 #59868 #70591 #70906 #70909 #70910 #70914 #70915 #70916 #70918 #70919 #70920 #70921 #70922 #9190 #70924 #70925 cc @ezyang @...",https://github.com/pytorch/pytorch/issues/58743,2c2537fdcdf92c566f381a00edf4d5c35a407994f73efdb154474d3c2d568b49 references,issue,58743,issue,59868,medium,issue.body,"090 #58745 torch.nonzero diverges from the specification (it returns a tensor, not a tuple, by default) - see gh-64502 #59786 #59787 #59867 #59868 #70591 #70906 #70909 #70910 #70914 #70915 #70916 #70918 #70919 #70920 #70921 #70922 #9190 #70924 #70925 cc @ezyang @gchanan @zou35...",https://github.com/pytorch/pytorch/issues/58743,1d8176f9cc0a8c9f0f9b9cd08044035809f4810f396c6bd1d7c1d262d0d24331 references,issue,58743,issue,70906,medium,issue.body,"ch.nonzero diverges from the specification (it returns a tensor, not a tuple, by default) - see gh-64502 #59786 #59787 #59867 #59868 #70591 #70906 #70909 #70910 #70914 #70915 #70916 #70918 #70919 #70920 #70921 #70922 #9190 #70924 #70925 cc @ezyang @gchanan @zou3519 @bdhirsh @j...",https://github.com/pytorch/pytorch/issues/58743,228bd2b1c62d9338ce3c14aa97bd546f4858f5b0771d0ac281ccbbfcdfaaaafd references,issue,58743,issue,70910,medium,issue.body,"erges from the specification (it returns a tensor, not a tuple, by default) - see gh-64502 #59786 #59787 #59867 #59868 #70591 #70906 #70909 #70910 #70914 #70915 #70916 #70918 #70919 #70920 #70921 #70922 #9190 #70924 #70925 cc @ezyang @gchanan @zou3519 @bdhirsh @jbschlosser @an...",https://github.com/pytorch/pytorch/issues/58743,28d825236e93b382005c9559d116afee71fabf218b8aaec1b29c0b4fdd4cbb63 references,issue,58743,issue,70914,medium,issue.body,"rom the specification (it returns a tensor, not a tuple, by default) - see gh-64502 #59786 #59787 #59867 #59868 #70591 #70906 #70909 #70910 #70914 #70915 #70916 #70918 #70919 #70920 #70921 #70922 #9190 #70924 #70925 cc @ezyang @gchanan @zou3519 @bdhirsh @jbschlosser @anjali411",https://github.com/pytorch/pytorch/issues/58743,cdbbf65b31f9978262ea2f5755505fe9d286500c02829ff8f394a52ef701586d references,issue,58743,issue,70915,medium,issue.body,"specification (it returns a tensor, not a tuple, by default) - see gh-64502 #59786 #59787 #59867 #59868 #70591 #70906 #70909 #70910 #70914 #70915 #70916 #70918 #70919 #70920 #70921 #70922 #9190 #70924 #70925 cc @ezyang @gchanan @zou3519 @bdhirsh @jbschlosser @anjali411",https://github.com/pytorch/pytorch/issues/58743,c6c11c6ad5133d0f3dd8080e466e42ffa214ab67c0b4a7d02bc5ae14e10289a4 references,issue,58743,issue,70916,medium,issue.body,"ication (it returns a tensor, not a tuple, by default) - see gh-64502 #59786 #59787 #59867 #59868 #70591 #70906 #70909 #70910 #70914 #70915 #70916 #70918 #70919 #70920 #70921 #70922 #9190 #70924 #70925 cc @ezyang @gchanan @zou3519 @bdhirsh @jbschlosser @anjali411",https://github.com/pytorch/pytorch/issues/58743,8767f2ba4193ef96cda61510bdcb6b7879ce177eb6ea34f841de3370908e63b7 references,issue,58743,issue,70919,medium,issue.body,"turns a tensor, not a tuple, by default) - see gh-64502 #59786 #59787 #59867 #59868 #70591 #70906 #70909 #70910 #70914 #70915 #70916 #70918 #70919 #70920 #70921 #70922 #9190 #70924 #70925 cc @ezyang @gchanan @zou3519 @bdhirsh @jbschlosser @anjali411",https://github.com/pytorch/pytorch/issues/58743,d96f988c080199a0ec18529714b5e23d8640b3e68097646ae16ee36127ccdcce references,issue,58743,issue,70920,medium,issue.body,"tensor, not a tuple, by default) - see gh-64502 #59786 #59787 #59867 #59868 #70591 #70906 #70909 #70910 #70914 #70915 #70916 #70918 #70919 #70920 #70921 #70922 #9190 #70924 #70925 cc @ezyang @gchanan @zou3519 @bdhirsh @jbschlosser @anjali411",https://github.com/pytorch/pytorch/issues/58743,6797b9b458448c19f92435de628e05b11e34a9fc9a5224448ed41426e98d7823 references,issue,58743,issue,70921,medium,issue.body,", not a tuple, by default) - see gh-64502 #59786 #59787 #59867 #59868 #70591 #70906 #70909 #70910 #70914 #70915 #70916 #70918 #70919 #70920 #70921 #70922 #9190 #70924 #70925 cc @ezyang @gchanan @zou3519 @bdhirsh @jbschlosser @anjali411",https://github.com/pytorch/pytorch/issues/58743,a44e96f598a42d19417a2fade3bfaf69d10f8dcafaed11a13515c70ff6a61a1a references,issue,58743,issue,70924,medium,issue.body,fault) - see gh-64502 #59786 #59787 #59867 #59868 #70591 #70906 #70909 #70910 #70914 #70915 #70916 #70918 #70919 #70920 #70921 #70922 #9190 #70924 #70925 cc @ezyang @gchanan @zou3519 @bdhirsh @jbschlosser @anjali411,https://github.com/pytorch/pytorch/issues/58743,f893da338dfc3dea9f5a67bb4e7ec1bc6e7f1eeb092d39967d236a45d603bd97 references,issue,58743,issue,70925,medium,issue.body,- see gh-64502 #59786 #59787 #59867 #59868 #70591 #70906 #70909 #70910 #70914 #70915 #70916 #70918 #70919 #70920 #70921 #70922 #9190 #70924 #70925 cc @ezyang @gchanan @zou3519 @bdhirsh @jbschlosser @anjali411,https://github.com/pytorch/pytorch/issues/58743,c9ec884328317ba5bbc00805f19aa1919a4072e9605724bd5b40923f353eda99 references,issue,174481,issue,125245,medium,issue.comments[0].body,This is partially related to #125245. Are there any updates with respect to that issue?,https://github.com/pytorch/pytorch/issues/174481,541cbaf56eb28cdbed39523a70a2066c3e1b67e745e3251de8ea8d42edeaf753 references,issue,125678,issue,60252,medium,issue.comments[0].body,".) cpu = _conversion_method_template(device=torch.device(""cpu"")) Apart from #115638 that already mentioned by OP, I also found an abandoned #60252 as well. I'm not installing Numpy since I don't use it. But the amount of warning from PyTorch is concerning. If PyTorch without N...",https://github.com/pytorch/pytorch/issues/125678,9eb71ecbdc5cdeabdf9c39c3611544893c0a44793b0693bda9baacf323989efb references,issue,175557,issue,154824,medium,issue.body,"_weakrefs() in cudagraph_trees.py — it frees storages that are still needed as inputs to subsequent warmup nodes. This is the same issue as #154824. Crashes at step 1 with mode=""reduce-overhead"", reproduces on PyTorch 2.8 and 2.10. I also found a secondary issue on the same co...",https://github.com/pytorch/pytorch/issues/175557,a24fd8b40ef7bcb7204c9f5231f730f8ed9ea97078fe3eb2e507f40877525b82 references,issue,168820,issue,168819,medium,issue.body,est/test_scaled_matmul_cuda.py Test class: TestFP8Matmul Test name: test_scaled_mm_vs_emulated_row_wise Defined at line: 1267 Parent issue: #168819 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_scaled_matmul_cuda.py cc @sunway513 @jithunnair-amd @pruth...,https://github.com/pytorch/pytorch/issues/168820,365644794ea33392fec4811930c081796d146540d918083870a55a5fc4533e90 references,issue,176592,issue,139500,medium,issue.body,Function | custom op 1 threads: ----------------------------------- f(x) | 25.0 | 88.3 Times are in microseconds (us). Potential related to #139500 (but #139500 seems tailored more towards inference setting). cc: @riccardofelluga Versions main cc @jerryzh168 @chauhang @penguin...,https://github.com/pytorch/pytorch/issues/176592,f6b26425d697e9291a2108357679d049bbfd3be3aee72cc82b433b02387d0fea references,issue,176592,issue,139500,medium,issue.comments[1].body,"known issue, duplicate of #139500",https://github.com/pytorch/pytorch/issues/176592,a578d2e75c2852327b4e567887404bf2a93b0e5ca4ba1ce7d345e58e5ac0fac8 references,issue,168823,issue,168819,medium,issue.body,st_scaled_matmul_cuda.py Test class: TestFP8Matmul Test name: test_blockwise_mxfp8_nvfp4_mxfp4_numerics Defined at line: 1894 Parent issue: #168819 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_scaled_matmul_cuda.py cc @sunway513 @jithunnair-amd @pruth...,https://github.com/pytorch/pytorch/issues/168823,525700b12b4a0b489b8f17a85dcdf7ba1cd0e42bbb6510b74cf9fa5f5a5e21a5 references,issue,155044,issue,146719,medium,issue.body,"al export program has an addition constant lifted_tensor_0 that appears to be equal to the value of torch.sqrt(x). This error is related to #146719, where I added a comment with some speculation about this issue. I decided to open a new issue here because the old one was quite...",https://github.com/pytorch/pytorch/issues/155044,81c6d759acc1a2ebfe0402ce5690a94b07679cd70ba7c4f1917ab504f2149e1f references,issue,174608,issue,170127,medium,issue.body,"tor.exc.InductorError: AcceleratorError: CUDA error: an illegal memory access was encountered Related Issue, which advised using math SDPA: #170127 Expected behavior AOTInductor should compile and package the exported model successfully. Actual behavior AOTInductor crashes dur...",https://github.com/pytorch/pytorch/issues/174608,0c8695242982f27036a366ae5e71a2c950b4fdfa384545977a5753a6768e7d37 references,issue,176730,issue,175873,medium,issue.body,"ttn_mask, dropout_p, is_causal, std::nullopt, /*dropout_mask*/ scale, enable_gqa)); } where we immediately drop the other outputs. Based on #175873 this does however have a cost at least for the MPS backend. The second output tensor is the attention weight tensor that can be l...",https://github.com/pytorch/pytorch/issues/176730,b05e9fd717c5ae297278c0dd00eb1d5d89ce9867892b76dcaf4c90ec721d345d references,issue,117122,issue,10454,medium,issue.body,"eeze(-1) SpinvUhBB = Spinv * UhBB return Vh.adjoint() @ SpinvUhBB X_svd= svd_lstsq(A, B) print(""X_svd"",X_svd) Related issues: #88101 #85021 #10454 cc @ptrblck @jianyuh @nikitaved @pearu @mruberry @walterddr @xwang233 @lezcano",https://github.com/pytorch/pytorch/issues/117122,8d8281f11a85bcdeabbee693c53218a5270fcbd184fe3d3161d23e62a08ff1c2 references,issue,172987,pr,176317,medium,issue.comments[1].body,Part 1 of this change landed in PR #175818 Part 2 up for review in PR #176317,https://github.com/pytorch/pytorch/issues/172987,2418e53f4ddb5bd70ab4551b2710e06f2444cf730125bcb27dc07dab90811cd0 references,issue,158371,issue,58997,medium,issue.body,"nd call init_group or something similar). Essentially, the state are the weights / running_stats if optimizer is considered as a nn.Module: #58997 Alternatives No response Additional context No response",https://github.com/pytorch/pytorch/issues/158371,3bd01f56cfa79d1cba488768e24c2f6a8469b0eda2ae7a45941bd4ef0e448757 references,issue,158371,issue,104849,medium,issue.comments[1].body,"Alternative option might be if some sort of sufficiently generic update rule-like swiss-army knife eager fused op would be introduced like #104849 , supporting passing multiple numerators and denomenators... I think if they are good, there would need to be some exposure - mayb...",https://github.com/pytorch/pytorch/issues/158371,6d681cf7d25f252eaafd414703c893bd042c9784fff4971b3a0b8b2eba5af2d7 references,issue,175978,pr,176070,medium,issue.comments[1].body,"Hi @mikaylagawarecki — tagging you since you marked the issue as actionable. I opened PR #176070, but CI hasn’t started because the fork workflows are Awaiting Approval. Could you please approve the workflow runs so CI can proceed? Than",https://github.com/pytorch/pytorch/issues/175978,37543ac925bf43b38baa92e0a33f0494635a669871205133bcda95a753bbca21 references,issue,142836,issue,142515,medium,issue.body,"oduce support for convolutions with out_channels > 2**16. While this appeared to work for Conv1d, it introduced a regression in Conv2d (see #142515 (comment)). It remains unclear whether Conv3d is affected. The issue results in a silent correctness bug. A guard was previously...",https://github.com/pytorch/pytorch/issues/142836,41ede47b46e4185100fb299c534cbacb069a1d7c024a5c6be33fdeacd1a8cefa references,issue,72388,issue,170426,medium,issue.comments[0].body,"torch.stack(indices, dim=1) # A.shape = torch.Size([2, 2, 2, 2, 2, 2]) global_topk(A, 1) # (tensor(7), tensor([1, 0, 0, 1,0 ,1])) Related: #170426",https://github.com/pytorch/pytorch/issues/72388,bc3f3faa56c08286ed21a52857bb9b9fd22e8bef55a108e1942362b20619ec2f references,issue,110080,issue,108046,medium,issue.body,"because, at the moment, each device type has its own API. I would be happy to implement such an API if there is support from the community. #108046 is loossly related and could be put under the same API namespace (cc team: @AnthonyBarbier @hmellor) Alternatives No response Add...",https://github.com/pytorch/pytorch/issues/110080,0826bfc9894f51acc1a690cdad084999a7c5d85e301f2f947e2ff86a11412341 references,issue,110080,issue,108046,medium,issue.comments[1].body,('xla') d.is_built() d.is_available() etc It might be also useful to have a generic interface to work with any device type (similar to what #108046 proposed): torch.devices.is_available() # if any device type is available devices = torch.devices.available() # returns ['X'] (al...,https://github.com/pytorch/pytorch/issues/110080,0aaa35109158d449153e851ebda402af259dd22922be029d555bb07257a63e55 references,issue,62530,issue,50688,medium,issue.body,search/NeuralCompression/blob/e1702923c97cc6c00cd8af4db2eb59f587c09e44/neuralcompression/entropy_coders/jax_arithemetic_coder.py ? Related: #50688 pytorch/ao#292 ? cc @albanD @mruberry @jbschlosser,https://github.com/pytorch/pytorch/issues/62530,d8c3a446b430bc26be27b7ee47c129c15306573fe48bf5aad7b90935ba4f7f84 references,issue,62530,issue,50688,medium,issue.comments[1].body,"ers/jax_arithemetic_coder.py is a good case study for these functional utils. I'll keep this issue specialized on lax.cond / lax.switch and #50688 on scans/for loops Wrt this arithmetic coder, is torch.cond sufficiently reimplementing lax.cond? Then maybe torch.switch (to reim...",https://github.com/pytorch/pytorch/issues/62530,8f41dea56cae40ceae791f8472b773ad5d0716b023ed6130c8560384581fb659 references,issue,139518,issue,139500,medium,issue.body,This came up when I was investigating #139500 (and in parallel @ezyang hypothesized about boxing vs unboxing performance). Experiment: calling torch.stack on 5 tensors. We can vary the,https://github.com/pytorch/pytorch/issues/139518,10d545100ae53b8e0cd4e7394fd2bebb114d4ed8d53a854db786543299dc2aa8 references,issue,115626,issue,111441,medium,issue.body,compiles for previously encountered frameNums and using them whenever it encounters them again From my limited understanding of the issues #111441 and #105279 (comment) torch.compile() unrolls the for loop. My issue here is that the dataloader serves different input sizes (tem...,https://github.com/pytorch/pytorch/issues/115626,37246add482be2038574143a3efb6496eb563a994ac94f275bed41407ed82b71 references,issue,115626,issue,50688,medium,issue.comments[0].body,Maybe related? #50688,https://github.com/pytorch/pytorch/issues/115626,ea51ed04fdf121861fe63ce1525ca48888977297be93c129222f7a85a9143c60 references,issue,122540,issue,125958,medium,issue.comments[0].body,Just as a reference in the case we have some fresh news to comment about point 2 #125958,https://github.com/pytorch/pytorch/issues/122540,a62978c7f9ca3250c556df1de651e28f3974c5da6a19cc298cadb48336f6fbbf references,issue,164299,issue,145374,medium,issue.body,"Describe the bug The MPS backend leaks memory. This has been a long-standing issue as mentioned in Issue #155060, Issue #154329, and issue #145374. I've observed it on my machines for the past 2 years and decided to do some RCA work over the past week. Memory growth is not rec...",https://github.com/pytorch/pytorch/issues/164299,7e25dd890e96a228ee16aa7788282be88d3f75dc5847818d6f9d40a568482c10 references,issue,164299,issue,154329,medium,issue.body,"Describe the bug 🐛 Describe the bug The MPS backend leaks memory. This has been a long-standing issue as mentioned in Issue #155060, Issue #154329, and issue #145374. I've observed it on my machines for the past 2 years and decided to do some RCA work over the past week. Memor...",https://github.com/pytorch/pytorch/issues/164299,880960c6c3f037346939893aa7a47c3b92c610aacb9cc80ab32145b6355c52a2 references,issue,164299,issue,155060,medium,issue.body,"🐛 Describe the bug 🐛 Describe the bug The MPS backend leaks memory. This has been a long-standing issue as mentioned in Issue #155060, Issue #154329, and issue #145374. I've observed it on my machines for the past 2 years and decided to do some RCA work over the past week.",https://github.com/pytorch/pytorch/issues/164299,0bc50de95d56f1a5fb3ee026b46a984a4779091c169e605be7b526dc9f050980 references,issue,174541,pr,170092,medium,issue.comments[0].body,It's possible that our work on LazyConstantVariable might resolve this? #170092 we're still in the process of merging,https://github.com/pytorch/pytorch/issues/174541,442372614b55f2962e548f3d485f39c67b40ccf995a526a391107f94b5f439d9 references,issue,174541,pr,170092,medium,issue.comments[1].body,"@williamwen42 Thanks for pointing this out... I checked #170092 and its stacked #170644. They improve LazyConstantVariable behavior and lazy-key dict assignment, but they don’t change the ConstDictVariab",https://github.com/pytorch/pytorch/issues/174541,258580cd6c6c67cc655a605f1080d13bad61ababe6d8dbe4666c7db398a87962 references,issue,174541,pr,170644,medium,issue.comments[1].body,"@williamwen42 Thanks for pointing this out... I checked #170092 and its stacked #170644. They improve LazyConstantVariable behavior and lazy-key dict assignment, but they don’t change the ConstDictVariable read path (getitem_co",https://github.com/pytorch/pytorch/issues/174541,4166a2052d0ce281d40498c9004c830ac9ecc9a60edfc91b9de62aa3bfc6f79e references,issue,101082,issue,67761,medium,issue.comments[0].body,A bit related: #72146 #68332 #67761,https://github.com/pytorch/pytorch/issues/101082,68799d3799dd66fa226aed29de32318d5e5957b4b559cca691a89b7693c4ad62 references,issue,101082,issue,68332,medium,issue.comments[0].body,A bit related: #72146 #68332 #67761,https://github.com/pytorch/pytorch/issues/101082,8e1919858384a00767886e19bf8ca59d27094a403d7858f7b46f30f2fe029e40 references,issue,101082,issue,72146,medium,issue.comments[0].body,A bit related: #72146 #68332 #67761,https://github.com/pytorch/pytorch/issues/101082,a08979d20f6ce6ec7c4d770f72788f56561899cb9a8815161ce8a43f9f9ae27a references,issue,113914,issue,50688,medium,issue.body,"atives We can try to work with the fx.graph, but feels like more of a hack. Additional context This is also related to the earlier request, #50688",https://github.com/pytorch/pytorch/issues/113914,7bed2e65151f2d90af4a6cb22693273f995373fe4e325577640704ba13622bff references,issue,114412,issue,114397,medium,issue.body,74 (improved dynamic shapes support) #114389 (need to proxy subclass constructor calls/methods) #117596 (need to proxy subclass attributes) #114397 (run functionalization before the subclass) #114411 (higher order op support + infra) #114413 (support subclasses that do not des...,https://github.com/pytorch/pytorch/issues/114412,2d634bd73adb5093b08a7fcf1e51bd5899c70ac4f11aa3f5872493353d7a428d references,issue,114412,issue,114403,medium,issue.body,"testing #114398 (testing for compositions of subclasses, modes, etc) #114399 (refactor ViewAndMutationMeta to not need an is_training flag) #114403 (not exactly specific to subclasses but came up during the initial work) #114414 (PyInterpreter.cpp cleanup that came up during s...",https://github.com/pytorch/pytorch/issues/114412,940b7a9e24e1cf5434c6291662e1264bc0ce80ffca722b906c3be3d49c5a6781 references,issue,114412,issue,114411,medium,issue.body,"o proxy subclass constructor calls/methods) #117596 (need to proxy subclass attributes) #114397 (run functionalization before the subclass) #114411 (higher order op support + infra) #114413 (support subclasses that do not desugar, go directly to the compiler) #114410 (backward...",https://github.com/pytorch/pytorch/issues/114412,c0c4ea4edb2cd8209f547a340eb58d5b353eb76ca9897f27980182756fe011aa references,issue,114412,issue,114413,medium,issue.body,") #117596 (need to proxy subclass attributes) #114397 (run functionalization before the subclass) #114411 (higher order op support + infra) #114413 (support subclasses that do not desugar, go directly to the compiler) #114410 (backward guards) #114975 (metadata mutation on sub...",https://github.com/pytorch/pytorch/issues/114412,8ec7cf87ae450d5c3fe787e2b3b010cad473811f1a5bec8094c87a97f4d0dc43 references,issue,114412,issue,114414,medium,issue.body,actor ViewAndMutationMeta to not need an is_training flag) #114403 (not exactly specific to subclasses but came up during the initial work) #114414 (PyInterpreter.cpp cleanup that came up during subclass support) cc @Chillee @ezyang @albanD @samdow @chauhang @penguinwu @bobren...,https://github.com/pytorch/pytorch/issues/114412,93988a2c1574d9bc20cfc4dc2c7d59c0fe70deef7b7089a02fdb40f94dc5bd92 references,issue,114412,issue,114415,medium,issue.body,to the compiler) #114410 (backward guards) #114975 (metadata mutation on subclass inputs) better docs Known bugs: #114405 (lack of guards) #114415 (inputs that alias each other below subclass desugaring) #114400 (input aliases input w subclasses) #114401 (input dupes another i...,https://github.com/pytorch/pytorch/issues/114412,6b2bb3032af28b6ab43f354faed919fa61213b3644c69dfad8cc612c8932bce5 references,issue,110175,issue,106596,medium,issue.body,"🚀 The feature, motivation and pitch Some details across: #109240 (comment) #106596 (comment) In summary: It is likely that since we do not analyze buffer aliases (read/write) which may occur in graph breaks, we require cop",https://github.com/pytorch/pytorch/issues/110175,48c0a0a343ddf2c7b688aff56931998cfd97996a3e400e273d5fcaf1219eb999 references,issue,82098,issue,78489,medium,issue.body,"🚀 The feature, motivation and pitch Sometimes it's useful for debugging linking / library loading (e.g. for related: #78489) to be able to find actually loaded shared dependency library full paths (libcudnn.so, libcublas.so and others) at runtime. In Unix systems",https://github.com/pytorch/pytorch/issues/82098,6c64500007f4df669089bae16723b8c26f9bd213a9eeb5b5d9093a0ffe600211 references,issue,61686,issue,58997,medium,issue.comments[1].body,"I agree it may be brittle, but it would enable scenarios like #58997 (comment) - integrating with existing almost-module code. Still worth technical discussion I think (even if decided against at the end) - e",https://github.com/pytorch/pytorch/issues/61686,33bc51bdcbcfb9ec9b25d150d363a8e03faba26e791a35aed769a27dfefd8391 references,issue,92141,issue,51720,medium,issue.comments[1].body,"h/blob/master/aten/src/ATen/native/BatchLinearAlgebraKernel.cpp#L261 Supporting these large matrices is possible, but requires implementing #51720.",https://github.com/pytorch/pytorch/issues/92141,289e27e4fa3a00fbef0c4168f067177bae00773628e6047b2e358ae8bbbd242f references,issue,169188,issue,169111,medium,issue.comments[0].body,Same comment: #169111 (comment),https://github.com/pytorch/pytorch/issues/169188,8e353c5f11add5b988ef7667166a1a3bf6de4c7a307fd4c5f180ab5d06d940ad references,issue,169188,issue,169111,medium,issue.comments[1].body,Please use Relax dynamo backend instead. #169111 (comment) There're two bugs in tvm codebase but should be fixed in apache/tvm#18725 and apache/tvm#18726.,https://github.com/pytorch/pytorch/issues/169188,b29c2fabf8466837f32509f7b2f9bd09dbc9f11090d23519f19813ed1983ddd4 references,issue,124275,issue,76410,medium,issue.comments[0].body,This seems to be caused by the call to torch.ops.profiler._record_function_enter_new. Might be related to #76410.,https://github.com/pytorch/pytorch/issues/124275,e61be68e67d04bb4fbc26d11fe4eba92bbc795cc9088c548396e5cd18bf47acf references,issue,79197,issue,23756,medium,issue.comments[0].body,"fer lifting the requirement of calling super's init: #61686 also, sometimes NOT supporting hooks can enable certain scenarios and simplify: #23756",https://github.com/pytorch/pytorch/issues/79197,6d2ebfd055e9800dcc23ec64b4928224138de2b84aa1ed91d0bae77337ffd6fe references,issue,79197,issue,61686,medium,issue.comments[0].body,"ents): https://mobile.twitter.com/kevin_zakka/status/1538634474107314176 although i prefer lifting the requirement of calling super's init: #61686 also, sometimes NOT supporting hooks can enable certain scenarios and simplify: #23756",https://github.com/pytorch/pytorch/issues/79197,98a6a40f1450eda296fb3a9fd318d14f949d35314f811c40db6f388bc7d576c9 references,issue,174081,issue,173546,medium,issue.body,Similar issue: #173546 CI Run: https://buildkite.com/vllm/ci/builds/49311 31.01.2026 Details [2026-01-31T04:40:19Z] ____________ test_ngram_and_suffix_correctness,https://github.com/pytorch/pytorch/issues/174081,cb95c027f211075eb3b96af7a998ff5e4aaf91a02b589cd1a3051bf520f86fb5 references,issue,173912,issue,139019,medium,issue.body,"enjoy the benefit of yielding nan's in the exact spots without having to manually mask for entries in which $a<1, b<1, a=b=1$. Issue #139019 might be related. Versions Collecting environment information... PyTorch version: 2.10.0+cpu Is debug build: False CUDA used to bu...",https://github.com/pytorch/pytorch/issues/173912,6f50342632a1eac5f8ccc7968636238ede8860a46db9da672b22379c6a317858 references,issue,163408,issue,149534,medium,issue.comments[1].body,Related: #163774 #149534 (comment),https://github.com/pytorch/pytorch/issues/163408,9b0118b69a98b9a15fa80b77f9be3d15228d8748f518276de2a4698ef8bfd2d2 references,issue,168393,issue,170834,medium,issue.comments[1].body,This should be fixed by #171305 (which addresses the same root cause as #170834). The Generated class in _library/autograd.py now uses _SingleLevelFunction which bypasses the functorch check in Function.apply().,https://github.com/pytorch/pytorch/issues/168393,d345c883fdecfbe91490002bdf72abdc6f465b463685d1bec3923e1492523642 references,issue,173514,issue,164281,medium,issue.body,"a memory pool is used : it would be good to add this to that section as well (maybe through a link). While looking for this, I noticed that #164281 is still open but I believe it can be closed. CC: @kwen2501 Suggest a potential alternative/fix No response cc @svekars @sekyonda...",https://github.com/pytorch/pytorch/issues/173514,90738ec6f434e771517c7524d36b9acaf89088434509270f799980ccc5cc49ca references,issue,173514,issue,164281,medium,issue.comments[0].body,"nsistency Would this address the concerns raised in the issue? I'd be happy to open a draft PR if this approach looks good! Also, regarding #164281 - should that be closed as part of this work, or handled separately? Thanks!",https://github.com/pytorch/pytorch/issues/173514,2dba3956831ccc029c2ea28f5f1a5794a200fa08f0f724b2eac22e3d66e3154b references,issue,112398,issue,108567,medium,issue.body,"in #115749) Non-copying constructor- construct a jagged layout NT from the (values, offsets) components directly (@jbschlosser in #121518) #108567: Slicing of arbitrary dims for NTs of both layouts Resolve rough edges of SDPA support for jagged layout NTs (internal only: link)...",https://github.com/pytorch/pytorch/issues/112398,22fbb92e311222382e6d663ed9e705527932eff53f909d91caba5bd1a2eb2212 references,issue,112398,pr,119977,medium,issue.body,n #122836) Data-dependent output shape support without graph breaks (in progress by @jbschlosser in e.g. #142063) Factory function support (#119977) Composition of NJTs with different metadata (e.g. different offsets tensors) but same conceptual ragged structure Documentation...,https://github.com/pytorch/pytorch/issues/112398,062ec2f27be9650aa88b5913c0b33ecc1a6e4e5022a2cb4eb6182683f8763c60 references,issue,173559,issue,148819,medium,issue.comments[0].body,#148819 (comment) I imagine this can be accomplished on top of a broader OptimizerDict idea. First we can build an OptimizerDict that's more abstra,https://github.com/pytorch/pytorch/issues/173559,ba1bddbababc009e2fa1c9f164b8f3ea4a927af7d47a9d58a72164bc47e30366 references,issue,121343,issue,120597,medium,issue.comments[0].body,"You're correct, torch.lcm internally calls a % b at some point, which is pretty much the same issue as #120597.",https://github.com/pytorch/pytorch/issues/121343,4cfdc9f55f77d60c70958637e6fecc58c8932c37b16188af2245dcf40f28603f references,issue,172500,issue,171804,medium,issue.body,"this may be expected behavior, not a bug, so please feel free to triage/re-label accordingly Context While debugging the issue reported in ##171804 (on demand profiling causing early CUDA init) I encountered some interesting/counter-intuitive behavior - I wanted to check wheth...",https://github.com/pytorch/pytorch/issues/172500,359c99fa322b5b3b32b7851853b53b3a1fb3a949ad3daaa4d6c9e92a072fac33 references,issue,170127,issue,124502,medium,issue.comments[0].body,"I think similar to #124502, if we change to the SDPBackend.MATH this succeeds from torch.nn.attention import SDPBackend with torch.nn.attention.sdpa_kernel([SDPBacken",https://github.com/pytorch/pytorch/issues/170127,6925d1e9e9933a517a1e2d1c884ff9777af58ccd40e91303609ff0d38c027405 references,issue,151705,issue,142315,medium,issue.body,"templates are effective, they make it difficult to fuse surrounding operations, even though Inductor supports prologue/epilogue fusion (see #142315 and #142315). Proposed Feature This proposal enables Inductor to generate performant matrix multiplication kernels directly, with...",https://github.com/pytorch/pytorch/issues/151705,a8368a0010789d44b90f1be5f8705aa9f4020ffb72bb0d4bdb38b9762ad4b867 references,issue,92250,issue,41950,medium,issue.comments[0].body,Related: #41950 #52675 #62931,https://github.com/pytorch/pytorch/issues/92250,56852af4beea014f9fd23f515cf3f22b27f317d0a416f4b8ca7ad0f34f1f2a6e references,issue,92250,issue,52675,medium,issue.comments[0].body,Related: #41950 #52675 #62931,https://github.com/pytorch/pytorch/issues/92250,d6b481085a53cf2bf5f1086dc72140354d6582f8ce3c8f086d95738bdd23c1f9 references,issue,92250,issue,62931,medium,issue.comments[0].body,Related: #41950 #52675 #62931,https://github.com/pytorch/pytorch/issues/92250,7dec02cf402a3dd39c2c0568e366b11d5e4cb5ed6746c0cf644a86b0f51352d3 references,issue,147662,issue,158657,medium,issue.comments[0].body,willing to work on this one here is one related issue #158657,https://github.com/pytorch/pytorch/issues/147662,0f7df4a03be824e9ac43feefd91d42190076ff81c4ed49eea246adf569b0a137 references,issue,76580,issue,25104,medium,issue.comments[1].body,"nd for right now if your use case is able to, dropout respects torch.manual_seed More generally, this is something that has come up before (#25104), but I'm not seeing a full response on that. From talking with @jbschlosser a couple months ago, this seems like something we wou...",https://github.com/pytorch/pytorch/issues/76580,6b8893012ea6d2ebbb131ddc935d4af7652b255ab74ac794b7a96422c1959949 references,issue,172632,pr,166876,medium,issue.body,4 https://gcc.gnu.org/bugzilla/show_bug.cgi?id=123457 Versions Base commit https://github.com/pytorch/pytorch/commits/5778f6ff + changes in #166876 cc @malfet @seemethere @snadampal @milpuz01 @aditew01 @nikhil-arm @fadara01 @nWEIdia,https://github.com/pytorch/pytorch/issues/172632,ea95376e7f6508492290040f6e5526e0678742651174315567d0ae9d655cf2a3 references,issue,172445,pr,166876,medium,issue.body,🐛 Describe the bug We have an open PR - #166876 to upgrade AArch64 CI images to ubuntu noble and to use GCC14 for building and testing pytorch. However we have run into several GCC and li,https://github.com/pytorch/pytorch/issues/172445,c0f850f0b3d1b85cde713ed7a492249b2c8a2f9ae09f5de784c68f06fd4cf26c references,issue,167881,issue,167283,medium,issue.body,"New Feature for Release Related to Pytorch compute platform quality levels proposal RFC #167283 This proposal is for changes to the PyTorch ""Getting Started"" page to better promote XPU builds and increase platform visibility. XPU PyTor",https://github.com/pytorch/pytorch/issues/167881,1843ecad7127be06f73cdbcdc98d3daca798a51f10010efc2c804253e5ad5f38 references,issue,167881,issue,166563,medium,issue.comments[1].body,Related request for 2.10.0 release to include wheel-variant installation instructions: #166563,https://github.com/pytorch/pytorch/issues/167881,4f35ca82947c8c14369be78b8cc6c54ead1be5c8a1965d87aa25ea8777987ee5 references,issue,172318,issue,158917,medium,issue.body,"sts() to OpenReg, providing a documented reference for backend authors on how to implement device-agnostic tests. Background As outlined in #158917, OpenReg aims to serve as the official reference implementation for new device integration into PyTorch. A key part of this is de...",https://github.com/pytorch/pytorch/issues/172318,483efe0c79212278e80eb1766e235dd4d48def02faf2a67d7a662d106b07f430 references,issue,117844,issue,95100,medium,issue.comments[1].body,I found this old issue (#95100) that runs into a similar problem w.r.t. missing complex32_bfloat16 support. @ezyang Would you still be open for adding such a new dtype?,https://github.com/pytorch/pytorch/issues/117844,e8e17f12ec102d8a37ed1fdfbcce23bcc22f9621373627d9f9e6b6e95cdfa000 references,issue,172034,issue,71379,medium,issue.body,"rom a commit (c373387) more than 4 years ago. I strongly suspect that the situation has changed in the meantime. There's an existing issue (#71379) in the same vein, but IMO broader in scope than what the aim is for this issue: respecting (or at least not overriding) CMAKE_CUD...",https://github.com/pytorch/pytorch/issues/172034,69cfa1fb6d90b81960ae8653faf0a083cc9db2084d49fadb7bf65d1231946344 references,issue,171302,issue,166208,medium,issue.comments[0].body,I wrote an issue for this exact thing here! #166208 If you have thoughts on what you're looking for/any feedback on design please share there.,https://github.com/pytorch/pytorch/issues/171302,bee966444a599c588b6ad744beaa0b300d9e46d533d2c7712d01066dbae981fc references,issue,168696,issue,168692,medium,issue.body,test/nn/test_convolution.py Test class: TestConvolutionNNDeviceType Test name: test_conv3d_cudnn_broken Defined at line: 4207 Parent issue: #168692 Code reference: https://github.com/pytorch/pytorch/blob/main/test/nn/test_convolution.py cc @sunway513 @jithunnair-amd @pruthvist...,https://github.com/pytorch/pytorch/issues/168696,644317af61a4a2b69251459a9f314a6669030e00aed2a447db0aeed51291372f references,issue,155992,issue,68320,medium,issue.body,hot on supporting in natively: meta-pytorch/torchsnapshot#102 meta-pytorch/torchsnapshot#114 And some older (before HF) discussions: #91965 #68320 I wonder if some HF utils on checkpointing / HF Hub blob caching structure could be upstreamed to PyTorch. E.g. for loading pretra...,https://github.com/pytorch/pytorch/issues/155992,e6dd070458b64c1dcc780747f19f2ab46d86059189a56f931bdd3078808ba8cf references,issue,155992,issue,91965,medium,issue.body,chsnapshot on supporting in natively: meta-pytorch/torchsnapshot#102 meta-pytorch/torchsnapshot#114 And some older (before HF) discussions: #91965 #68320 I wonder if some HF utils on checkpointing / HF Hub blob caching structure could be upstreamed to PyTorch. E.g. for loading...,https://github.com/pytorch/pytorch/issues/155992,d7dfe1d1efeb7230fe24f3f1af5c19d4baa2b9b66d83b5c84bd12bc4f7331c59 references,issue,91965,issue,68320,medium,issue.body,"access) for faster downloads and avoiding overloading public servers, especially given problems with over-downloading in all DDP replicas: #68320 It would be nice to have some mechanisms / hooks in torch domain libraries for rewriting the URLs. (maybe in torch.hub?) It would a...",https://github.com/pytorch/pytorch/issues/91965,a533102f77f7d6552fa0f9de0b9e032967ea25df97215bbc2871fc0954607164 references,issue,154613,issue,61585,medium,issue.comments[0].body,"Related to #61585. In addition to norm, split and unique in torch/functional.py also do not support axis alias.",https://github.com/pytorch/pytorch/issues/154613,499272fbfa2aaf6c8c72e413ee315c92ed83c370ddd04f713ee506ce03191ec9 references,issue,170781,issue,170770,medium,issue.body,"cribe the bug Suppose you have some model code where you initialize a meta module, and then you to_empty to instantiate it for real. Due to #170770 you can't just wrap the meta and the to_empty inside one big FakeTensorMode region. So it is tempting to do the meta conversion o...",https://github.com/pytorch/pytorch/issues/170781,07129e7186880e0062d1e5ff3eff262e5b8240b404d788272c1acdb4c0310b3b references,issue,170770,issue,145529,medium,issue.body,weakref associated with it The weakref here is almost certainly from MetaConverter. Not sure what a good approach here is. Related #141548 #145529 cc @chauhang @penguinwu @eellison @bdhirsh @azahed98 Versions main,https://github.com/pytorch/pytorch/issues/170770,e293322480234d060b2313bb2a60c13bda74460a28a0fc123624bcf92e8efe5a references,issue,170496,issue,167729,medium,issue.body,"rtunately, torch.autograd.grad is currently not supported in Dynamo, and also setting requires_grad_(True) is also not supported in Dynamo: #167729 #170487 So having to use torch.func.grad to be able to compute model grads wrt inputs (which can potentially be without requires_...",https://github.com/pytorch/pytorch/issues/170496,75a349b62d61b2cb14208d99a16005a59b2f028a2f274c8e5b367fe90cf466bc references,issue,170496,issue,170487,medium,issue.body,"y, torch.autograd.grad is currently not supported in Dynamo, and also setting requires_grad_(True) is also not supported in Dynamo: #167729 #170487 So having to use torch.func.grad to be able to compute model grads wrt inputs (which can potentially be without requires_grad) Ve...",https://github.com/pytorch/pytorch/issues/170496,b0c00603cc5e843c4c40dee0468a153e8fa2df50fbab2a7bc2e7bbf0adacd9f3 references,issue,170496,issue,170834,medium,issue.comments[0].body,Currently this seems to be the problem: #170834,https://github.com/pytorch/pytorch/issues/170496,c0ddbede38050f0404ad4b9ff6b1da97d3e164d9d0947da0ad4ca460ccbff300 references,issue,163504,issue,160553,medium,issue.comments[0].body,"This is a good old error handling problem, and fixing it is not easy, as Metal does not any assert method I.e. it's very similar to #160553",https://github.com/pytorch/pytorch/issues/163504,fe2be1e57e488038beef8ce35c2457b25aa8487cd2bc018d959580b9b8b35fa1 references,issue,168722,issue,168720,medium,issue.body,ntext Source file: test/test_cuda.py Test class: TestCudaMallocAsync Test name: test_device_memory_used Defined at line: 4847 Parent issue: #168720 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_cuda.py cc @sunway513 @jithunnair-amd @pruthvistony @ROCmS...,https://github.com/pytorch/pytorch/issues/168722,1f8b5e1db4fb8d21a1b655ad380c82a691781f49b51fb27d2e2e299cd71329ab references,issue,50112,issue,49961,medium,issue.body,it's the latter that's supposed to set the default device. If possible I would like to ask for a clarification of what @ngimel shared here: #49961 (comment) quote: Default device is the device you are setting with torch.cuda.set_device(). It's possible to set device to 1 and t...,https://github.com/pytorch/pytorch/issues/50112,1ac6d35e01b0bd305ca66fb0fe005efc29d47db99a2f3c38e89fcaacd609078e references,issue,168774,issue,168773,medium,issue.body,Context Source file: test/test_matmul_cuda.py Test class: TestMatmulCuda Test name: test_cublas_addmm Defined at line: 174 Parent issue: #168773 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_matmul_cuda.py cc @sunway513 @jithunnair-amd @pruthvistony @R...,https://github.com/pytorch/pytorch/issues/168774,70aa656af3b6c898a19c821f6db20ffb782a85c5c51939ea39f526d07c33f26b references,issue,168807,issue,168801,medium,issue.body,rce file: test/test_nn.py Test class: TestNNDeviceType Test name: test_upsamplingNearest2d_launch_fail Defined at line: 11319 Parent issue: #168801 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_nn.py cc @sunway513 @jithunnair-amd @pruthvistony @ROCmSup...,https://github.com/pytorch/pytorch/issues/168807,cdc2cdd29d305a60f707c1d2388819d8a8efc1b26e7577ba40d304cc1f26a1a0 references,issue,170168,issue,137635,medium,issue.body,"escription torch.linspace produces different results on CPU vs CUDA when using integer dtypes. This is similar to the MPS issue reported in #137635 but affects CUDA devices. Minimal reproduction import torch # Test case 1 cpu_result = torch.linspace(4.3, -3, 50, dtype=torch.in...",https://github.com/pytorch/pytorch/issues/170168,e641391db8f0a6342912f1fe79364d9fc5ae2337178bce732365e922ddbbda92 references,issue,168450,issue,168449,medium,issue.body,rce file: test/distributed/tensor/test_tensor_ops.py Test class: DistTensorOpsTest Test name: test_index Defined at line: 574 Parent issue: #168449 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/tensor/test_tensor_ops.py cc @sunway513 @jithunnair...,https://github.com/pytorch/pytorch/issues/168450,75aeaf53fb10d594b70961253cf65650ab168a605de40df417dc8c147ff4e4c0 references,issue,168448,issue,168447,medium,issue.body,ile: test/distributed/tensor/test_matrix_ops.py Test class: DistMatrixOpsTest Test name: test_grouped_mm Defined at line: 551 Parent issue: #168447 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/tensor/test_matrix_ops.py cc @sunway513 @jithunnair...,https://github.com/pytorch/pytorch/issues/168448,2f30ab4ba5c9967cecef46a04f72cfc3b89d72f3fd132212b4d8e0588f3dd1c1 references,issue,168858,issue,168855,medium,issue.body,ing.py Test class: TestTesting Test name: test_cuda_assert_should_not_stop_common_distributed_test_suite Defined at line: 367 Parent issue: #168855 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_testing.py cc @sunway513 @jithunnair-amd @pruthvistony @RO...,https://github.com/pytorch/pytorch/issues/168858,c41269542ae6d3bea75919cebdda0c9162570a9a6f9f001bb651c9e9b1f66a24 references,issue,136586,issue,136446,medium,issue.comments[1].body,Yes: #136446. but this is an additional request for flops,https://github.com/pytorch/pytorch/issues/136586,9efb3d29cf4f03c3a8d201dd59cbbfc4223a1398bf58b148353e9f35bf28f52f references,issue,166512,issue,13130,medium,issue.body,"has already been successfully integrated into several open-source projects, including: pyca/cryptography issue #13086 pyca/cryptography PR #13130 The runners are provided by IBM and can be seamlessly integrated into existing CI workflows without requiring any infrastructure ch...",https://github.com/pytorch/pytorch/issues/166512,50f7ee11f82b791eb459226f2b0afbca4d5637815d774bf3e905df0f2cdc4b62 references,issue,170038,issue,104587,medium,issue.comments[0].body,"is is not a bug in PyTorch - it's fundamental floating-point behavior that affects all numerical computing frameworks. Similar discussions: #104587, #133006",https://github.com/pytorch/pytorch/issues/170038,e92b09e075203c4d90cca0789af91735e73dea6eb41cb49bc24366a18c0dedbf references,issue,170038,issue,133006,medium,issue.comments[0].body,"a bug in PyTorch - it's fundamental floating-point behavior that affects all numerical computing frameworks. Similar discussions: #104587, #133006",https://github.com/pytorch/pytorch/issues/170038,fbb7a07dcd84c0c3729999e70a1f921e9e79db38d5a1f0f1fd099141cf7f0e77 references,issue,168863,issue,168861,medium,issue.body,Context Source file: test/test_torch.py Test class: TestTorchDeviceType Test name: test_cov Defined at line: 2260 Parent issue: #168861 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_torch.py cc @sunway513 @jithunnair-amd @pruthvistony @ROCmSupport @j,https://github.com/pytorch/pytorch/issues/168863,210323b7d34194a5872a1cffc2a4b78501c65111d72604b4fe80b5f000e41ef6 references,issue,119215,issue,75097,medium,issue.comments[0].body,ks very similar to: #119215 - NCCL abort hangs (collective nature of ncclCommAbort) #115388 - destroy_process_group hangs after CUDA graphs #75097 - Random destroy_process_group hangs The difference is I'm hitting this with: Signal interrupt during training 4-bit quantized mod...,https://github.com/pytorch/pytorch/issues/119215,449d58d039f296f78d15320b7e52e7366d3c2e0ba70e76bddc9f735df3cffa88 references,issue,119215,issue,115388,medium,issue.comments[0].body,uda.empty_cache() Connection to Existing Issues This looks very similar to: #119215 - NCCL abort hangs (collective nature of ncclCommAbort) #115388 - destroy_process_group hangs after CUDA graphs #75097 - Random destroy_process_group hangs The difference is I'm hitting this wi...,https://github.com/pytorch/pytorch/issues/119215,db261f0281d2732b8873cce9ca9064bffd9d37766533c3900bb6d64c07acead8 references,issue,168743,issue,168742,medium,issue.body,xt Source file: test/test_cuda_primary_ctx.py Test class: TestCudaPrimaryCtx Test name: test_set_device_0 Defined at line: 38 Parent issue: #168742 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_cuda_primary_ctx.py cc @sunway513 @jithunnair-amd @pruthvi...,https://github.com/pytorch/pytorch/issues/168743,073e93810fd21dc6fb2ff5093eb24dadf298b51a27dc49c7cd6bbed0e94180eb references,issue,168800,issue,168798,medium,issue.body,Context Source file: test/test_nn.py Test class: TestNN Test name: test_batchnorm Defined at line: 5195 Parent issue: #168798 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_nn.py cc @sunway513 @jithunnair-amd @pruthvistony @ROCmSupport @jata,https://github.com/pytorch/pytorch/issues/168800,fbf3a8ab081c66761537e4ce6ac65f223a7a8aa15a541097bda5ddddfebaabb6 references,issue,112583,issue,65766,medium,issue.comments[0].body,e the same issue with Pytorch 2.0. Some related issues: https://discuss.pytorch.org/t/autocast-and-torch-no-grad-unexpected-behaviour/93475 #65766,https://github.com/pytorch/pytorch/issues/112583,353eef2c457e9047fc3d64add26717646d2ba243b2c94affd4acb7d4e00f4890 references,issue,169934,issue,112583,medium,issue.comments[0].body,"This is a duplicate of #112583 There was a PR to fix it, we should get it through the finish line!",https://github.com/pytorch/pytorch/issues/169934,6ef2a0798e05deec3f4542c330c52a9613fa7312fc7e5d57c0ae811b35ea5c0a references,issue,157648,issue,141287,medium,issue.body,torchvision on linux. I also get the issue if I fiurst set mps as the default device. I was asked to submit a new issue after commenting in #141287 Versions Collecting environment information... PyTorch version: 2.7.1 Is debug build: False CUDA used to build PyTorch: None ROCM...,https://github.com/pytorch/pytorch/issues/157648,430677c1f901c59181cce470b3d94e1cb6aa3d781df41739ca3f053c64cc075f references,issue,168529,issue,168522,medium,issue.body,e file: test/distributed/test_symmetric_memory.py Test class: AsyncTPTest Test name: test_optimal_layout Defined at line: 610 Parent issue: #168522 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/test_symmetric_memory.py cc @sunway513 @jithunnair-...,https://github.com/pytorch/pytorch/issues/168529,670ab92026b206f962b2ab3f6bbfa192c24494f8946ffd470b7818d496700156 references,issue,168528,issue,168522,medium,issue.body,uted/test_symmetric_memory.py Test class: AsyncTPTest Test name: test_fused_scaled_matmul_reduce_scatter Defined at line: 560 Parent issue: #168522 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/test_symmetric_memory.py cc @sunway513 @jithunnair-...,https://github.com/pytorch/pytorch/issues/168528,736012e1dfebaf0216264a45b9274adc43410cb0a3826a2527d649a19152e0e6 references,issue,168527,issue,168522,medium,issue.body,distributed/test_symmetric_memory.py Test class: AsyncTPTest Test name: test_fused_matmul_reduce_scatter Defined at line: 527 Parent issue: #168522 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/test_symmetric_memory.py cc @sunway513 @jithunnair-...,https://github.com/pytorch/pytorch/issues/168527,0e2d09aed4d1ae17f350bd694ab65f6911121befd369e02f9d780e66c36e1e8b references,issue,168526,issue,168522,medium,issue.body,tributed/test_symmetric_memory.py Test class: AsyncTPTest Test name: test_fused_all_gather_scaled_matmul Defined at line: 442 Parent issue: #168522 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/test_symmetric_memory.py cc @sunway513 @jithunnair-...,https://github.com/pytorch/pytorch/issues/168526,4f02161f240b6ab451161dd01480071b8d08fbf77952924d1cf1229bc31bc56a references,issue,168525,issue,168522,medium,issue.body,/distributed/test_symmetric_memory.py Test class: AsyncTPTest Test name: test_multimem_all_gather_matmul Defined at line: 397 Parent issue: #168522 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/test_symmetric_memory.py cc @sunway513 @jithunnair-...,https://github.com/pytorch/pytorch/issues/168525,b4c52c2cddf48304c698fe8a1e6e25ed8bcf7d6e5198647413267aa01f329ab3 references,issue,168524,issue,168522,medium,issue.body,tributed/test_symmetric_memory.py Test class: AsyncTPTest Test name: test_fused_all_gather_matmul_native Defined at line: 344 Parent issue: #168522 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/test_symmetric_memory.py cc @sunway513 @jithunnair-...,https://github.com/pytorch/pytorch/issues/168524,a80be3f2568c10520330b6dd566ff2ff9cf8d89b87d13dff010354818ff6f60c references,issue,168523,issue,168522,medium,issue.body,est/distributed/test_symmetric_memory.py Test class: AsyncTPTest Test name: test_fused_all_gather_matmul Defined at line: 302 Parent issue: #168522 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/test_symmetric_memory.py cc @sunway513 @jithunnair-...,https://github.com/pytorch/pytorch/issues/168523,f868a535be32201caac9d71f1002a665db81a5a5eaa13ae962baf02fbd08acf6 references,issue,168460,issue,168453,medium,issue.body,istributed/test_c10d_gloo.py Test class: ProcessGroupGlooTest Test name: test_block_current_stream_cuda Defined at line: 1628 Parent issue: #168453 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/test_c10d_gloo.py cc @sunway513 @jithunnair-amd @pr...,https://github.com/pytorch/pytorch/issues/168460,e9cbd7921ba2644002d675b02fc5721215153dfcc050a8866a37fb6a01359187 references,issue,168459,issue,168453,medium,issue.body,test/distributed/test_c10d_gloo.py Test class: ProcessGroupGlooTest Test name: test_reduce_stress_cuda Defined at line: 1559 Parent issue: #168453 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/test_c10d_gloo.py cc @sunway513 @jithunnair-amd @pru...,https://github.com/pytorch/pytorch/issues/168459,41970a5948e44789c2d9791889fbc0ff44a4aa803e7186e43d64224e99f4ff56 references,issue,168458,issue,168453,medium,issue.body,st/distributed/test_c10d_gloo.py Test class: ProcessGroupGlooTest Test name: test_allgather_stress_cuda Defined at line: 1372 Parent issue: #168453 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/test_c10d_gloo.py cc @sunway513 @jithunnair-amd @pr...,https://github.com/pytorch/pytorch/issues/168458,6df546d7fa6aeb98bf3dd81de0cb2b30fdd242b25a5ca2a1617585ef53094886 references,issue,168457,issue,168453,medium,issue.body,test/distributed/test_c10d_gloo.py Test class: ProcessGroupGlooTest Test name: test_gather_stress_cuda Defined at line: 1236 Parent issue: #168453 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/test_c10d_gloo.py cc @sunway513 @jithunnair-amd @pru...,https://github.com/pytorch/pytorch/issues/168457,920255e2fb48cb5a04306577c2c4a602774b5ecf9d1c5cd2d8ecde6ea4fc0027 references,issue,168456,issue,168453,medium,issue.body,test/distributed/test_c10d_gloo.py Test class: ProcessGroupGlooTest Test name: test_scatter_stress_cuda Defined at line: 1061 Parent issue: #168453 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/test_c10d_gloo.py cc @sunway513 @jithunnair-amd @pr...,https://github.com/pytorch/pytorch/issues/168456,f30da7c91d46f0018e5efb49e3e34172dab699445fdd04ca17d040428c85a866 references,issue,168455,issue,168453,medium,issue.body,est/distributed/test_c10d_gloo.py Test class: ProcessGroupGlooTest Test name: test_allreduce_stress_cuda Defined at line: 607 Parent issue: #168453 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/test_c10d_gloo.py cc @sunway513 @jithunnair-amd @pr...,https://github.com/pytorch/pytorch/issues/168455,5d9bafb3d5c7b4d007d6faed58becc3fb09190b32f0a6e58b488037a9fa269d6 references,issue,168454,issue,168453,medium,issue.body,est/distributed/test_c10d_gloo.py Test class: ProcessGroupGlooTest Test name: test_broadcast_stress_cuda Defined at line: 449 Parent issue: #168453 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/test_c10d_gloo.py cc @sunway513 @jithunnair-amd @pr...,https://github.com/pytorch/pytorch/issues/168454,48858909e1fe412bbf884b559407e47d755fd6c6dcce1925ddf6596c46f3b96c references,issue,168415,issue,168411,medium,issue.body,tensor/test_sharded_tensor.py Test class: TestShardedTensorChunked Test name: test_multiple_local_shards Defined at line: 947 Parent issue: #168411 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/_shard/sharded_tensor/test_sharded_tensor.py cc @su...,https://github.com/pytorch/pytorch/issues/168415,9e9fd659528ff1b312d200e777ac59dcb6e95fc3d4b90e016c5c547e4c0617b2 references,issue,168414,issue,168411,medium,issue.body,ard/sharded_tensor/test_sharded_tensor.py Test class: TestShardedTensorChunked Test name: test_new_group Defined at line: 891 Parent issue: #168411 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/_shard/sharded_tensor/test_sharded_tensor.py cc @su...,https://github.com/pytorch/pytorch/issues/168414,ff6018737481ff767358a968c23f5968839e3c19713c0fc0cdeb4967843da0a4 references,issue,168413,issue,168411,medium,issue.body,ed_tensor/test_sharded_tensor.py Test class: TestShardedTensorChunked Test name: test_partial_world_size Defined at line: 837 Parent issue: #168411 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/_shard/sharded_tensor/test_sharded_tensor.py cc @su...,https://github.com/pytorch/pytorch/issues/168413,29d7390323d10a3dd2eba1ed58ebcab02836b33230dc96a9bbc76706ea95972d references,issue,168412,issue,168411,medium,issue.body,d_tensor/test_sharded_tensor.py Test class: TestShardedTensorChunked Test name: test_complete_world_size Defined at line: 516 Parent issue: #168411 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/_shard/sharded_tensor/test_sharded_tensor.py cc @su...,https://github.com/pytorch/pytorch/issues/168412,a1340774f10b1670a133aa62e6611ddfdac0fca64873ed3ccaddad1b5ea2f121 references,issue,168431,issue,168424,medium,issue.body,Test class: TestShardedTensorFromLocalShards Test name: test_init_from_local_shards_and_global_metadata Defined at line: 2850 Parent issue: #168424 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/_shard/sharded_tensor/test_sharded_tensor.py cc @su...,https://github.com/pytorch/pytorch/issues/168431,05fd12ab657439c8408b004ec7ac6e6899511c6ad21f00651a947ddf62f4c02a references,issue,168430,issue,168424,medium,issue.body,ShardedTensorFromLocalShards Test name: test_init_from_local_shards_and_global_metadata_with_local_view Defined at line: 2753 Parent issue: #168424 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/_shard/sharded_tensor/test_sharded_tensor.py cc @su...,https://github.com/pytorch/pytorch/issues/168430,d898e0e9b5096763ab8d4fcc7d3caca9077ef1e7fd6ea21e90609802f0ece929 references,issue,168429,issue,168424,medium,issue.body,tShardedTensorFromLocalShards Test name: test_init_from_local_shards_and_global_metadata_with_all_zeros Defined at line: 2678 Parent issue: #168424 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/_shard/sharded_tensor/test_sharded_tensor.py cc @su...,https://github.com/pytorch/pytorch/issues/168429,cdad58bb3525e337057ba3fafd80d332dc3cabbf231267ef695cb2ece1e5a29d references,issue,168428,issue,168424,medium,issue.body,nsor.py Test class: TestShardedTensorFromLocalShards Test name: test_non_rw_sharded_recalc_for_metadata Defined at line: 2583 Parent issue: #168424 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/_shard/sharded_tensor/test_sharded_tensor.py cc @su...,https://github.com/pytorch/pytorch/issues/168428,ea70e331f80bb5278066d2b1703299c21205ef3d8be695291f7ac13a0b163738 references,issue,168427,issue,168424,medium,issue.body,class: TestShardedTensorFromLocalShards Test name: test_init_from_local_shards_with_different_glb_size Defined at line: 2549 Parent issue: #168424 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/_shard/sharded_tensor/test_sharded_tensor.py cc @sun...,https://github.com/pytorch/pytorch/issues/168427,0a296cbd933c836f3ed6d2dfb37c3fa6b3bcfc41c3b484d682fead4f8855a85c references,issue,168426,issue,168424,medium,issue.body,test_sharded_tensor.py Test class: TestShardedTensorFromLocalShards Test name: test_recalc_for_metadata Defined at line: 2494 Parent issue: #168424 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/_shard/sharded_tensor/test_sharded_tensor.py cc @su...,https://github.com/pytorch/pytorch/issues/168426,080466edb9a5acf0cdc169d0acc8f7134e4683360027ce27e24196581df314c6 references,issue,168425,issue,168424,medium,issue.body,t_sharded_tensor.py Test class: TestShardedTensorFromLocalShards Test name: test_init_from_local_shards Defined at line: 2435 Parent issue: #168424 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/_shard/sharded_tensor/test_sharded_tensor.py cc @su...,https://github.com/pytorch/pytorch/issues/168425,cda6d1c02eca80d436f6fc700c896afab2cef43a4c6d3485cd0741b9da402edf references,issue,168423,issue,168422,medium,issue.body,t_sharded_tensor.py Test class: TestShardedTensorFromLocalTensor Test name: test_init_from_local_tensor Defined at line: 2359 Parent issue: #168422 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/_shard/sharded_tensor/test_sharded_tensor.py cc @su...,https://github.com/pytorch/pytorch/issues/168423,b2e6d0dae71350e65114da43142c3a57db99c621e1044297b58a5dd8ffda0a18 references,issue,168421,issue,168416,medium,issue.body,ed_tensor/test_sharded_tensor.py Test class: TestShardedTensorEnumerable Test name: test_with_rpc_names Defined at line: 2225 Parent issue: #168416 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/_shard/sharded_tensor/test_sharded_tensor.py cc @su...,https://github.com/pytorch/pytorch/issues/168421,a230ba4dc7e34aebe5ec1446421d63201af2f27f90a9ac4dd4b0bbb8b5562575 references,issue,168420,issue,168416,medium,issue.body,or/test_sharded_tensor.py Test class: TestShardedTensorEnumerable Test name: test_multiple_local_shards Defined at line: 2141 Parent issue: #168416 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/_shard/sharded_tensor/test_sharded_tensor.py cc @su...,https://github.com/pytorch/pytorch/issues/168420,07260e7c5874ab99bfff7a23c06b8910234c3c3d19c210e9c22e248c817feb32 references,issue,168419,issue,168416,medium,issue.body,sharded_tensor/test_sharded_tensor.py Test class: TestShardedTensorEnumerable Test name: test_new_group Defined at line: 2072 Parent issue: #168416 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/_shard/sharded_tensor/test_sharded_tensor.py cc @su...,https://github.com/pytorch/pytorch/issues/168419,3961e56075559b0108ec4e33faec37d616d803211a15c7e1614137ff2d2a78fd references,issue,168418,issue,168416,medium,issue.body,ensor/test_sharded_tensor.py Test class: TestShardedTensorEnumerable Test name: test_partial_world_size Defined at line: 2005 Parent issue: #168416 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/_shard/sharded_tensor/test_sharded_tensor.py cc @su...,https://github.com/pytorch/pytorch/issues/168418,77eb7da1c15cad9d862b48adb030257e4f631177120c08ade07a909fc23f13b2 references,issue,168417,issue,168416,medium,issue.body,ded_tensor/test_sharded_tensor.py Test class: TestShardedTensorEnumerable Test name: test_grid_sharding Defined at line: 1487 Parent issue: #168416 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/_shard/sharded_tensor/test_sharded_tensor.py cc @su...,https://github.com/pytorch/pytorch/issues/168417,d3f1457467e6c731f6c33fa4dddfe3e6802c8de441b2867f92cb52a0b0819ee1 references,issue,169429,issue,166628,medium,issue.body,"tion routine failed. Error loading ""...\Lib\site-packages\torch\lib\torch_python.dll"" or one of its dependencies. The failure is similar to #166628 however in this case import torch always fails even when importing it as the first package. This is also mentioned by some users...",https://github.com/pytorch/pytorch/issues/169429,1c5199e382e337f9b466f480844aecc788ed172d1bd91660e63985648ba7d9a9 references,issue,168624,issue,168623,medium,issue.body,tor/test_minifier_isolate.py Test class: MinifierIsolateTests Test name: test_after_aot_gpu_runtime_error Defined at line: 46 Parent issue: #168623 Code reference: https://github.com/pytorch/pytorch/blob/main/test/inductor/test_minifier_isolate.py cc @sunway513 @jithunnair-amd...,https://github.com/pytorch/pytorch/issues/168624,ebf245bc36741b4d02836f5446d3d4c7d007fc85d1ab0c952ed79c066e4a1ea3 references,issue,160553,issue,153378,medium,issue.comments[0].body,"Feels almost like a duplicate of #153378 Metal does not have a concept of exceptions, but may be some sort of shared memory flag would be a reasonable solution...",https://github.com/pytorch/pytorch/issues/160553,8a6c5e2b0a97e431d1b596aa1a8e071367d7a10ef260ab3d88dae184b7fa31b6 references,issue,137767,issue,135860,medium,issue.body,es: torch.linalg.norm & torch.norm: #132634 torch.acos: #134487 torch.nn.functional.normalize: #135428 torch.sigmoid: #135777 torch.lobpcg: #135860 torch.exp: #136063 torch.asin: #138327 #141487 #156152 cc @jgong5 @mingfeima @XiaobingSuper @sanchitintel @ashokei @jingxu10 @ezy...,https://github.com/pytorch/pytorch/issues/137767,0ff3e657cd71330f9619146804d900de8fb0b6b15e797ffd4bd748e502ae1561 references,issue,137767,issue,156152,medium,issue.body,#134487 torch.nn.functional.normalize: #135428 torch.sigmoid: #135777 torch.lobpcg: #135860 torch.exp: #136063 torch.asin: #138327 #141487 #156152 cc @jgong5 @mingfeima @XiaobingSuper @sanchitintel @ashokei @jingxu10 @ezyang @anjali411 @dylanbespalko @mruberry @nikitaved @amjames,https://github.com/pytorch/pytorch/issues/137767,afabec771be9db81117ee9af189d2cbbb61f71f47215c03192cc895dc00ac20c references,issue,165612,issue,40471,medium,issue.body,"xtension based on NumPy's dtype registration system (and therefore provides proper NumPy dtype types and objects) Related past discussions: #40471, #40568 cc @albanD @rgommers (NumPy/SciPy) @lucascolley @ev-br (SciPy) @kmaehashi (CuPy) @aterrel @rparolin (CUDA Python) @jrhemst...",https://github.com/pytorch/pytorch/issues/165612,2384859ecdda4e171992e61c690370fe5d901544a7614ea7470da60e819dcd9a references,issue,165612,issue,40568,medium,issue.body,"based on NumPy's dtype registration system (and therefore provides proper NumPy dtype types and objects) Related past discussions: #40471, #40568 cc @albanD @rgommers (NumPy/SciPy) @lucascolley @ev-br (SciPy) @kmaehashi (CuPy) @aterrel @rparolin (CUDA Python) @jrhemstad (CUDA...",https://github.com/pytorch/pytorch/issues/165612,37636a42e49c220a19f442b1379222734ab0a6ab526bd17e7631f456cb7178c3 references,issue,65151,issue,62032,medium,issue.body,"🚀 Feature As originally pointed out here: #62032 The boxed fallback has an added benefit that we could then consider, at a later point in time, to register it as the universal Autograd fal",https://github.com/pytorch/pytorch/issues/65151,1f89b05e066bd9df1e2d0fd1725fbc32b32b731d52f1eb2acc7f8779beca3aef references,issue,167656,issue,91692,medium,issue.comments[0].body,"tly filled by the user, and can then help with interpreting outputs of various debugging/tracing/profiling scenarios... e.g. as I tried in: #91692 (comment)",https://github.com/pytorch/pytorch/issues/167656,ce616722a73a27f355b6feb3e3b3d166edf83f498153f18e3ff77d946a4642f3 references,issue,167656,issue,104247,medium,issue.comments[0].body,"Regarding identity tracking, I'm proposing to add a .name attribute to tensors/modules etc #104247 Such attributes can be explicitly filled by the user, and can then help with interpreting outputs of various debugging/tracing/profiling sc",https://github.com/pytorch/pytorch/issues/167656,5bc8ec3709d62a64e3ceeb7b02e0b4b896faddc89bf2d0eb040eb3ba2d2972f3 references,issue,167656,issue,168976,medium,issue.comments[1].body,@yushangdi seems similar to #168976? I think there's settings we should flip as default: record_nn_module_stack=True record_stack_trace=True record_output=True record_ids=True,https://github.com/pytorch/pytorch/issues/167656,73d18e7a4d124910b492219d98f3a7002e6b119102dbb7c6d3abacfbd72ce55a references,issue,168403,issue,168402,medium,issue.body,est_fully_shard_overlap.py Test class: TestFullyShardOverlap Test name: test_fully_shard_training_overlap Defined at line: 57 Parent issue: #168402 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/_composable/fsdp/test_fully_shard_overlap.py cc @su...,https://github.com/pytorch/pytorch/issues/168403,2e13b350b98792b5be3a1dd683b1ad33ba67fe7043df1f8f08781c729d32355e references,issue,168405,issue,168404,medium,issue.body,p/test_fully_shard_training.py Test class: TestFullyShardNDTraining Test name: test_2d_mlp_with_nd_mesh Defined at line: 1214 Parent issue: #168404 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/_composable/fsdp/test_fully_shard_training.py cc @s...,https://github.com/pytorch/pytorch/issues/168405,e576ca80f6fa45fcfd55d2ca218341dc3bc68d47d71c10cb900b02e467de8663 references,issue,168407,issue,168406,medium,issue.body,bility/test_2d_composability.py Test class: TestFullyShard2DTraining Test name: test_train_parity_2d_mlp Defined at line: 130 Parent issue: #168406 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/_composable/test_composability/test_2d_composabilit...,https://github.com/pytorch/pytorch/issues/168407,d8f03951cf23612d217b059b079c12df169bd35712c3214fd74f2e15016f7adf references,issue,168476,issue,168475,medium,issue.body,ource file: test/distributed/test_composability.py Test class: ComposabilityTest Test name: test_pp_fsdp Defined at line: 288 Parent issue: #168475 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/test_composability.py cc @sunway513 @jithunnair-amd...,https://github.com/pytorch/pytorch/issues/168476,362b4eb05552186a355830bff700b2325ffa5bae87270cfc56bd7f9a1c3ceb9d references,issue,168452,issue,168451,medium,issue.body,pute_reordering.py Test class: TestComputeCommReorderingMultiProc Test name: test_grouped_scheduler_node Defined at line: 333 Parent issue: #168451 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/test_aten_comm_compute_reordering.py cc @sunway513...,https://github.com/pytorch/pytorch/issues/168452,d5a96373859e0a08f1dcbc4a667cd078db5ae4a8a7379f64901804ef1fece22e references,issue,168477,issue,168475,medium,issue.body,uted/test_composability.py Test class: ComposabilityTest Test name: test_pp_fsdp_unshard_reshard_runtime Defined at line: 389 Parent issue: #168475 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/test_composability.py cc @sunway513 @jithunnair-amd...,https://github.com/pytorch/pytorch/issues/168477,2df565ea0bc586d6e1da01674374106bfa515a69d3b06e5dd07a86904fa1322c references,issue,168537,issue,168536,medium,issue.body,tributed/test_symmetric_memory.py Test class: LoweringTest Test name: test_lowering_one_shot_all_reduce Defined at line: 1167 Parent issue: #168536 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/test_symmetric_memory.py cc @sunway513 @jithunnair-...,https://github.com/pytorch/pytorch/issues/168537,9aaad870f60daea91c7af8ef0c842273fbfb47e58b29ef6d77f7b291b02052af references,issue,168568,issue,168563,medium,issue.body,uctor.py Test class: AOTInductorTestsTemplate Test name: test__weight_int4pack_mm_with_scales_and_zeros Defined at line: 6799 Parent issue: #168563 Code reference: https://github.com/pytorch/pytorch/blob/main/test/inductor/test_aot_inductor.py cc @sunway513 @jithunnair-amd @pr...,https://github.com/pytorch/pytorch/issues/168568,b99327355f513d3ca75c91e7e29e01eda8b81c639481f9c887310b32104024d5 references,issue,168567,issue,168563,medium,issue.body,/inductor/test_aot_inductor.py Test class: AOTInductorTestsTemplate Test name: test__weight_int4pack_mm Defined at line: 6762 Parent issue: #168563 Code reference: https://github.com/pytorch/pytorch/blob/main/test/inductor/test_aot_inductor.py cc @sunway513 @jithunnair-amd @pr...,https://github.com/pytorch/pytorch/issues/168567,d5d963cf7fc761b99c465e10d49322aec34c85e96cb27b208e83fe5cbe1ad7b2 references,issue,168566,issue,168563,medium,issue.body,uctor.py Test class: AOTInductorTestsTemplate Test name: test_autotune_int64_user_defined_triton_kernel Defined at line: 5861 Parent issue: #168563 Code reference: https://github.com/pytorch/pytorch/blob/main/test/inductor/test_aot_inductor.py cc @sunway513 @jithunnair-amd @pr...,https://github.com/pytorch/pytorch/issues/168566,4f18a46807ae634ccf14c6b7beef4dafa0b34487f0e98d6b298c2e33d4d5042d references,issue,168646,issue,168644,medium,issue.body,e: test/inductor/test_multi_kernel.py Test class: MultiKernelTest Test name: test_triton_relu_fused_gemm Defined at line: 144 Parent issue: #168644 Code reference: https://github.com/pytorch/pytorch/blob/main/test/inductor/test_multi_kernel.py cc @sunway513 @jithunnair-amd @pr...,https://github.com/pytorch/pytorch/issues/168646,08ab8281c31fabfdfe79a58e0026521535ad53c0cc3eddc66343f96a0ad4b4ee references,issue,168645,issue,168644,medium,issue.body,Source file: test/inductor/test_multi_kernel.py Test class: MultiKernelTest Test name: test_triton_gemm Defined at line: 115 Parent issue: #168644 Code reference: https://github.com/pytorch/pytorch/blob/main/test/inductor/test_multi_kernel.py cc @sunway513 @jithunnair-amd @pru...,https://github.com/pytorch/pytorch/issues/168645,c65abffbc68d4bb9ad29b71a5ae27c9b7ad5a4155e1a1a72a509b1d86167ba1f references,issue,168707,issue,168706,medium,issue.body,e file: test/test_cpp_extensions_aot.py Test class: TestCppExtensionAOT Test name: test_cublas_extension Defined at line: 132 Parent issue: #168706 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_cpp_extensions_aot.py cc @sunway513 @jithunnair-amd @pruth...,https://github.com/pytorch/pytorch/issues/168707,d91e3a690630e46a27b411ac2b7c75e00706374061abe654beca34a77da7066b references,issue,168708,issue,168706,medium,issue.body,file: test/test_cpp_extensions_aot.py Test class: TestCppExtensionAOT Test name: test_cusolver_extension Defined at line: 142 Parent issue: #168706 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_cpp_extensions_aot.py cc @sunway513 @jithunnair-amd @pruth...,https://github.com/pytorch/pytorch/issues/168708,049b57490056681ffd0905155002e26a025ba944e9e76c4a6eecc10dda392819 references,issue,168709,issue,168706,medium,issue.body,ce file: test/test_cpp_extensions_aot.py Test class: TestCppExtensionAOT Test name: test_cuda_dlink_libs Defined at line: 176 Parent issue: #168706 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_cpp_extensions_aot.py cc @sunway513 @jithunnair-amd @pruth...,https://github.com/pytorch/pytorch/issues/168709,d9a58dadc139b48db70df390691574ed744b8be017a3270ab9fa32068c902136 references,issue,168766,issue,168756,medium,issue.body,ext Source file: test/test_linalg.py Test class: TestLinalg Test name: test__dyn_quant_pack_4bit_weight Defined at line: 7844 Parent issue: #168756 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_linalg.py cc @sunway513 @jithunnair-amd @pruthvistony @ROC...,https://github.com/pytorch/pytorch/issues/168766,594b8f578dfcd198463b048f8b47ed1ffee6484e8762c2b57586b1b3f6f2f322 references,issue,168767,issue,168756,medium,issue.body,Context Source file: test/test_linalg.py Test class: TestLinalg Test name: test__dyn_quant_matmul_4bit Defined at line: 7872 Parent issue: #168756 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_linalg.py cc @sunway513 @jithunnair-amd @pruthvistony @ROCm...,https://github.com/pytorch/pytorch/issues/168767,a714b1f6ddc88cb892d6aa83ff5de0fa4fb953bfa047e6e02e038414f706b939 references,issue,168768,issue,168756,medium,issue.body,t Source file: test/test_linalg.py Test class: TestLinalg Test name: test_compile_dyn_quant_matmul_4bit Defined at line: 7944 Parent issue: #168756 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_linalg.py cc @sunway513 @jithunnair-amd @pruthvistony @ROC...,https://github.com/pytorch/pytorch/issues/168768,597c57306b241b5d961e2760aa92d04b7fe1598bcbff9e91d46468621db2470b references,issue,168771,issue,168756,medium,issue.body,Context Source file: test/test_linalg.py Test class: TestLinalg Test name: test_ldl_solve Defined at line: 9904 Parent issue: #168756 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_linalg.py cc @sunway513 @jithunnair-amd @pruthvistony @ROCmSupport @,https://github.com/pytorch/pytorch/issues/168771,7ca558fc0854c3dd06c0aa4602ffb45a20e37c5735ab48389d6df024d021227e references,issue,168772,issue,168756,medium,issue.body,Context Source file: test/test_linalg.py Test class: TestLinalg Test name: test_ck_blas_library Defined at line: 9977 Parent issue: #168756 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_linalg.py cc @sunway513 @jithunnair-amd @pruthvistony @ROCmSupport @,https://github.com/pytorch/pytorch/issues/168772,8ba351bda6b39ee6bd2227325f6c6badfcd1d9d36562262214842ae4f2fb5980 references,issue,168777,issue,168773,medium,issue.body,t Source file: test/test_matmul_cuda.py Test class: TestMatmulCuda Test name: test_grouped_gemm_compiled Defined at line: 569 Parent issue: #168773 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_matmul_cuda.py cc @sunway513 @jithunnair-amd @pruthvistony...,https://github.com/pytorch/pytorch/issues/168777,fd6bc77e759194d605e951771dc2a920bfcc8cdeee6e23a91738c8b11d061453 references,issue,168778,issue,168773,medium,issue.body,t Source file: test/test_matmul_cuda.py Test class: TestMatmulCuda Test name: test_mm_bmm_dtype_overload Defined at line: 690 Parent issue: #168773 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_matmul_cuda.py cc @sunway513 @jithunnair-amd @pruthvistony...,https://github.com/pytorch/pytorch/issues/168778,1ca25dc04424d8e69baabe7e7d0d83690b45184a73bca80d22211fb3dc73e2f2 references,issue,168776,issue,168773,medium,issue.body,e: test/test_matmul_cuda.py Test class: TestMatmulCuda Test name: test_cublas_batch_invariance_blackwell Defined at line: 368 Parent issue: #168773 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_matmul_cuda.py cc @sunway513 @jithunnair-amd @pruthvistony...,https://github.com/pytorch/pytorch/issues/168776,9156b0a3e6df22f548b8df0a6bb8321cce4752adf55865358059ae2f4e1e21b5 references,issue,168781,issue,168780,medium,issue.body,file: test/test_matmul_cuda.py Test class: TestMixedDtypesLinearCuda Test name: test_mixed_dtypes_linear Defined at line: 955 Parent issue: #168780 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_matmul_cuda.py cc @sunway513 @jithunnair-amd @pruthvistony...,https://github.com/pytorch/pytorch/issues/168781,1e18b5b3a092ae5ebd2173b6ba057efd9fd04c8006fdcd5c516563445a112265 references,issue,168832,issue,168831,medium,issue.body,Test class: SparseSemiStructuredTensorCompileTest Test name: test_mlp_contiguous_relu_compile_cusparselt Defined at line: 228 Parent issue: #168831 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_sparse_semi_structured.py cc @nikitaved @pearu @cpuhrsch @...,https://github.com/pytorch/pytorch/issues/168832,3308d9e793568655634ab8b98172e66d491aae6d00d708cf6adf9ea3dc9b86dc references,issue,168833,issue,168831,medium,issue.body,py Test class: SparseSemiStructuredTensorCompileTest Test name: test_mlp_contiguous_relu_compile_cutlass Defined at line: 239 Parent issue: #168831 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_sparse_semi_structured.py cc @nikitaved @pearu @cpuhrsch @...,https://github.com/pytorch/pytorch/issues/168833,45db4405e9e94f50645f7e568956629beb371d60af896394e55e0713ddd70ae7 references,issue,168834,issue,168831,medium,issue.body,sparse_semi_structured.py Test class: SparseSemiStructuredTensorCompileTest Test name: test_sp24_compile Defined at line: 254 Parent issue: #168831 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_sparse_semi_structured.py cc @nikitaved @pearu @cpuhrsch @...,https://github.com/pytorch/pytorch/issues/168834,21fe603a169f17c49269d04de9e410bc877e150fe71421a09a1268c544802b1a references,issue,168836,issue,168835,medium,issue.body,_semi_structured.py Test class: TestSparseSemiStructuredTraining Test name: test_prune_dense_static_sort Defined at line: 587 Parent issue: #168835 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_sparse_semi_structured.py cc @nikitaved @pearu @cpuhrsch @...,https://github.com/pytorch/pytorch/issues/168836,d9f473759b9768742174d2aa249f399327013788b51d32bcc9eead78cfcd7c92 references,issue,168837,issue,168835,medium,issue.body,d.py Test class: TestSparseSemiStructuredTraining Test name: test_pruning_algo_largest_abs_values_greedy Defined at line: 636 Parent issue: #168835 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_sparse_semi_structured.py cc @nikitaved @pearu @cpuhrsch @...,https://github.com/pytorch/pytorch/issues/168837,29a1e3b8d424885b424dd6fd2d4ebd80b9818636e7e7a47df09e0ac71872b112 references,issue,168838,issue,168835,medium,issue.body,ructured.py Test class: TestSparseSemiStructuredTraining Test name: test_pack_both_ways_meta_correctness Defined at line: 677 Parent issue: #168835 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_sparse_semi_structured.py cc @nikitaved @pearu @cpuhrsch @...,https://github.com/pytorch/pytorch/issues/168838,15fd5c11390336adf63fc921a2a7aea80b5bbdc52d0d69c71ede4ef4da16b26c references,issue,168840,issue,168835,medium,issue.body,emi_structured.py Test class: TestSparseSemiStructuredTraining Test name: test_pack_both_ways_edge_case1 Defined at line: 756 Parent issue: #168835 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_sparse_semi_structured.py cc @nikitaved @pearu @cpuhrsch @...,https://github.com/pytorch/pytorch/issues/168840,2abeadd5a85af7401f8fc62e3dc759741a719288fd9f7a53fa21e0c8dbb78dda references,issue,168839,issue,168835,medium,issue.body,sparse_semi_structured.py Test class: TestSparseSemiStructuredTraining Test name: test_pack_both_ways_id Defined at line: 715 Parent issue: #168835 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_sparse_semi_structured.py cc @nikitaved @pearu @cpuhrsch @...,https://github.com/pytorch/pytorch/issues/168839,3463726337f2ba3c3ce0914933e67a75f8f7b092b1638233c516e70562ac0c2f references,issue,168842,issue,168835,medium,issue.body,_sparse_semi_structured.py Test class: TestSparseSemiStructuredTraining Test name: test_sp24_apply_dense Defined at line: 805 Parent issue: #168835 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_sparse_semi_structured.py cc @nikitaved @pearu @cpuhrsch @...,https://github.com/pytorch/pytorch/issues/168842,07b4f67ba86c3a80d4fce5989292e89480c806e7fd7b44680504616c74c99eab references,issue,168841,issue,168835,medium,issue.body,t/test_sparse_semi_structured.py Test class: TestSparseSemiStructuredTraining Test name: test_sp24_apply Defined at line: 785 Parent issue: #168835 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_sparse_semi_structured.py cc @nikitaved @pearu @cpuhrsch @...,https://github.com/pytorch/pytorch/issues/168841,47bc97e8751ffadb0d230e46261a10f5963d9196f0f4b16343075bb681a7e9d2 references,issue,168843,issue,168835,medium,issue.body,test_sparse_semi_structured.py Test class: TestSparseSemiStructuredTraining Test name: test_sp24_matmuls Defined at line: 847 Parent issue: #168835 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_sparse_semi_structured.py cc @nikitaved @pearu @cpuhrsch @...,https://github.com/pytorch/pytorch/issues/168843,6d305bacb30fb3b0effb76a81e35720cf0d4baaf212524a07d48319a273abe6a references,issue,168844,issue,168835,medium,issue.body,rse_semi_structured.py Test class: TestSparseSemiStructuredTraining Test name: test_sp24_matmuls_mat_vec Defined at line: 886 Parent issue: #168835 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_sparse_semi_structured.py cc @nikitaved @pearu @cpuhrsch @...,https://github.com/pytorch/pytorch/issues/168844,c9b37ff801c8938358fee822390ae5bf19fb34e31d00c46828b6938acecbce10 references,issue,168845,issue,168835,medium,issue.body,_sparse_semi_structured.py Test class: TestSparseSemiStructuredTraining Test name: test_sp24_matmuls_bmm Defined at line: 900 Parent issue: #168835 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_sparse_semi_structured.py cc @nikitaved @pearu @cpuhrsch @...,https://github.com/pytorch/pytorch/issues/168845,a43b9949db23847a9aeeb9f3eddcf9919d6951aab1e768a531b4e6206de0767a references,issue,168848,issue,168846,medium,issue.body,ctured.py Test class: TestSparseSemiStructuredCUTLASS Test name: test_sparse_semi_structured_ops_cutlass Defined at line: 984 Parent issue: #168846 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_sparse_semi_structured.py cc @nikitaved @pearu @cpuhrsch @...,https://github.com/pytorch/pytorch/issues/168848,16d1925109a0dcc6d924c031980e92af51874003b4ebb71d505a0a81cbcd293c references,issue,168847,issue,168846,medium,issue.body,est_sparse_semi_structured.py Test class: TestSparseSemiStructuredCUTLASS Test name: test_linear_cutlass Defined at line: 925 Parent issue: #168846 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_sparse_semi_structured.py cc @nikitaved @pearu @cpuhrsch @...,https://github.com/pytorch/pytorch/issues/168847,4ab1324aeaf980026aa2975d662041a83d2fe3235074c7b9259b943d7028601c references,issue,168850,issue,168846,medium,issue.body,semi_structured.py Test class: TestSparseSemiStructuredCUTLASS Test name: test_conversions_all_patterns Defined at line: 1080 Parent issue: #168846 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_sparse_semi_structured.py cc @nikitaved @pearu @cpuhrsch @...,https://github.com/pytorch/pytorch/issues/168850,1f760b85f198743d10750c0a9743a6f510b2d0dee35b81f332608218119e78b8 references,issue,168849,issue,168846,medium,issue.body,/test_sparse_semi_structured.py Test class: TestSparseSemiStructuredCUTLASS Test name: test_conversions Defined at line: 1051 Parent issue: #168846 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_sparse_semi_structured.py cc @nikitaved @pearu @cpuhrsch @...,https://github.com/pytorch/pytorch/issues/168849,827acc21487fd51812e26f2b2e4ad294e7ad46e94be946532793e462fc5213ea references,issue,168852,issue,168851,medium,issue.body,_semi_structured.py Test class: TestSparseSemiStructuredCUSPARSELT Test name: test_cslt_sparse_mm_alpha Defined at line: 1200 Parent issue: #168851 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_sparse_semi_structured.py cc @nikitaved @pearu @cpuhrsch @...,https://github.com/pytorch/pytorch/issues/168852,24553d3937da9d8cfd0a617703595c18c36369be794d91b576103b9b1a9ca6d3 references,issue,168853,issue,168851,medium,issue.body,py Test class: TestSparseSemiStructuredCUSPARSELT Test name: test_cslt_sparse_mm_alpha_compile_autotune Defined at line: 1217 Parent issue: #168851 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_sparse_semi_structured.py cc @nikitaved @pearu @cpuhrsch @...,https://github.com/pytorch/pytorch/issues/168853,5abbf29fa4122abf7b3e354a42d03012f6ab447a5dd8197a332e7988fba5af97 references,issue,168854,issue,168851,medium,issue.body,ured.py Test class: TestSparseSemiStructuredCUSPARSELT Test name: test_cslt_sparse_mm_alpha_mixed_dtype Defined at line: 1239 Parent issue: #168851 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_sparse_semi_structured.py cc @nikitaved @pearu @cpuhrsch @...,https://github.com/pytorch/pytorch/issues/168854,8100450a2f8b47059f061759d02c254cd00e6515d872dac96936bff7d2cb3f52 references,issue,168446,issue,168445,medium,issue.body,t/distributed/tensor/test_attention.py Test class: RingAttentionTest Test name: test_ring_attention_sdpa Defined at line: 112 Parent issue: #168445 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/tensor/test_attention.py cc @sunway513 @jithunnair-...,https://github.com/pytorch/pytorch/issues/168446,7b43ef17637d68085a912d1bb8da8b54eb7fea7a02555680f659251172bd1802 references,issue,168788,issue,168784,medium,issue.body,rce file: test/test_nestedtensor.py Test class: TestNestedTensorSubclass Test name: test_sdpa_backwards Defined at line: 7208 Parent issue: #168784 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_nestedtensor.py cc @sunway513 @jithunnair-amd @pruthviston...,https://github.com/pytorch/pytorch/issues/168788,6ce9e3689a0334908151ab327a6f7d8c5393631a04f249934ee3db09db433eef references,issue,168789,issue,168784,medium,issue.body,urce file: test/test_nestedtensor.py Test class: TestNestedTensorSubclass Test name: test_sdpa_autocast Defined at line: 7234 Parent issue: #168784 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_nestedtensor.py cc @sunway513 @jithunnair-amd @pruthviston...,https://github.com/pytorch/pytorch/issues/168789,68f0efdc86afd53f5a0ee60c19593d294f8a878d34ddb9e07819f64ee1ad31b5 references,issue,168872,issue,168871,medium,issue.body,e: test/test_transformers.py Test class: TestSDPACudaOnly Test name: test_cudnn_attention_broken_166211 Defined at line: 2856 Parent issue: #168871 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_transformers.py cc @sunway513 @jithunnair-amd @pruthviston...,https://github.com/pytorch/pytorch/issues/168872,eb92b6bccc65eddc9d34b1effc1e883bbf5bbe12843f87c93a63cfe2a9ae4286 references,issue,168876,issue,168871,medium,issue.body,est/test_transformers.py Test class: TestSDPACudaOnly Test name: test_fused_kernels_nested_broadcasting Defined at line: 3977 Parent issue: #168871 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_transformers.py cc @sunway513 @jithunnair-amd @pruthviston...,https://github.com/pytorch/pytorch/issues/168876,02425c3044ee7d25545b51f20460db149a41c502e399cc4a605f05a784db4809 references,issue,168875,issue,168871,medium,issue.body,transformers.py Test class: TestSDPACudaOnly Test name: test_fused_backwards_throws_determinism_warning Defined at line: 3311 Parent issue: #168871 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_transformers.py cc @sunway513 @jithunnair-amd @pruthviston...,https://github.com/pytorch/pytorch/issues/168875,afb4976f5781eeb755f9df86ee5492752c7d6c6d64cb254f97e7ec2337ce9bd9 references,issue,168877,issue,168871,medium,issue.body,nsformers.py Test class: TestSDPACudaOnly Test name: test_fused_kernels_nested_broadcasting_query_dense Defined at line: 4061 Parent issue: #168871 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_transformers.py cc @sunway513 @jithunnair-amd @pruthviston...,https://github.com/pytorch/pytorch/issues/168877,d327ef29ea56a53825ae383f054602f30c3d757df9d1f1ff656549fa7e5cb9aa references,issue,168481,issue,168480,medium,issue.body,ile: test/distributed/test_nccl.py Test class: NCCLSymmetricMemoryTest Test name: test_nccl_symmem_alloc Defined at line: 231 Parent issue: #168480 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/test_nccl.py cc @sunway513 @jithunnair-amd @pruthvi...,https://github.com/pytorch/pytorch/issues/168481,420596d509160d3a51432b04f8e3739085145b8d8701ebb7db1b6c7bc97a8887 references,issue,168533,issue,168532,medium,issue.body,est/distributed/test_symmetric_memory.py Test class: SymmMemNegativeTest Test name: test_barrier_timeout Defined at line: 787 Parent issue: #168532 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/test_symmetric_memory.py cc @sunway513 @jithunnair-...,https://github.com/pytorch/pytorch/issues/168533,192192f35fa3206c63dcb0b12f385dbfea3652c1b79812bc0fc4696c7be41edb references,issue,168534,issue,168532,medium,issue.body,/distributed/test_symmetric_memory.py Test class: SymmMemNegativeTest Test name: test_put_signal_timeout Defined at line: 813 Parent issue: #168532 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/test_symmetric_memory.py cc @sunway513 @jithunnair-...,https://github.com/pytorch/pytorch/issues/168534,344d69083c9eafdcae2e16cfd1f0fbcac8276db093c23c75d2e8c26a411d0d5e references,issue,168535,issue,168532,medium,issue.body,distributed/test_symmetric_memory.py Test class: SymmMemNegativeTest Test name: test_wait_signal_timeout Defined at line: 842 Parent issue: #168532 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/test_symmetric_memory.py cc @sunway513 @jithunnair-...,https://github.com/pytorch/pytorch/issues/168535,3525976c173cfdb9c4a833e5cab61004deabcf32326facff556ad0fd02632325 references,issue,168515,issue,168503,medium,issue.body,e: test/distributed/test_nvshmem_triton.py Test class: NVSHMEMTritonTest Test name: test_triton_alltoall Defined at line: 844 Parent issue: #168503 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/test_nvshmem_triton.py cc @sunway513 @jithunnair-am...,https://github.com/pytorch/pytorch/issues/168515,c8e64de4e0f953215e36d928b1cd2766597c367e4e0153ea88b4ea775cb65389 references,issue,168516,issue,168503,medium,issue.body,: test/distributed/test_nvshmem_triton.py Test class: NVSHMEMTritonTest Test name: test_triton_broadcast Defined at line: 891 Parent issue: #168503 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/test_nvshmem_triton.py cc @sunway513 @jithunnair-am...,https://github.com/pytorch/pytorch/issues/168516,879902a7c2b0e246493abd18007e9928b214a353a1219fcc7d1e1d7a2db19f24 references,issue,168518,issue,168503,medium,issue.body,t/distributed/test_nvshmem_triton.py Test class: NVSHMEMTritonTest Test name: test_triton_minmax_reduce Defined at line: 1021 Parent issue: #168503 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/test_nvshmem_triton.py cc @sunway513 @jithunnair-am...,https://github.com/pytorch/pytorch/issues/168518,416838559afeefcdd0be5f61cc38dda576eb79881e8d9c23ba8d45fd837c6924 references,issue,168517,issue,168503,medium,issue.body,test/distributed/test_nvshmem_triton.py Test class: NVSHMEMTritonTest Test name: test_triton_sum_reduce Defined at line: 959 Parent issue: #168503 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/test_nvshmem_triton.py cc @sunway513 @jithunnair-amd...,https://github.com/pytorch/pytorch/issues/168517,a5408a555a1a44ac9879dd36ffb1a0801d82fe254a44c0cef1bfbe46ae7b67c0 references,issue,168519,issue,168503,medium,issue.body,est/distributed/test_nvshmem_triton.py Test class: NVSHMEMTritonTest Test name: test_triton_prod_reduce Defined at line: 1107 Parent issue: #168503 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/test_nvshmem_triton.py cc @sunway513 @jithunnair-am...,https://github.com/pytorch/pytorch/issues/168519,b0527e75884c7468afaa701a07e60bc199de9e499e4026f28ffcdeaf7c2f76f6 references,issue,168435,issue,168434,medium,issue.body,/utils/distributed_test.py Test class: DistributedUtilTest Test name: test_port_already_in_use_on_worker Defined at line: 178 Parent issue: #168434 Code reference: https://github.com/pytorch/pytorch/blob/main/test/distributed/elastic/utils/distributed_test.py cc @sunway513 @ji...,https://github.com/pytorch/pytorch/issues/168435,b77e4ebc849d2ee4f818f5981077c2bd1c51040b447ae65128189bdf6bfe3a70 references,issue,168868,issue,168861,medium,issue.body,ontext Source file: test/test_torch.py Test class: TestTorchDeviceType Test name: test_pdist_norm_large Defined at line: 4066 Parent issue: #168861 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_torch.py cc @sunway513 @jithunnair-amd @pruthvistony @ROCm...,https://github.com/pytorch/pytorch/issues/168868,385ca7cdd086d2ea185f7101e5002a61f131da8688c40a8815127d620ab14962 references,issue,168806,issue,168801,medium,issue.body,: test/test_nn.py Test class: TestNNDeviceType Test name: test_upsampling_64bit_indexing_channels_last Defined at line: 10482 Parent issue: #168801 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_nn.py cc @sunway513 @jithunnair-amd @pruthvistony @ROCmSup...,https://github.com/pytorch/pytorch/issues/168806,afefdc4934dd8efe5bc3a22288bcd81e89fffe476c6edf01299474914204299a references,issue,168792,issue,168784,medium,issue.body,t_nestedtensor.py Test class: TestNestedTensorSubclass Test name: test_compile_preserves_metadata_cache Defined at line: 7743 Parent issue: #168784 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_nestedtensor.py cc @sunway513 @jithunnair-amd @pruthviston...,https://github.com/pytorch/pytorch/issues/168792,c1cbdb7373b1b64450b729f57adade2c16b4a79d58f5ee2ecc82e2fe5d487555 references,issue,168791,issue,168784,medium,issue.body,file: test/test_nestedtensor.py Test class: TestNestedTensorSubclass Test name: test_dummy_mha_with_nt Defined at line: 7388 Parent issue: #168784 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_nestedtensor.py cc @sunway513 @jithunnair-amd @pruthvistony...,https://github.com/pytorch/pytorch/issues/168791,7be3a2d8391e02933357e5247a50df45dcc32a84d94ff5da79ebc7cdb4a56323 references,issue,168721,issue,168720,medium,issue.body,text Source file: test/test_cuda.py Test class: TestCudaMallocAsync Test name: test_memory_profiler_viz Defined at line: 4139 Parent issue: #168720 Code reference: https://github.com/pytorch/pytorch/blob/main/test/test_cuda.py cc @sunway513 @jithunnair-amd @pruthvistony @ROCmS...,https://github.com/pytorch/pytorch/issues/168721,e7d708a80ffe867fbab363362176d42ce9b54253cf6e4ce4b556d8f17630c812 references,issue,150989,issue,27617,medium,issue.comments[0].body,A bit related (on fusing stack and padding - potentially with multiple): #65156 (comment) #27617,https://github.com/pytorch/pytorch/issues/150989,763ba065d25a9d768fa710821450da4c786a5c5154773fe6f1c67abdb821149e references,issue,150989,issue,65156,medium,issue.comments[0].body,A bit related (on fusing stack and padding - potentially with multiple): #65156 (comment) #27617,https://github.com/pytorch/pytorch/issues/150989,559c555dbb13aa64ad077162774996a35446703b12724bd7f1dabc14ef23de8c references,issue,168253,issue,159380,medium,issue.comments[0].body,"I just did more experiments, and it seems like it is a similar issue to #159380.",https://github.com/pytorch/pytorch/issues/168253,352357f7d26d0dded80ea5b33f744dac6b886f0c95396008e310682ec6cb7ff8 references,issue,168288,pr,166876,medium,issue.comments[0].body,"Thanks for raising this Rob, I'll try to land #166876 next week.",https://github.com/pytorch/pytorch/issues/168288,311365a4520fa1d74808e264c31ab4015fa7dd6d30f3a13ad3649742d6f6fe8d references,issue,161302,issue,41081,medium,issue.comments[0].body,ey are already implemented by the same _BatchNorm class) and also providing the SyncBatchNorm functionality as a constructor option: #89116 #41081 (comment) #41243,https://github.com/pytorch/pytorch/issues/161302,5e67df000831ac93fc94a4024dbff5bc2fb1f1c45ac0c5321f9058dcbd314949 references,issue,161302,issue,41243,medium,issue.comments[0].body,plemented by the same _BatchNorm class) and also providing the SyncBatchNorm functionality as a constructor option: #89116 #41081 (comment) #41243,https://github.com/pytorch/pytorch/issues/161302,c30c76d0acd1ac0998b9088e2c115a80cf3000d192e8cce317f45d9d3541b60a references,issue,161302,issue,89116,medium,issue.comments[0].body,ons (they are already implemented by the same _BatchNorm class) and also providing the SyncBatchNorm functionality as a constructor option: #89116 #41081 (comment) #41243,https://github.com/pytorch/pytorch/issues/161302,b383b0eda54aa8fb0a892c997ce6bef0b46660e0edb4086f09a5bbdebaa98b8e references,issue,166004,issue,102963,medium,issue.comments[0].body,"m the function's parameter check (possibly n*lda*batchSize>INT32_MAX, or lda) at /tmp/build/80754af9/python_1627392990942/work/Objects/call.c:433 --Type for more, q to quit, c to continue without paging-- #88 0x0000555555717155 in call_function (kwnames=0x0, oparg=, pp_stack=) at /tmp/build/80754af9/python_1...",https://github.com/pytorch/pytorch/issues/66247,9d68e1421088d76ef5f484d754039726bee149069c0a6724782b8ab3e537492f references,issue,71465,issue,51455,medium,issue.comments[0].body,Related nn.LayerNorm docs issue: #51455 (it would be nice to have a super-compact reference pytorch impl of layernorm in docs for it's much clearer on what's aggregated and even a,https://github.com/pytorch/pytorch/issues/71465,9d007e995cccffeeb01a0a1f1acbb788964ed4cf86c74d8d946baa903e9dc7a0 references,issue,61417,issue,18904,medium,issue.body,64008 Performance #59545 #62164 #51509 #51450 #40570 #31837 #16617 #57610 #55366 #63870 #63869 Feature Requests #61474 #35641 #32998 #32097 #18904 Allow specifying a range for dimensions to reduce over #54766 #32288 Miscellaneous #56764 #41793 #29758 #27312 #25775 cc @mruberry...,https://github.com/pytorch/pytorch/issues/61417,441ef9dd9d764e30935eb6f5bbaaa9f292be284b97cfa8465b0d86f464d3d787 references,issue,61417,issue,19149,medium,issue.body,cing over multiple dimensions. Python Array API Standard #44409 #61486 #61490 #61582 #61492 #50342 #35529 NumPy compatibility #50010 #29137 #19149 Scalar and empty tensors #46999 #28380 Variants #24041 #32570 Type Promotion #45833 #56132 Testing Reduction operators have additi...,https://github.com/pytorch/pytorch/issues/61417,2fc3de057899fe6811da9b6d0b65d34e35debccc1b86c1c200d4b44f6dd4a688 references,issue,61417,issue,23159,medium,issue.body,rectness by providing a reference implementation such as a NumPy equivalent operator. #49746 #59550 #59415 #53704 Bugs #50382 #46225 #28993 #23159 #61523 #61656 #48573 #64008 Performance #59545 #62164 #51509 #51450 #40570 #31837 #16617 #57610 #55366 #63870 #63869 Feature Reque...,https://github.com/pytorch/pytorch/issues/61417,cc2edde4868ad356bed6477789dcfcd9e19fce061e6c80aa1fce61b3b1ac1145 references,issue,61417,issue,24041,medium,issue.body,d #44409 #61486 #61490 #61582 #61492 #50342 #35529 NumPy compatibility #50010 #29137 #19149 Scalar and empty tensors #46999 #28380 Variants #24041 #32570 Type Promotion #45833 #56132 Testing Reduction operators have additional structure than other operators such as dim and kee...,https://github.com/pytorch/pytorch/issues/61417,fda80b43cb57a2d9df5019a6ae40b1a2df30ca4bb62abf20fdc3e355cb9b1722 references,issue,61417,issue,27312,medium,issue.body,#61474 #35641 #32998 #32097 #18904 Allow specifying a range for dimensions to reduce over #54766 #32288 Miscellaneous #56764 #41793 #29758 #27312 #25775 cc @mruberry @rgommers @heitorschueroff @gchanan,https://github.com/pytorch/pytorch/issues/61417,17ac99066154c750ff726d2399af88432c99fdb3f3f33649c46130e82e724a1e references,issue,61417,issue,29137,medium,issue.body,rt reducing over multiple dimensions. Python Array API Standard #44409 #61486 #61490 #61582 #61492 #50342 #35529 NumPy compatibility #50010 #29137 #19149 Scalar and empty tensors #46999 #28380 Variants #24041 #32570 Type Promotion #45833 #56132 Testing Reduction operators have...,https://github.com/pytorch/pytorch/issues/61417,6f906a298f72153d12f764dfc0592d1adb9627e249e9f8aeb1c3c343831280ae references,issue,61417,issue,29758,medium,issue.body,equests #61474 #35641 #32998 #32097 #18904 Allow specifying a range for dimensions to reduce over #54766 #32288 Miscellaneous #56764 #41793 #29758 #27312 #25775 cc @mruberry @rgommers @heitorschueroff @gchanan,https://github.com/pytorch/pytorch/issues/61417,61f0a3c36bda1d7c7cdaf250865298587c9d9c3e6c95fee811a1c9e5d2e21f89 references,issue,61417,issue,31837,medium,issue.body,or. #49746 #59550 #59415 #53704 Bugs #50382 #46225 #28993 #23159 #61523 #61656 #48573 #64008 Performance #59545 #62164 #51509 #51450 #40570 #31837 #16617 #57610 #55366 #63870 #63869 Feature Requests #61474 #35641 #32998 #32097 #18904 Allow specifying a range for dimensions to...,https://github.com/pytorch/pytorch/issues/61417,d109d9ccc91c54a67c7c956ea17ee44cab6b9facf8bcd8f4c56587ccb5d1134a references,issue,61417,issue,32097,medium,issue.body,48573 #64008 Performance #59545 #62164 #51509 #51450 #40570 #31837 #16617 #57610 #55366 #63870 #63869 Feature Requests #61474 #35641 #32998 #32097 #18904 Allow specifying a range for dimensions to reduce over #54766 #32288 Miscellaneous #56764 #41793 #29758 #27312 #25775 cc @m...,https://github.com/pytorch/pytorch/issues/61417,7ab1100d5f578c70cf7796199267b022310f3e497fb7531dfd6bccf5565413c4 references,issue,61417,issue,32288,medium,issue.body,7610 #55366 #63870 #63869 Feature Requests #61474 #35641 #32998 #32097 #18904 Allow specifying a range for dimensions to reduce over #54766 #32288 Miscellaneous #56764 #41793 #29758 #27312 #25775 cc @mruberry @rgommers @heitorschueroff @gchanan,https://github.com/pytorch/pytorch/issues/61417,f942f0aed90eac627f7ef50eba8fd2ffc7900f6b7dbf5d63885c4a8d74e2749b references,issue,61417,issue,32570,medium,issue.body,9 #61486 #61490 #61582 #61492 #50342 #35529 NumPy compatibility #50010 #29137 #19149 Scalar and empty tensors #46999 #28380 Variants #24041 #32570 Type Promotion #45833 #56132 Testing Reduction operators have additional structure than other operators such as dim and keepdim pa...,https://github.com/pytorch/pytorch/issues/61417,69ffdbe6ad1c0277f9dd4814c67ef7881a60653cb2594ad3d0b18274d3621ba5 references,issue,61417,issue,32998,medium,issue.body,61656 #48573 #64008 Performance #59545 #62164 #51509 #51450 #40570 #31837 #16617 #57610 #55366 #63870 #63869 Feature Requests #61474 #35641 #32998 #32097 #18904 Allow specifying a range for dimensions to reduce over #54766 #32288 Miscellaneous #56764 #41793 #29758 #27312 #2577...,https://github.com/pytorch/pytorch/issues/61417,92f87db7a2015f35cd8437dc8abae8ee46ad56d2e8654319bfa8ef832f903dfd references,issue,61417,issue,35529,medium,issue.body,"here applicable, they should support reducing over multiple dimensions. Python Array API Standard #44409 #61486 #61490 #61582 #61492 #50342 #35529 NumPy compatibility #50010 #29137 #19149 Scalar and empty tensors #46999 #28380 Variants #24041 #32570 Type Promotion #45833 #5613...",https://github.com/pytorch/pytorch/issues/61417,cbd5bea744650330657c3f11ef04415a46a3698fe863dad084463fe759cd674a references,issue,61417,issue,35641,medium,issue.body,61523 #61656 #48573 #64008 Performance #59545 #62164 #51509 #51450 #40570 #31837 #16617 #57610 #55366 #63870 #63869 Feature Requests #61474 #35641 #32998 #32097 #18904 Allow specifying a range for dimensions to reduce over #54766 #32288 Miscellaneous #56764 #41793 #29758 #2731...,https://github.com/pytorch/pytorch/issues/61417,43e1b9a6713c72cc3ff3e23c463ee0b6a3b8d9fe237b317765eb48486cae6d73 references,issue,61417,issue,40570,medium,issue.body,operator. #49746 #59550 #59415 #53704 Bugs #50382 #46225 #28993 #23159 #61523 #61656 #48573 #64008 Performance #59545 #62164 #51509 #51450 #40570 #31837 #16617 #57610 #55366 #63870 #63869 Feature Requests #61474 #35641 #32998 #32097 #18904 Allow specifying a range for dimensio...,https://github.com/pytorch/pytorch/issues/61417,cb04af7340b28ea1405f1029d3cc17aef95fd358f38743038a0d0eaa3ab7dfb9 references,issue,61417,issue,41793,medium,issue.body,ature Requests #61474 #35641 #32998 #32097 #18904 Allow specifying a range for dimensions to reduce over #54766 #32288 Miscellaneous #56764 #41793 #29758 #27312 #25775 cc @mruberry @rgommers @heitorschueroff @gchanan,https://github.com/pytorch/pytorch/issues/61417,bc92d77784beca1d9875a0fcdfca86d4fd37e04ab78b4645f51ff3ca20ab853a references,issue,61417,issue,45833,medium,issue.body,#61492 #50342 #35529 NumPy compatibility #50010 #29137 #19149 Scalar and empty tensors #46999 #28380 Variants #24041 #32570 Type Promotion #45833 #56132 Testing Reduction operators have additional structure than other operators such as dim and keepdim parameters which can be e...,https://github.com/pytorch/pytorch/issues/61417,089616c28d950b45c7abd58b30449d3b2bd39505d4b5ee5d01e605e54a5ff060 references,issue,61417,issue,46225,medium,issue.body,esting for correctness by providing a reference implementation such as a NumPy equivalent operator. #49746 #59550 #59415 #53704 Bugs #50382 #46225 #28993 #23159 #61523 #61656 #48573 #64008 Performance #59545 #62164 #51509 #51450 #40570 #31837 #16617 #57610 #55366 #63870 #63869...,https://github.com/pytorch/pytorch/issues/61417,cc78acb9efffbabaf47b831233d1582e702f03b9ee2c08e2ec0cbe0c4624758d references,issue,61417,issue,46999,medium,issue.body,ython Array API Standard #44409 #61486 #61490 #61582 #61492 #50342 #35529 NumPy compatibility #50010 #29137 #19149 Scalar and empty tensors #46999 #28380 Variants #24041 #32570 Type Promotion #45833 #56132 Testing Reduction operators have additional structure than other operat...,https://github.com/pytorch/pytorch/issues/61417,4bf133342ebda80d41c246390ce9ef4e877fef6fe718b869410a70a4732f72bc references,issue,61417,issue,50010,medium,issue.body,d support reducing over multiple dimensions. Python Array API Standard #44409 #61486 #61490 #61582 #61492 #50342 #35529 NumPy compatibility #50010 #29137 #19149 Scalar and empty tensors #46999 #28380 Variants #24041 #32570 Type Promotion #45833 #56132 Testing Reduction operato...,https://github.com/pytorch/pytorch/issues/61417,3c038b8adc4bf97830f3ac1ae780d1ef5d80c391d1022f04fdb741fab0a41e2d references,issue,61417,issue,50342,medium,issue.body,"ard). Where applicable, they should support reducing over multiple dimensions. Python Array API Standard #44409 #61486 #61490 #61582 #61492 #50342 #35529 NumPy compatibility #50010 #29137 #19149 Scalar and empty tensors #46999 #28380 Variants #24041 #32570 Type Promotion #4583...",https://github.com/pytorch/pytorch/issues/61417,c788f4b4728f98211240075667335c4aa585fa37691e9918ce4ff5afa3359183 references,issue,61417,issue,50382,medium,issue.body,even testing for correctness by providing a reference implementation such as a NumPy equivalent operator. #49746 #59550 #59415 #53704 Bugs #50382 #46225 #28993 #23159 #61523 #61656 #48573 #64008 Performance #59545 #62164 #51509 #51450 #40570 #31837 #16617 #57610 #55366 #63870...,https://github.com/pytorch/pytorch/issues/61417,191376b583fa527e156e7c2cf038dfb224862282b97928d53712134e635ba927 references,issue,61417,issue,51450,medium,issue.body,ivalent operator. #49746 #59550 #59415 #53704 Bugs #50382 #46225 #28993 #23159 #61523 #61656 #48573 #64008 Performance #59545 #62164 #51509 #51450 #40570 #31837 #16617 #57610 #55366 #63870 #63869 Feature Requests #61474 #35641 #32998 #32097 #18904 Allow specifying a range for...,https://github.com/pytorch/pytorch/issues/61417,379755e5b1da379f39244e56355d2a1a7a36026d533129d77cc2ac1cde3f5c51 references,issue,61417,issue,51509,medium,issue.body,mPy equivalent operator. #49746 #59550 #59415 #53704 Bugs #50382 #46225 #28993 #23159 #61523 #61656 #48573 #64008 Performance #59545 #62164 #51509 #51450 #40570 #31837 #16617 #57610 #55366 #63870 #63869 Feature Requests #61474 #35641 #32998 #32097 #18904 Allow specifying a ran...,https://github.com/pytorch/pytorch/issues/61417,3b1dda70e59c1c6b04e37d08a5be67cb7a6877e88e3299ac8910ef916ea21aab references,issue,61417,issue,53704,medium,issue.body,"tensors, and even testing for correctness by providing a reference implementation such as a NumPy equivalent operator. #49746 #59550 #59415 #53704 Bugs #50382 #46225 #28993 #23159 #61523 #61656 #48573 #64008 Performance #59545 #62164 #51509 #51450 #40570 #31837 #16617 #57610 #...",https://github.com/pytorch/pytorch/issues/61417,0fac792a29eeebba62844a25cbe8c992d6ac3b67d971ad53a0c78735af5aee7c references,issue,61417,issue,54766,medium,issue.body,6617 #57610 #55366 #63870 #63869 Feature Requests #61474 #35641 #32998 #32097 #18904 Allow specifying a range for dimensions to reduce over #54766 #32288 Miscellaneous #56764 #41793 #29758 #27312 #25775 cc @mruberry @rgommers @heitorschueroff @gchanan,https://github.com/pytorch/pytorch/issues/61417,c30914f7fda634d273c27f7c50e0b1782a01926d15f7ac719abc551846a324ee references,issue,61417,issue,55366,medium,issue.body,415 #53704 Bugs #50382 #46225 #28993 #23159 #61523 #61656 #48573 #64008 Performance #59545 #62164 #51509 #51450 #40570 #31837 #16617 #57610 #55366 #63870 #63869 Feature Requests #61474 #35641 #32998 #32097 #18904 Allow specifying a range for dimensions to reduce over #54766 #3...,https://github.com/pytorch/pytorch/issues/61417,99dbd4f451d24c9ed6ebed6bdba1119dc769c6916a78cdac5e94227daa126534 references,issue,61417,issue,56764,medium,issue.body,3869 Feature Requests #61474 #35641 #32998 #32097 #18904 Allow specifying a range for dimensions to reduce over #54766 #32288 Miscellaneous #56764 #41793 #29758 #27312 #25775 cc @mruberry @rgommers @heitorschueroff @gchanan,https://github.com/pytorch/pytorch/issues/61417,301488db30f0d61ef0c3525c94a16ed6773340292565aacb059b79e92798c501 references,issue,61417,issue,57610,medium,issue.body,550 #59415 #53704 Bugs #50382 #46225 #28993 #23159 #61523 #61656 #48573 #64008 Performance #59545 #62164 #51509 #51450 #40570 #31837 #16617 #57610 #55366 #63870 #63869 Feature Requests #61474 #35641 #32998 #32097 #18904 Allow specifying a range for dimensions to reduce over #5...,https://github.com/pytorch/pytorch/issues/61417,52cbe39a488cefa4da984e23c7a039cc93ec7cd41a450d01c0bfc5496844bfa1 references,issue,61417,issue,59545,medium,issue.body,n such as a NumPy equivalent operator. #49746 #59550 #59415 #53704 Bugs #50382 #46225 #28993 #23159 #61523 #61656 #48573 #64008 Performance #59545 #62164 #51509 #51450 #40570 #31837 #16617 #57610 #55366 #63870 #63869 Feature Requests #61474 #35641 #32998 #32097 #18904 Allow sp...,https://github.com/pytorch/pytorch/issues/61417,462d0b4ba700cf560b15ad94c8de28b6742b3cfbabb3e76cb9a15edb0d16ac93 references,issue,61417,issue,61474,medium,issue.body,23159 #61523 #61656 #48573 #64008 Performance #59545 #62164 #51509 #51450 #40570 #31837 #16617 #57610 #55366 #63870 #63869 Feature Requests #61474 #35641 #32998 #32097 #18904 Allow specifying a range for dimensions to reduce over #54766 #32288 Miscellaneous #56764 #41793 #2975...,https://github.com/pytorch/pytorch/issues/61417,3ccd000df9c0624677e3bfcff8588c250d28f0ddd0c34e87bec1e4e78049c93b references,issue,61417,issue,61486,medium,issue.body,"o the Python Array API Standard). Where applicable, they should support reducing over multiple dimensions. Python Array API Standard #44409 #61486 #61490 #61582 #61492 #50342 #35529 NumPy compatibility #50010 #29137 #19149 Scalar and empty tensors #46999 #28380 Variants #24041...",https://github.com/pytorch/pytorch/issues/61417,51c7c14b39daf138673d75d8f52dd39e837e4e2da28352c0dbaf5a2946e6182b references,issue,61417,issue,61490,medium,issue.body,"ython Array API Standard). Where applicable, they should support reducing over multiple dimensions. Python Array API Standard #44409 #61486 #61490 #61582 #61492 #50342 #35529 NumPy compatibility #50010 #29137 #19149 Scalar and empty tensors #46999 #28380 Variants #24041 #32570...",https://github.com/pytorch/pytorch/issues/61417,f34151ad5a94ab301579388c4fa31be74426770cb27f8bc800578e5562e1477a references,issue,61417,issue,61523,medium,issue.body,s by providing a reference implementation such as a NumPy equivalent operator. #49746 #59550 #59415 #53704 Bugs #50382 #46225 #28993 #23159 #61523 #61656 #48573 #64008 Performance #59545 #62164 #51509 #51450 #40570 #31837 #16617 #57610 #55366 #63870 #63869 Feature Requests #61...,https://github.com/pytorch/pytorch/issues/61417,105330f99b87d1014548d8e4d58278b82c99b3b72a725389267bf3af4bfa271c references,issue,61417,issue,61582,medium,issue.body,"rray API Standard). Where applicable, they should support reducing over multiple dimensions. Python Array API Standard #44409 #61486 #61490 #61582 #61492 #50342 #35529 NumPy compatibility #50010 #29137 #19149 Scalar and empty tensors #46999 #28380 Variants #24041 #32570 Type P...",https://github.com/pytorch/pytorch/issues/61417,b96bb432c6705407642b97ab3cb98d466073e0b1e30c76ba3a796a3695991830 references,issue,61417,issue,63870,medium,issue.body,704 Bugs #50382 #46225 #28993 #23159 #61523 #61656 #48573 #64008 Performance #59545 #62164 #51509 #51450 #40570 #31837 #16617 #57610 #55366 #63870 #63869 Feature Requests #61474 #35641 #32998 #32097 #18904 Allow specifying a range for dimensions to reduce over #54766 #32288 Mi...,https://github.com/pytorch/pytorch/issues/61417,0daf570fc6becb8ef9e36731963826f73314169065563117816ee35e276555df references,issue,20367,issue,5388,medium,issue.body,"specific op regressions, that almost certainly overlap, but which we should track separately to make sure we cover all the cases. To start: #5388 #16717 #2560 cc @ezyang @gchanan @zou3519 @VitalyFedyunin @ngimel @mruberry",https://github.com/pytorch/pytorch/issues/20367,57fb82a21c7f2c02775676b001b12bdcc644d2ce2aab8ba3c48b3354fdd9bd65 references,issue,20367,issue,16717,medium,issue.body,"ic op regressions, that almost certainly overlap, but which we should track separately to make sure we cover all the cases. To start: #5388 #16717 #2560 cc @ezyang @gchanan @zou3519 @VitalyFedyunin @ngimel @mruberry",https://github.com/pytorch/pytorch/issues/20367,6c8bb43e11180b3473a7d469999ffcc248d771a328e4f740597b3f55bc673e1c references,issue,18631,issue,13716,medium,issue.body,"ite. I'm using 1.1.0a0+65d6f10_2_ged1fa68 ,cuda10, driver:410.78,Titan Xp btw, there is an issue of depthwise convolution being slow in CPU #13716",https://github.com/pytorch/pytorch/issues/18631,c7e5dfce2ebf24052d7905cc069416f26a134842a19d741df5e1c031f8a290e7 references,issue,68070,issue,25478,medium,issue.comments[0].body,What was the recommended resolution in triage meeting? options() is a bit wonky due to #25478,https://github.com/pytorch/pytorch/issues/68070,f716e684752642701a046390678057653c35404758decb07edf0845d4a29ab9c references,issue,64412,issue,26165,medium,issue.comments[0].body,This was also reported in #26165 and #24237,https://github.com/pytorch/pytorch/issues/64412,ab9a15d1f8b250e3cc78e458f82d249c384ecf8a6ccf53cfd054a24808da18e5 competes with,issue,68174,issue,50012,medium,issue.body,"different from NumPy’s split because it asks what size the resulting chunks should be, instead of how many there should be (see this issue (#50012) for details) nonzero → argwhere PyTorch’s nonzero is different from NumPy’s nonzero but identical to NumPy’s argwhere max → amax,...",https://github.com/pytorch/pytorch/issues/68174,1000de02c125c7922981280433a349317a5d98dd4aca93f03827ef109a03afdd references,issue,161713,issue,160322,medium,issue.body,mprove types exported by PyTorch so people who install torch can better navigate and enable type checking in their code. This follows up on #160322 to help define how we'll track external type coverage improvements and rules for how types are added externally. (Note this is no...,https://github.com/pytorch/pytorch/issues/161713,e7d70574005afee3589f0cd3a09db42201e292ec1317de2d5bd86022a8764d8d references,issue,162820,issue,162178,medium,issue.body,🐛 Describe the bug Tracked in umbrella #162178 Job link: https://github.com/pytorch/pytorch/actions/runs/17660052730/job/50193312091 Failure message: 2025-09-12T05:47:07.8805304Z expect_,https://github.com/pytorch/pytorch/issues/162820,8d92380b38951c0f15d3f7d569f38dabc120264d936dd4142afc96d4c088190e references,issue,162918,issue,61819,medium,issue.comments[0].body,Related: #61819 #103580,https://github.com/pytorch/pytorch/issues/162918,4fb35611c7c0234ed0c578c489bac82a57b5e4b36116ca76b9b5ca614ba3bbf0 references,issue,162918,issue,103580,medium,issue.comments[0].body,Related: #61819 #103580,https://github.com/pytorch/pytorch/issues/162918,5bceabf0d9f4b9abccf154fe4e28ed325299c4d571a80653baeba9a28c9bf727 references,issue,162178,issue,162748,medium,issue.body,"ch, but not merged yet] python test/distributed/test_distributed_spawn.py TestDistBackendWithSpawn.test_3_level_hierarchical_model_averager #162748 [New failure after 09/11/2025 merge main] python test/distributed/tensor/test_attention.py RingFlexAttentionTest.test_ring_flex_a...",https://github.com/pytorch/pytorch/issues/162178,7409cc5fb1bc29607044c2c9256072a1643dcbe58d2d10dd09f94d0e73ace58a references,issue,162178,issue,162820,medium,issue.body,8 [New failure after 09/11/2025 merge main] python test/distributed/tensor/test_attention.py RingFlexAttentionTest.test_ring_flex_attention #162820 #162871 #162897 #162917 #163462 TOT cc @ptrblck @msaroufim @eqy @jerryzh168 @tinglvv @drisspg @atalman @malfet,https://github.com/pytorch/pytorch/issues/162178,bf04550ffce8761a62eedbad5546551aeb42426849d3dfe91ab936b09f1f434b references,issue,162178,issue,162871,medium,issue.body,ailure after 09/11/2025 merge main] python test/distributed/tensor/test_attention.py RingFlexAttentionTest.test_ring_flex_attention #162820 #162871 #162897 #162917 #163462 TOT cc @ptrblck @msaroufim @eqy @jerryzh168 @tinglvv @drisspg @atalman @malfet,https://github.com/pytorch/pytorch/issues/162178,d8d274dafdf5066a6ea66998ec97aa4aadaa5ff9f10523bd6504cba4f0baeb47 references,issue,162178,issue,162897,medium,issue.body,fter 09/11/2025 merge main] python test/distributed/tensor/test_attention.py RingFlexAttentionTest.test_ring_flex_attention #162820 #162871 #162897 #162917 #163462 TOT cc @ptrblck @msaroufim @eqy @jerryzh168 @tinglvv @drisspg @atalman @malfet,https://github.com/pytorch/pytorch/issues/162178,d9bb82e11f928d1bd4c1dbdfbf238eada4e0c4e6a1d136e5447492c59fb4dbd3 references,issue,162178,issue,162917,medium,issue.body,11/2025 merge main] python test/distributed/tensor/test_attention.py RingFlexAttentionTest.test_ring_flex_attention #162820 #162871 #162897 #162917 #163462 TOT cc @ptrblck @msaroufim @eqy @jerryzh168 @tinglvv @drisspg @atalman @malfet,https://github.com/pytorch/pytorch/issues/162178,4a270ba155ac5b735946d332487341f7eb85b69c990e7176ed25c87741af9160 references,issue,162178,issue,163429,medium,issue.body,"compose_k_dynamic_False_bfloat16_sizes2, python test/inductor/test_max_autotune.py TestMaxAutotune.test_non_contiguous_input_mm_plus_mm etc #163429 (PRs that fix them: #162187 ) Small and ""Big"" mismatch: python test/test_matmul_cuda.py TestMatmulCudaCUDA.test_cublas_addmm_redu...",https://github.com/pytorch/pytorch/issues/162178,58a7779b486e8f2e6114aaf0647b7596d3de2c2aeb934213e36901ad27079241 references,issue,163142,issue,152032,medium,issue.body,For some more context you can check out #152032 torch.cuda._compile_kernel is an experimental API that leverages NVRTC under the hood for blazing fast compilation times. @gau-nernst has a,https://github.com/pytorch/pytorch/issues/163142,0f840dca6b1b5a3d3c28f38f9da1c776f6178b9474e045b28d206abb8f704261 references,issue,146828,issue,32351,medium,issue.comments[1].body,"Could be connected to #32351, where the same error is caused by an explicit __getstate__ returning freshly created tensor, which seems to be garbage collected before be",https://github.com/pytorch/pytorch/issues/146828,9d13cf426afe40b2936cf7951f63c416e5739058a2eabb7c145a139c7ab80abe references,issue,103580,issue,61819,medium,issue.comments[0].body,Maybe related: #61819,https://github.com/pytorch/pytorch/issues/103580,f5b4f7a01a1c31c185fdd28789325d8c262ae503ff7e5486e0fd78b811e02084 references,issue,69532,issue,69531,medium,issue.body,"ility. The current implementation of SVD makes backward unstable when the input has repeated and/or zero singular values. See, for example: #69531. This kills the possibility to use SVD.backward on low-rank inputs which sure have (repeated) zero singular values. torch.linalg.s...",https://github.com/pytorch/pytorch/issues/69532,78c426e7cb2d100e0172caad417a8f5b767da3d753e37e69cac393d623c4a344 references,issue,106667,issue,34651,medium,issue.body,ries of image/audio encoding/decoding Trying to distill from above concrete ideas on improving FFI story of PyTorch: Fixing interop issues: #34651 #69491 CuPy-like features: API like RawKernel RawModule nvrtc-compilation (supposed to be faster than dealing with full nvcc / wor...,https://github.com/pytorch/pytorch/issues/106667,8023d638d14976f8d20c7533aaf00968faf97813c698b99cb95aef602544a094 references,issue,106667,issue,69491,medium,issue.body,image/audio encoding/decoding Trying to distill from above concrete ideas on improving FFI story of PyTorch: Fixing interop issues: #34651 #69491 CuPy-like features: API like RawKernel RawModule nvrtc-compilation (supposed to be faster than dealing with full nvcc / working wit...,https://github.com/pytorch/pytorch/issues/106667,ba7a0afc4be31ce30fed55f7a513ac879eb9004be59c2aea4f48224efbe6e0dc references,issue,83163,issue,36524,medium,issue.comments[0].body,"Related: #36524 #64038 As discussed there, this _stacklevel is a remnant of migrating to explicit-only dim and to control warnings about not setting dim an",https://github.com/pytorch/pytorch/issues/83163,18fa004746e1ac8e2ffcdf49bf31412d84b06f0d0e707fb063401878b10b0077 references,issue,83163,issue,64038,medium,issue.comments[0].body,"Related: #36524 #64038 As discussed there, this _stacklevel is a remnant of migrating to explicit-only dim and to control warnings about not setting dim and the d",https://github.com/pytorch/pytorch/issues/83163,c94f2cc2f2def840993082279f96e560e5597d5541c694129aaec994cf7331aa references,issue,83163,issue,36524,medium,issue.comments[1].body,"Related: #36524. As discussed there, this _stacklevel is a remnant of migrating to explicit-only dim and to control warnings about not setting dim and the",https://github.com/pytorch/pytorch/issues/83163,16c9d5c8f4cbcfa2b0251840d57e4f9a563e64a57787cc9e809e4628f8204d20 references,issue,124505,issue,106614,medium,issue.body,"🚀 The feature, motivation and pitch Originally discussed in: #106614 As I understood, any torch.compile call with passed options will modify global compiler options (and thus require torch._dynamo.reset() or",https://github.com/pytorch/pytorch/issues/124505,097826b054f217d70a852e472de5b092afb6f828d9713da50bc99454d61a5171 references,issue,142358,issue,124505,medium,issue.comments[0].body,I wonder if it's related to global-ness of compiler options... (#124505),https://github.com/pytorch/pytorch/issues/142358,95cde47a03d27d5b41d50fd3b38629ea3bba46e581b90aed4b33c017ec2a74d3 references,issue,156308,issue,124505,medium,issue.comments[0].body,Related idea and discussion: #124505,https://github.com/pytorch/pytorch/issues/156308,1d50d28b0bbc5b8683b5e7998db012226af4b2685fa5775c8fc90e96b9d7582f references,issue,90560,issue,49016,medium,issue.comments[0].body,implemented in https://github.com/mlomnitz/DiffJPEG) as will make these things more exposed to public / foster exploration Related on DCT: #49016 also lossless integer DCT would be nice,https://github.com/pytorch/pytorch/issues/90560,9e6f64129c5c5172edaa64a31cf96a09a07c6c069924dcd2ed0224a850a18905 references,issue,162476,issue,116254,medium,issue.body,"ibe the bug A heap-buffer-overflow can be triggered in torch.quantized_max_pool2d via the Python API, similar to the issue reported in C++ (#116254). When converting it's C++ snippet to Python, the same input parameters cannot trigger the overflow: import torch print(torch.__v...",https://github.com/pytorch/pytorch/issues/162476,0a9953422758f635e2a493556c57b69830b3142b1ce9a3f237e1049e97d2c3a0 references,issue,162451,issue,162365,medium,issue.comments[0].body,"Linked to other Windows inductor regression UTs: #162365 , #162366",https://github.com/pytorch/pytorch/issues/162451,bce1d11cfa75c242f5034070a4143becac104b1b4ff94f39d5b3d3c843cdc5f5 references,issue,160697,issue,159843,medium,issue.body,None [rank0]: torch._dynamo.exc.BackendCompilerFailed: backend='inductor' raised: [rank0]: AssertionError: This is also somewhat related to #159843 Versions Details PyTorch version: 2.9.0a0+git02e09d1 Is debug build: False CUDA used to build PyTorch: 12.6 ROCM used to build Py...,https://github.com/pytorch/pytorch/issues/160697,a1a155f4604272481da34904ab9fd4e9d1570bf623cb64be99972d66b53cabda references,issue,69687,issue,75586,medium,issue.body,"v addr baddbmm bmm cat #75590 stack #75590 cross dot vdot index_add_ index_copy_ index_fill_.int_Tensor index_put_ _index_put_impl_ max min #75586 mul.Tensor sub.Tensor remainder.Tensor To be verified: pow.Tensor_Tensor, linalg_solve should special case (even zero tensor can’t...",https://github.com/pytorch/pytorch/issues/69687,2ab0d32d861a713af7e78aad9f0fa842b6426110268318ff314fc32cf4455e07 references,issue,158793,issue,159013,medium,issue.body,ts and Proposals We want to use RFC to track all RFCs since each topic is an orthogonal one so discussions will be more focused. Topic one: #159013 (Mostly wanted and asked) Topic two: #159014 Topic three: #159015. Topic four: #159017 Topic five: #159018 Topic six: #159019 Alt...,https://github.com/pytorch/pytorch/issues/158793,bd653ef77e05f7be06fe8ac8716f090f442d17e93cfdb50434419d7d781844c5 references,issue,158793,issue,159014,medium,issue.body,all RFCs since each topic is an orthogonal one so discussions will be more focused. Topic one: #159013 (Mostly wanted and asked) Topic two: #159014 Topic three: #159015. Topic four: #159017 Topic five: #159018 Topic six: #159019 Alternatives No response Additional context No r...,https://github.com/pytorch/pytorch/issues/158793,ad001de66602c64ea993094f1b77f4c0cecbffc533f03a2dbeb2b71249ab8c3c references,issue,158793,issue,159015,medium,issue.body,opic is an orthogonal one so discussions will be more focused. Topic one: #159013 (Mostly wanted and asked) Topic two: #159014 Topic three: #159015. Topic four: #159017 Topic five: #159018 Topic six: #159019 Alternatives No response Additional context No response cc @H-Huang @...,https://github.com/pytorch/pytorch/issues/158793,2686a9a3cee084a024056089dbd893137241e4a466f28c3de3b0d2242c1bc077 references,issue,158793,issue,159017,medium,issue.body,one so discussions will be more focused. Topic one: #159013 (Mostly wanted and asked) Topic two: #159014 Topic three: #159015. Topic four: #159017 Topic five: #159018 Topic six: #159019 Alternatives No response Additional context No response cc @H-Huang @awgu @wanchaol @fegin...,https://github.com/pytorch/pytorch/issues/158793,84c0a372b5be08d448ec22252b4df9e81270fe61074f69059dfec11281083146 references,issue,158793,issue,159018,medium,issue.body,will be more focused. Topic one: #159013 (Mostly wanted and asked) Topic two: #159014 Topic three: #159015. Topic four: #159017 Topic five: #159018 Topic six: #159019 Alternatives No response Additional context No response cc @H-Huang @awgu @wanchaol @fegin @wz337 @wconstab @d...,https://github.com/pytorch/pytorch/issues/158793,e85b50a9fe67a7691e5edce3e24fd5250077705631c1dc1d5ed0e65859ffedcc references,issue,158793,issue,159019,medium,issue.body,d. Topic one: #159013 (Mostly wanted and asked) Topic two: #159014 Topic three: #159015. Topic four: #159017 Topic five: #159018 Topic six: #159019 Alternatives No response Additional context No response cc @H-Huang @awgu @wanchaol @fegin @wz337 @wconstab @d4l3k @pragupta,https://github.com/pytorch/pytorch/issues/158793,2d3dc0ea89d0f689039ed830c8c89af2a10d9d0cf6a74edd7643126791d98489 references,issue,159015,issue,153302,medium,issue.body,"roup when possible. For example, users create a new device mesh for every fully_shard but no new PG needs to be created. This is tracked in #153302. Alternatives No response Additional context No response cc @H-Huang @awgu @wanchaol @fegin @wz337 @wconstab @d4l3k @pragupta",https://github.com/pytorch/pytorch/issues/159015,eb2f065315848391d8220f05170d65804354820a94a7045a57ddb6f7471ef857 references,issue,161671,issue,128554,medium,issue.body,weight_obs' # Failed to load: _ModuleStackTracer.__init__() missing 1 required positional argument: 'scope_root' Potentially related issue: #128554 Versions PyTorch version: 2.7.1+cpu Is debug build: False CUDA used to build PyTorch: None ROCM used to build PyTorch: N/A OS: De...,https://github.com/pytorch/pytorch/issues/161671,9cc8fe0d54ece639c295c62f4f966d55cbf86e0ac703cf4b14573c53f4b2be78 references,issue,114299,issue,120003,medium,issue.body,"nt), then they should use fully_shard directly. We are preferring to not do this. Progress Tracker The main code has landed in PyTorch. See #120003 for the remaining gaps, which may take more time to fill. Some notes: The current non-test code sits at ~3k lines, whereas the ex...",https://github.com/pytorch/pytorch/issues/114299,d8001f9322243c2abc85fa609a47319ee4dfd607767b97a4ee88011aab0ee46b references,issue,159906,issue,160420,medium,issue.comments[0].body,ithm than querying cudaEvents. That work was started in #146924 Allocations with particular NUMA affinity cannot be done with cudaHostAlloc #160420 It's pretty clear to me that there is a need for the interface to change to allow for more than one caching host allocator within...,https://github.com/pytorch/pytorch/issues/159906,ceb04b11c88ca2e42567416d70871ca477733b044062bf385843e7fd87eaad13 references,issue,60466,issue,45009,medium,issue.comments[0].body,Maybe related: #45009,https://github.com/pytorch/pytorch/issues/60466,1bb89f853afb800fd62c2e5f850d443dcd27eabe6c58ab94412e8c28db05b95b references,issue,89116,issue,41081,medium,issue.body,"BatchNorm would be the minimally invasive approach. Other issues have also proposed this solution before. Also, I am willing to contribute. #41081 #82464 cc @mrshenli @pritamdamania87 @zhaojuanmao @satgera @rohan-varma @gqchen @aazzolini @osalpekar @jiayisuse @H-Huang @kwen250...",https://github.com/pytorch/pytorch/issues/89116,ad04ea376d70a901984cdb0b72a2264087cc551f67f4eb8e1b2cf6a2bc7fc91b references,issue,89116,issue,82464,medium,issue.body,"rm would be the minimally invasive approach. Other issues have also proposed this solution before. Also, I am willing to contribute. #41081 #82464 cc @mrshenli @pritamdamania87 @zhaojuanmao @satgera @rohan-varma @gqchen @aazzolini @osalpekar @jiayisuse @H-Huang @kwen2501 @awgu...",https://github.com/pytorch/pytorch/issues/89116,0d14ec2d48f380520627d10dac21660493c99f92653da4ed19bcaa924bc78c32 references,issue,89116,issue,26288,medium,issue.comments[0].body,Also related: #41243 #66073 #68648 #26288,https://github.com/pytorch/pytorch/issues/89116,96a13406cdbfd9e82f5023597ce4e6d750153e7a68212bf16fdd92e125dc63a8 references,issue,89116,issue,41243,medium,issue.comments[0].body,Also related: #41243 #66073 #68648 #26288,https://github.com/pytorch/pytorch/issues/89116,c49d945f6aa418df4af35151c5bc27ca039effcedfd880a3bc3e3d8d6994d1de references,issue,89116,issue,66073,medium,issue.comments[0].body,Also related: #41243 #66073 #68648 #26288,https://github.com/pytorch/pytorch/issues/89116,6606e1da17e7be3b78a7747978c358ca02049973e94396f5d1de343936e365f6 references,issue,89116,issue,68648,medium,issue.comments[0].body,Also related: #41243 #66073 #68648 #26288,https://github.com/pytorch/pytorch/issues/89116,a6eb08e8d731dbd282c65607bc2c573ac2134b36c572627af32d42b5d5e7605c references,issue,160508,issue,138800,medium,issue.body,issue if you come across this error otherwise.') This seems to be triggered by a similar set of conditions but the error is different from #138800. Details: Given this little bit of code: from jaxtyping import Float from torch import Tensor def _radial_and_tangential_undistort...,https://github.com/pytorch/pytorch/issues/160508,f451019cafdc941be9df42d085020c69ddd1ac1fa644c40d9c277749b73f71a0 references,issue,159447,issue,152335,medium,issue.comments[1].body,See #152335,https://github.com/pytorch/pytorch/issues/159447,f0e482a6ce9484417c3b38ae5fc964d0e1ce4792d2b0001af9bb766b00391111 references,issue,93859,issue,73479,medium,issue.body,"rating problems with these, so a concise text-based format (typically for small tensors) would be useful! Original context and proposal in: #73479 (comment) Related on general use of utility for base64 torch.save/torch.load (?) format for repro purposes: #93366 (comment) cc @m...",https://github.com/pytorch/pytorch/issues/93859,a340c4dd0a6651e849fcaf674cd0b88b6a6a9bc7484665dccbe700872ae01fcc references,issue,93859,issue,90560,medium,issue.comments[0].body,ed my example of base64<>uint8tensor encoding in PyTorch in https://gist.github.com/vadimkantorov/ea989c75f79961fe46182845b40d5f31 Related: #90560,https://github.com/pytorch/pytorch/issues/93859,df5c521e60cc2a7e897fe587404fcbc37178f88924466c41a6ca78c8356b6fb0 references,issue,159765,issue,50122,medium,issue.comments[0].body,A bit related on gradient mod util: gradient reversal module / autograd.Function and an util replacing infs/nans in gradients: #50122,https://github.com/pytorch/pytorch/issues/159765,afa61b41fd84ca4edd02c67eea5340e8fa358c395e1cee0aca63ffc90d1fc338 references,issue,159852,issue,159353,medium,issue.comments[1].body,"Hi, I think this is duplicated with #159353 (comment). The args is already unrolled. So could just do: import torch from torch import nn class MyModel(torch.nn.Module): def __init__(s",https://github.com/pytorch/pytorch/issues/159852,194b6bfe0f5afa957100024b3c4357b9293a34d8af1a6db026a7df41f8be9892 references,issue,121219,issue,64345,medium,issue.body,"mes. However, there seems to be a discrepancy between Python 3.10 and Python 3.11 and how imports are handled. I believe this is related to #64345 and #113564. The following snippet produces good trace files with Python 3.10, but occasionally produces corrupted output on Pytho...",https://github.com/pytorch/pytorch/issues/121219,b4f9b26178cb616742202b4904f804a98b88acbec6dec942e57d9fccf5f5166d references,issue,159169,issue,134191,medium,issue.comments[1].body,Some prior discussion #134191,https://github.com/pytorch/pytorch/issues/159169,d9bdb62b54dbca05efadb4de58b0df0debca532b8b289a22c08477b27d4909e4 references,issue,121969,issue,121894,medium,issue.comments[0].body,"Related: #121894 It looks like #84609 introduced the flag, but searching for TORCH_USE_CUDA_DSA in our docs reveals no results.",https://github.com/pytorch/pytorch/issues/121969,be982636b24802dff85517df09986b4cf0063552769e9c28510bb7322c92461d references,issue,158970,issue,92927,medium,issue.comments[0].body,"place safely, or used as buffers (after their value has been ingested into the optimizer's moments) Kind of the same question discussed in: #92927, where we discussed that torch.compile should be able to generate non-allocating, fused, inplace code for ReLU+Dropout #158643 or...",https://github.com/pytorch/pytorch/issues/158970,ca697b722d1721e75384b55b94eb91d9294e0df7ff58f3093f690fdce565006c references,issue,158970,issue,158643,medium,issue.comments[0].body,"discussed in: #92927, where we discussed that torch.compile should be able to generate non-allocating, fused, inplace code for ReLU+Dropout #158643 or even BatchNorm+Relu+Dropout and that it would be greate to be able to assert such properties somehow, maybe via args to torch....",https://github.com/pytorch/pytorch/issues/158970,01812a220bd2b3efe612159065a148533e3dfebb0d2cc994872126781d33e007 references,issue,75862,issue,51455,medium,issue.comments[0].body,I guess GroupNorm would also benefit from a high-level reference impl in Python (at least for docs purposes) as discussed for LayerNorm in #51455 ...,https://github.com/pytorch/pytorch/issues/75862,6dde2e8e977b3379d8488ca8d8079c61f8b2ecb2970d3060be3a8e4ca4941f78 references,issue,129668,issue,129658,medium,issue.body,"es if they pop up) Thanks for the support team, I was pleasantly suprised that a a lot of the items I had in this list were already closed! #129658 #129659 #130533 #129666 #127173 #129667 #130537 #129665 #130720 #129664 #129662 #130719 #129660 #129661 cc @ezyang @anijain2305 @...",https://github.com/pytorch/pytorch/issues/129668,e6178c4eb7a96357be3ba428ecda22948c1f23d73337e9e78bc5da8bd88c7643 references,issue,112303,issue,111441,medium,issue.comments[0].body,Related #111441,https://github.com/pytorch/pytorch/issues/112303,62d250c68ff45378cb5d96479d5a228989b52500470a2b2e21808fda322c42d2 references,issue,112901,issue,94652,medium,issue.body,create resume_at at end of natural loop (dominator node of natural loop). Optimization: - skip resume_at while no torch.* in the bytecode. #94652 The difficulty with this is knowing the stack and locals at the time of calling resume_at. Graph breaking while inlining - we curre...,https://github.com/pytorch/pytorch/issues/112901,1d15594fe45ba703f8e52987a60b35d905b0ab35acbb9adca2d62303ee5d634a references,issue,112901,issue,111003,medium,issue.body,"e stack and locals at the time of calling resume_at. Graph breaking while inlining - we currently restore to base-level tx before compiling #111003 Requires proper restoration of stacks at all levels, including ferrying the graph outs from the base level back to the top-level...",https://github.com/pytorch/pytorch/issues/112901,99962859ea8f5f41fa8c152b27f49a310ff3def9d647a584270a670bdd4d411a references,issue,112901,issue,112303,medium,issue.body,"level back to the top-level function call (which can then be consumed as locals). Cache calls to inline, guard, and reuse recorded graphs - #112303 Undiagnosed (bug?): Graph breaking in if-statement doesn't continue tracing #111918 Abstractions and Infra There is also an inter...",https://github.com/pytorch/pytorch/issues/112901,f5f66c9f81338c2a054ae1fd0ff7976cb967b574aa1a476dfd70eb1f251dadd5 references,issue,112901,issue,94652,medium,issue.comments[0].body,Issue scrubbing. I'll leave this open. I think #94652 is the only one that's not really worked on.,https://github.com/pytorch/pytorch/issues/112901,37474bd478e4f995448e51644c05c4c2aca918db11b4d63b01f38018f7054ede references,issue,159354,issue,130513,medium,issue.body,"ally clause of TestRPCPickler.test_case to timeout, likely because the RPC init call already failed. The reason is the same as described in #130513 and can be fixed by patching tensorpipe, which unfortunately is archived. Versions PyTorch version: 2.7.1 Is debug build: False C...",https://github.com/pytorch/pytorch/issues/159354,42ef12de23f121a4af204c258ef0c2dd94b05266a9af19ec161981d27aed3407 references,issue,102673,issue,99176,medium,issue.comments[1].body,"1980 these are examples why we need Python debug builds! From quick local testing, some of these also fail in 3.10 debug. Linking the issue #99176",https://github.com/pytorch/pytorch/issues/102673,6e60f3156f7adf32d814cebbeddeecdc3179ea95e4f897e54bfbc2e1d8a3f4e0 references,issue,103800,issue,103761,medium,issue.body,🐛 Describe the bug This is a sub-task of #103761. This is not user-facing as this only involves functions prefixed with an underscore. _no_grad_embedding_renorm_ has mismatched typing on t,https://github.com/pytorch/pytorch/issues/103800,ca15a344204b63500cca41cea68d496d9556ffcf14981669914438e2c889a2b3 references,issue,108984,issue,108934,medium,issue.body,"🐛 Describe the bug While debugging a compile issue on PPC64le in #108934, we realized there is a linker issue in pytorch 2.1.0-rc3 with GCC 11.2.1: python3 -m pip install -r requirements.txt rm -rf build CC=gcc C",https://github.com/pytorch/pytorch/issues/108984,64cbff5a6e978ecd600acd0e5be90428e079a76252da6ca1d3c79944b04417fe references,issue,114859,issue,111384,medium,issue.body,"the numeric differences between torch.ops.aten._native_batch_norm_legit.default and torch.ops.aten.cudnn_batch_norm.default as point out in #111384, since if we move model to cuda before export also gives similar accuracy as * Acc@1 61.326 Acc@5 83.250. Versions (quantization)...",https://github.com/pytorch/pytorch/issues/114859,06d09b1ed2f90c8b7512e2842e417f105cfa787f0f0b5b96df433600862d15d4 references,issue,104849,issue,71683,medium,issue.body,mul in that it should allow inplace multiplication of the accumulator by scalar/tensor prior to adding the rhs summand. Some older context: #71683 (comment) #79352 (comment) #104781 (comment) Here is that line from Adam https://github.com/pytorch/pytorch/blob/main/torch/optim/...,https://github.com/pytorch/pytorch/issues/104849,98885b97298dd8b6b58b70426791b175c29be2525fb2f45f7addf6ee3ee2f150 references,issue,104849,issue,50122,medium,issue.comments[1].body,"uce 0 (then this maybe can also be a generalization of torch.where). Currently this is a problem with torch.lerp: #71701 and is related to: #50122 such pseudo-functions or functions can probably be very well-fused with inductor-generated code Btw, have there been attempts to t...",https://github.com/pytorch/pytorch/issues/104849,a22e28e4edd9993522548d1b67ac6a7bc3c4cd7a1db85bc13c7a308cf911235e references,issue,91692,issue,1529,medium,issue.body,"and pitch It's useful for understanding memory usage and if memory can be saved by refatoring to fusion + inplace The parent issue would be #1529, here's the scope is only on stored intermediate values in autograd graph. But overall, being able to get a list of all tensors and...",https://github.com/pytorch/pytorch/issues/91692,5fde2ae3f19872b8a8ef9f878df74f44a317901c003eff75dfda2989bcd0f0cd references,issue,159229,issue,159230,medium,issue.body,"data/users/jjwu/a/pytorch/torch/_dynamo/utils.py"", line 2046, in create assert name not in scope AssertionError: I think this is related to #159230. Basically, because there's a recompile, we save a new dynamo cache entry of some sort. This is the cache info (i.e., what we sav...",https://github.com/pytorch/pytorch/issues/159229,09fa4c080508a4f30e19bd080ec304900baf05ee3cb0b6b7c9dff83ea85b9d82 references,issue,61582,issue,53900,medium,issue.body,tiple dimensions is blocked on the following issue #61490 torch.max torch.min torch.median torch.nanmedian torch.mode Related Issues #56586 #53900 #61453 cc @mruberry @rgommers @heitorschueroff @pmeier @asmeurer @leofang @AnirudhDagar @asi1024 @emcastillo @kmaehashi,https://github.com/pytorch/pytorch/issues/61582,6b1e12be29a274e93353d34f3f4581466131c6554922fa22c6b1113c511758ad references,issue,61582,issue,56586,medium,issue.body,ver multiple dimensions is blocked on the following issue #61490 torch.max torch.min torch.median torch.nanmedian torch.mode Related Issues #56586 #53900 #61453 cc @mruberry @rgommers @heitorschueroff @pmeier @asmeurer @leofang @AnirudhDagar @asi1024 @emcastillo @kmaehashi,https://github.com/pytorch/pytorch/issues/61582,167f7c898c415c0fc2f7acb931f4b52429c580f3054a70c821836838cf8782cc references,issue,61582,issue,61453,medium,issue.body,imensions is blocked on the following issue #61490 torch.max torch.min torch.median torch.nanmedian torch.mode Related Issues #56586 #53900 #61453 cc @mruberry @rgommers @heitorschueroff @pmeier @asmeurer @leofang @AnirudhDagar @asi1024 @emcastillo @kmaehashi,https://github.com/pytorch/pytorch/issues/61582,e159afd1b51545dacc377b6072b03a123b2b442d813cab5af59a02e9c3f75412 references,issue,61582,issue,61490,medium,issue.body,"nquantile torch.aminmax For the following operators, adding support for reducing over multiple dimensions is blocked on the following issue #61490 torch.max torch.min torch.median torch.nanmedian torch.mode Related Issues #56586 #53900 #61453 cc @mruberry @rgommers @heitorschu...",https://github.com/pytorch/pytorch/issues/61582,dbf9691a9aea0dc06756b792e5d1b5a04238408005826952567d951864970de7 references,issue,155325,issue,97990,medium,issue.comments[1].body,Sounds almost like a duplicate of #97990 Reason for pinning the minor version is that thru out testing we've discovered that relaxing those dependencies leads to unexpected behavio,https://github.com/pytorch/pytorch/issues/155325,c2c925936bf2355c2c4d99efc8bf360bf104560b84dd1dc60d374a25a9bdd0b0 references,issue,155066,issue,152790,medium,issue.body,"ic setups of compiling these libraries from scratch, this would already be useful Discussed with @malfet at: Dao-AILab/flash-attention#1644 #152790 (comment)",https://github.com/pytorch/pytorch/issues/155066,2d0d711713a015cfd40c4e7c4751a01ec1f3189caa178692282e136868888197 references,issue,158610,issue,134784,medium,issue.body,"t: tensor([1., 1.], grad_fn=) Expected output: tensor([-1., -1.], grad_fn=) I believe this is the root cause of #134784 Versions Collecting environment information... PyTorch version: 2.7.1+cu126 Is debug build: False CUDA used to build PyTorch: 12....",https://github.com/pytorch/pytorch/issues/158610,ea0d9f7dc2905ac64dbe34167d5c60cfd284db3f2f75b246c733b25662f648d5 competes with,issue,152220,issue,129140,medium,issue.body,"🚀 The feature, motivation and pitch As discussed in some previous PR/RFC (#129147, #129140), passing in device_id into init_process_group will eagerly init the parent NCCL communicator, and subsequent P2P calls will use that inste",https://github.com/pytorch/pytorch/issues/152220,e32874cd6f6f3d0117324bd90b734cf693d6eef21221adebfcae3b4f6ae64b4e references,issue,31252,issue,1529,medium,issue.body,"ions, I'm happy to try to implement this feature provided some guidance on what direction to take. Linking some of the people who discussed #1529. @vadimkantorov @ezyang @ssnl @VitalyFedyunin @soumith cc @ngimel",https://github.com/pytorch/pytorch/issues/31252,2186be1b078537686eca89b7747f9387e7a8505d0dab11c3d7966569d82cc49b references,issue,106664,issue,105318,medium,issue.body,"torch/tensor functions (and such pages can list all functions with the same name e.g. torch.sigmoid, F.sigmoid which may be good by itself #105318 or https://discuss.pytorch.org/t/various-quantized-quantizable-intrinsic-modules-purpose/183562 for LinearReLU) If we go to https:...",https://github.com/pytorch/pytorch/issues/106664,efb753df431ffa0bf441694e3bf249d426e2599823087525274f4f65ea4df6c8 references,issue,102517,issue,10950,medium,issue.body,estion would be of zero-copy in-RAM IPC data exchange: maybe via named pipes? Some mmaps? shm? I found older issues asking the same: #23490 #10950 This might be important for working in hosted multi-threaded environments where every thread runs its own Python / PyTorch code. T...,https://github.com/pytorch/pytorch/issues/102517,759e2bfbc8242fb5d19aa3770797aa7305d0c5341a5f0bf73ae49dcd1c0c776d references,issue,102517,issue,23490,medium,issue.body,"o, a question would be of zero-copy in-RAM IPC data exchange: maybe via named pipes? Some mmaps? shm? I found older issues asking the same: #23490 #10950 This might be important for working in hosted multi-threaded environments where every thread runs its own Python / PyTorch...",https://github.com/pytorch/pytorch/issues/102517,0be6d995dcce665462244446b2c56975e199c9c489fc20a2682b443917048754 references,issue,112883,issue,113063,medium,issue.body,"s timm_efficientdet - at least locally, OOMs, needs to lower batch size - @eellison cudagraphs dynamic Deferring the dynamic failures until #113063 hf_BigBird Not all values of RelaxedUnspecConstraint(L['inputs'][0].size()[0]) are valid because L['inputs'][0].size()[0] was inf...",https://github.com/pytorch/pytorch/issues/112883,02bb93979fe1c6a6dc95b088bdec29f4646d2ead2622d588d92a9d16416fe4f8 competes with,issue,157547,issue,35527,medium,issue.comments[0].body,"A bit related for supporting dtype=torch.bool for rand* functions - produces similar errors :( : #35527 (comment) Instead of only returning float tensors, maybe torch.rand can be upgraded to do randint if specifically instructed so with dtype",https://github.com/pytorch/pytorch/issues/157547,db8f47673ab44afd532f1f7b5bfd973fbc39625935df23c4fda9b9b335c0b067 references,issue,108565,issue,33041,medium,issue.body,"ymorphic code: bytes(tensor.view('uint8')), currently we need to have to use backend specific np.uint8 or torch.uint8 Maybe related issues: #33041 #43949 Versions np.__version__ # '1.24.2' torch.__version__ # '2.1.0.dev20230802+cpu' cc @mruberry @rgommers",https://github.com/pytorch/pytorch/issues/108565,c962d348e529e18d9cee7735b216edd65409818cb299a7d4af21a8f136a4d1c9 references,issue,108565,issue,43949,medium,issue.body,"c code: bytes(tensor.view('uint8')), currently we need to have to use backend specific np.uint8 or torch.uint8 Maybe related issues: #33041 #43949 Versions np.__version__ # '1.24.2' torch.__version__ # '2.1.0.dev20230802+cpu' cc @mruberry @rgommers",https://github.com/pytorch/pytorch/issues/108565,fec1a3ac9059de94e04aa4099e9c83c3ca6da56abefedc5952a58290047a6e0c references,issue,118715,issue,118018,medium,issue.comments[0].body,This is a subset of #118018,https://github.com/pytorch/pytorch/issues/118715,294d05feffdfffd9c4910328432e8a99b4da30a6a5158c6bd964a26fe0f5bde2 references,issue,111693,issue,112043,medium,issue.comments[1].body,Similar issue as: #112043,https://github.com/pytorch/pytorch/issues/111693,70d204a99c341722b580d52eb0240f3eaeb0e817ae1d827bb53b2def969a64fb references,issue,133250,issue,134363,medium,issue.body,"tly O(# of nodes in graph) Inductor pattern matcher's replacement graph is assumed to be functional, but we don't check this. Also related: #134363 cc @ezyang @chauhang @penguinwu @voznesenskym @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv @ji...",https://github.com/pytorch/pytorch/issues/133250,1bc05dbaaed0889e58afe13959f4b371fe2a21667d4eb51d698edcc13679da2f references,issue,107112,issue,34646,medium,issue.comments[0].body,"Related: #34646 where I use some Storage.from_buffer workarounds As far as I know, there currently only exists completely undocumented TypedStorage/Untyped",https://github.com/pytorch/pytorch/issues/107112,c45790570f4d487eeffe958b0f5c83ce70a10d66cac8698ecbc4e63d34f0f5d3 references,issue,151740,issue,120237,medium,issue.comments[0].body,Could be related to this #120237,https://github.com/pytorch/pytorch/issues/151740,dc45a65da84572f540c044de4a4de2567f914aa6f145ae271ef1f0cdd04950b0 references,issue,68905,issue,55267,medium,issue.body,"/github/home/.cache/torch_extensions/py38_cu111/cpu_adam/cpu_adam.so: undefined symbol: curandCreateGenerator As I flagged some months back #55267 ideally the cuda extensions shouldn't be installed into a single shared dir, but to be installed into the virtual environment tree...",https://github.com/pytorch/pytorch/issues/68905,56c0fae124d9c99b2c1324fd6c293d81b58a588eda2e38119b69bbf772ba5f26 references,issue,156616,issue,132378,medium,issue.body,"re) but no one seems to have found a satisfying solution. I also looked at existing issues about this, and I found this very related issue: #132378. In this issue's case, a is_outputs_batched parameter would have elegantly solved the person's problem (although in their case th...",https://github.com/pytorch/pytorch/issues/156616,19c72e45e9c1f30c84ddc064fab01a89359481bc3174d26fa3874aa61bfd8292 references,issue,157431,issue,64947,medium,issue.body,"🐛 Describe the bug It seems anyway that this function would benefit from rewriting (#64947), but here's a small edge case problem. >>> tensor = torch.tensor([0.0, 1.0, 2.0]) >>> torch.quantile(tensor, 0.5000001, interpolation=""mid",https://github.com/pytorch/pytorch/issues/157431,cdc24e1974d54cf925aaa9f8abe6623ff8fba57a31063c59a551e644861b6374 references,issue,152877,issue,71117,medium,issue.body,"g training examples may also add community cycles to testing and tweaking the forward AD support (*_jvp operators) Pytorch has invested in: #71117 Textbook context, for example: (Chapter 6 of Deep Learning (Goodfellow, Bengio, Courville)): ""when the number of outputs of the gr...",https://github.com/pytorch/pytorch/issues/152877,f37902c7fae5d45042368fba1e62913612b2da2ff37e91d5705a26423bd2436f references,issue,94718,issue,91309,medium,issue.body,"t-device synchronization. This has been discussed in other issues regarding reconstruction with certain windows when center=False ( #62323, #91309), but my proposal is from a performance standpoint (https://discuss.pytorch.org/t/torch-istft-nola-check-causes-synchronization-an...",https://github.com/pytorch/pytorch/issues/94718,dd6b55c3438990476c88fe6e0ff991ba9d446206a46b6cbe92e0369c246e383b references,issue,89255,issue,28341,medium,issue.comments[0].body,"tps://github.com/cornellius-gp/linear_operator. For reference, here's the long-running feature request for LinearOperators in PyTorch core: #28341",https://github.com/pytorch/pytorch/issues/89255,8789db05eb34c5e34229c2d5c4253ec1102f8dce047613cf17e06c211e45ffec references,issue,155171,issue,154984,medium,issue.comments[0].body,@soulitzer are you interested in taking a look at AC problem for qwen2? The repro is on single device. This further helps AC + fsdp2: #154984,https://github.com/pytorch/pytorch/issues/155171,6212effc7069797d706ceb901ab87eb739100b855dbefdcc837e3be4dde87329 references,issue,155171,issue,151936,medium,issue.comments[1].body,hmm. AC was doing well in Apr #151936,https://github.com/pytorch/pytorch/issues/155171,9b99ab6d7b23e901698f6577e32b7ed604345876bba4b3d8d806b2b629e68f9e references,issue,150277,issue,152226,medium,issue.comments[0].body,exing can be formulated as a element wise multiplication with a sparse mask tensor having ones values and indices corresponding to idx. See #152226 (comment) that resolves an equivalent problem.,https://github.com/pytorch/pytorch/issues/150277,694cd6f4ec2f3379b97f8a10c0ea5ed79e9a7ed8cd0c2b34a3cfe381faf050c9 references,issue,155060,issue,154329,medium,issue.body,"is a few seconds into the script: On CPU, I don't get any memory pressure at all. I initially documented these observations as comments to #154329, which I think might be related. Versions PyTorch version: 2.7.0 Is debug build: False CUDA used to build PyTorch: None ROCM used...",https://github.com/pytorch/pytorch/issues/155060,82ee03eb89c37a6fb0e1225a9601bba84787fccc50ca8aa025f0a98e459bcef4 references,issue,155218,issue,95024,medium,issue.body,"31, it requires 22 seconds to complete the task. The application of the above operation is for video diffusion. Similar issues are found in #95024 and #152816 Versions PyTorch version: 2.7.1+cu126 Is debug build: False CUDA used to build PyTorch: 12.6 ROCM used to build PyTorc...",https://github.com/pytorch/pytorch/issues/155218,c32b716b1f4c3e3c27a63b18b9dc015cc27be689751099d4699d99ccd19f6fdd references,issue,155490,issue,145511,medium,issue.body,Add feature to PP to perform activation offloading (based on PipeOffload: https://arxiv.org/pdf/2503.01328). related: #145511 cc @awgu @wanchaol @fegin @fduwjj @wz337 @wconstab @d4l3k,https://github.com/pytorch/pytorch/issues/155490,cef59911cd9d0d8dff9dad207ae46355dd42c68ad047e3e887d140f8056d2101 references,issue,149325,issue,140570,medium,issue.body,"nd how to run tests, as allocating large tensors on machines we have right now simply not going to work Example of existing issues: #149261 #140570 #122916 #116769 #116769 #143859 #154828 Versions CI cc @kulinseth @albanD @DenisVieriu97 @jhavukainen",https://github.com/pytorch/pytorch/issues/149325,dd2da3e2dd3818bae027382c65e51091c57cfefc4d868089419178f80035b898 references,issue,149325,issue,149261,medium,issue.body,"en ops and how to run tests, as allocating large tensors on machines we have right now simply not going to work Example of existing issues: #149261 #140570 #122916 #116769 #116769 #143859 #154828 Versions CI cc @kulinseth @albanD @DenisVieriu97 @jhavukainen",https://github.com/pytorch/pytorch/issues/149325,757b9da16aaf74a6ac5339e2afc7823f89bc9895ac66430850b8c2e7b44a87db references,issue,133676,issue,55279,medium,issue.comments[0].body,"Related: https://github.com/facebookresearch/theseus #108977 #55279 #109017 (for easy, zero-copy conversion between PyTorch's COO/CSR and SciPy's COO/CSR)",https://github.com/pytorch/pytorch/issues/133676,50d11ca36a06da95f4f7b8635706d0f13ee9e527ecbdce374781e22d27c069c7 references,issue,133676,issue,108977,medium,issue.comments[0].body,"Related: https://github.com/facebookresearch/theseus #108977 #55279 #109017 (for easy, zero-copy conversion between PyTorch's COO/CSR and SciPy's COO/CSR)",https://github.com/pytorch/pytorch/issues/133676,95116a909d0fabadcdbb06dc679822fbe2c0d617045eb21d02ea3e575d849fc5 references,issue,133676,issue,109017,medium,issue.comments[0].body,"Related: https://github.com/facebookresearch/theseus #108977 #55279 #109017 (for easy, zero-copy conversion between PyTorch's COO/CSR and SciPy's COO/CSR)",https://github.com/pytorch/pytorch/issues/133676,8e348769dbc37ae40d57ce64324dbed54a3c9c2599705c0b948fecbaa6f39639 references,issue,150612,issue,13246,medium,issue.comments[1].body,"@ilyas-sirazitdinov-snkeos There has been discussion previous on this. Sharing in case that is useful, for eg. #13246 (comment). Assuming your setup works with num_workers=0, can you try using persistent_workers=False and/or using the default multiprocessin",https://github.com/pytorch/pytorch/issues/150612,dfde2e1160cd578efd8b29a5736ccce0bcc281bfaa1237669e919644e6411ac9 references,issue,101192,issue,147380,medium,issue.comments[0].body,"we had this issue come up for the joint-graph API, for tied weights: #147380",https://github.com/pytorch/pytorch/issues/101192,427ec9c368281f3f790dec42a39d69c7e7bfcaf6a5dece87f4328cf11686ebcc references,issue,99625,issue,91989,medium,issue.body,"🐛 Describe the bug This issue may be related to #98836 and #91989. Background The issue arise when using PyTorch with Ray. The raylet process is forked from the main process after importing torch, then ray",https://github.com/pytorch/pytorch/issues/99625,32651444fd62884afa2592aea15cdc1f718b0cf77137062f36036e5aacec4c4e references,issue,99625,issue,98836,medium,issue.body,🐛 Describe the bug This issue may be related to #98836 and #91989. Background The issue arise when using PyTorch with Ray. The raylet process is forked from the main process after importing torc,https://github.com/pytorch/pytorch/issues/99625,889086665113b35d81773fc47363d80651235aab16b715b557050c554da6e3eb references,issue,141133,issue,139298,medium,issue.body,"s track of known issues and the current state of the integration, with the goal to bump order in priority list with higher confidence Tasks #139298 Match Query's stride layout #138354 gradOutput vs Ouput stride mismatch, blocked by cudnn bump: #138354 (comment) Fill Value twid...",https://github.com/pytorch/pytorch/issues/141133,8994db566c06b7d14b0ec9d9d33da8824bd0a883ece2c9e4bfa7e90400ea4066 references,issue,109770,issue,108175,medium,issue.comments[0].body,See #108175,https://github.com/pytorch/pytorch/issues/109770,3975362494a4f483884242856435e1cb9610c2fc838bc5a85ba6628bbd61fe05 references,issue,127622,issue,28307,medium,issue.body,mespace)::CudaIPCSentDataDelete(void*) () from /home/tuero/miniconda3/envs/hptslite/lib/python3.11/site-packages/torch/lib/libtorch_cuda.so #28307 0x00007fff98c6047a in torch::CudaIPCSentData::~CudaIPCSentData() () from /home/tuero/miniconda3/envs/hptslite/lib/python3.11/site-...,https://github.com/pytorch/pytorch/issues/127622,67501023fe733d605219f9b3364bc390d055b979d4570508f69fdb13d577608d references,issue,127622,issue,28308,medium,issue.body,h::CudaIPCSentData::~CudaIPCSentData() () from /home/tuero/miniconda3/envs/hptslite/lib/python3.11/site-packages/torch/lib/libtorch_cuda.so #28308 0x00007fff98c62405 in torch::(anonymous namespace)::CudaIPCSentDataDelete(void*) () from /home/tuero/miniconda3/envs/hptslite/lib/...,https://github.com/pytorch/pytorch/issues/127622,d958099f2120216a1b4d014f0bbb2fe2b8353bd18af1339a25710c4cc0fa1299 references,issue,127622,issue,28318,medium,issue.body,fffe3d54cdb in c10::TensorImpl::~TensorImpl() () from /home/tuero/miniconda3/envs/hptslite/lib/python3.11/site-packages/torch/lib/libc10.so #28318 0x00007fffe3d54e89 in c10::TensorImpl::~TensorImpl() () from /home/tuero/miniconda3/envs/hptslite/lib/python3.11/site-packages/tor...,https://github.com/pytorch/pytorch/issues/127622,073803d774d20af473499d3ca5a17b57703122978ec3392d3b6cb32bfcb1ed14 references,issue,127622,issue,28319,medium,issue.body,fffe3d54e89 in c10::TensorImpl::~TensorImpl() () from /home/tuero/miniconda3/envs/hptslite/lib/python3.11/site-packages/torch/lib/libc10.so #28319 0x00007fffe2a7cf98 in THPVariable_clear(THPVariable*) () from /home/tuero/miniconda3/envs/hptslite/lib/python3.11/site-packages/to...,https://github.com/pytorch/pytorch/issues/127622,509a9f17f0358c44db290d7522b123cfefd6cd3814ccd7f9c15094490b8eaa3c references,issue,127622,issue,28320,medium,issue.body,8 in THPVariable_clear(THPVariable*) () from /home/tuero/miniconda3/envs/hptslite/lib/python3.11/site-packages/torch/lib/libtorch_python.so #28320 0x00007fffe2a7d2e6 in THPVariable_subclass_dealloc(_object*) () from /home/tuero/miniconda3/envs/hptslite/lib/python3.11/site-pack...,https://github.com/pytorch/pytorch/issues/127622,960fa3e297fcd0ddb12c7d0ee516d80df3854537f350abed7bbc4995b18c768f references,issue,127622,issue,28329,medium,issue.body,/conda/python-3.11.8/Objects/object.c:2390 #28328 Py_DECREF (op=) at /usr/local/src/conda/python-3.11.8/Include/object.h:538 #28329 Py_XDECREF (op=) at /usr/local/src/conda/python-3.11.8/Include/object.h:602 #28330 free_keys_object (keys=0x1547582...,https://github.com/pytorch/pytorch/issues/127622,b4461d9549926a42525ccf9b831cca3ce7a51d2369c57d371b93682860fdb406 references,issue,127622,issue,28341,medium,issue.body,/python-3.11.8/Include/object.h:602 #28340 dict_dealloc (mp=0x7fff2432fd00) at /usr/local/src/conda/python-3.11.8/Objects/dictobject.c:2371 #28341 0x000000000055ea02 in Py_DECREF (op=) at /usr/local/src/conda/python-3.11.8/Include/object.h:538 #28342 subtype_dea...,https://github.com/pytorch/pytorch/issues/127622,0d26534c96695bcf4c22eb1af6006869d6e58ba384c6b98565599750891781cb references,issue,154259,issue,120115,medium,issue.comments[0].body,"Thanks for taking this one! For detectron2* timeout, it's probably related to #120115, which can fixed by skipping setting non-deterministic here, pytorch/benchmarks/dynamo/common.py Line 3592 in 28af442 if args.only is not N",https://github.com/pytorch/pytorch/issues/154259,e70cba4c7c8c675f1c98f0bfeb7c63979aff572a8a0ae817e15936c872d43bb8 references,issue,64208,issue,62352,medium,issue.comments[1].body,"shape[I.ndim + 1:])).squeeze(I.ndim) In addition to being quite convoluted, this would also quite likely fail in JIT Related issues: #7991, #62352",https://github.com/pytorch/pytorch/issues/64208,f0907408d47b6999181f86c25043290206d48536bef339c189dd6639fadd5cee references,issue,153961,issue,145702,medium,issue.body,"🐛 Describe the bug Follow @Blackhex ‘s comments: #153480 (review) So, we open this issue. Original issue: #145702 Fixed by add /d2implyavx512upperregs- build option: #153480 Still need to consider, upgrade VS2022: Reason: #145702 (comment) Try to upgrad",https://github.com/pytorch/pytorch/issues/153961,098ca928797cc8bdee396a66f8e2bad8a42f8876f812083a7a119f6149663b03 references,issue,153938,issue,126523,medium,issue.body,"🐛 Describe the bug We are parsing the test-reports (XML) files after running tests, see also #126523 and noticed an error in a sanity check I added: The number of tests doesn't add up. E.g.: first) Convenient / consistent padding for a single dim or subset of dims Complete set of padding modes matching real use cases #57911 points out that PyTorch's circular padding mode only allows wrapping once, whereas numpy's wrap mode will wrap any number of times...",https://github.com/pytorch/pytorch/issues/153854,752ff75f8aa4d4ab84a307bb281f36dc21e64250d02ab40db8d17ebe80794382 references,issue,153854,issue,60294,medium,issue.body,"the last dimension and working back. Padding does not have to be specified for all dimensions. Proposed APIs: Match numpy's pad() API (see #60294 / ##61635). import numpy as np a = [1, 2, 3, 4, 5] np.pad(a, (2, 3), 'constant', constant_values=(4, 6)) array([4, 4, 1, ..., 6, 6,...",https://github.com/pytorch/pytorch/issues/153854,48c8bdee38f9b80e808ab0fd52c70491982cdede5eaf84a25dd07d615b6a81b4 references,issue,105325,issue,65156,medium,issue.body,"added (I can't think of a good way to make vmap do this.) Annoyance is avoiding wasted FLOPs by ""adjusting down"" the logical size. Related: #65156 Related: nested tensor cc @cpuhrsch @jbschlosser @bhosmer @drisspg @msaroufim @albanD Versions main",https://github.com/pytorch/pytorch/issues/105325,14673f019e4850ee6e37882ba252d51408412501ad21f84301b4e619c1e33ddd references,issue,151912,issue,124505,medium,issue.comments[0].body,"backend has previously been used.. to get rid of this persistence, one must clear the inductor cache Maybe related to this dynamo design :( #124505",https://github.com/pytorch/pytorch/issues/151912,f361dea8d0e36969742fb0d194b1618e433dffd2d4a73471d17fe0c2b4046d01 references,issue,153302,issue,147179,medium,issue.body,🐛 Describe the bug Got 2 OSS issues around GPU OOM for reshard_after_forward=int #147179 #149029 we are creating a new device mesh for every fully_shard. Each device mesh creates a new PG. Each PG takes 1GB memory. The problemat,https://github.com/pytorch/pytorch/issues/153302,217b300162d2afd97366b36ee17930d2bf8bda81d05a4d5d9a549fe9c0d54096 references,issue,123447,issue,95895,medium,issue.body,".async_save. Interestingly enough, this does not happen if we first call all_reduce (commented out in the example). Potentially related to: #95895 (comment) import os import torch import torch.distributed.checkpoint as dcp import torch.multiprocessing as mp import torch.distri...",https://github.com/pytorch/pytorch/issues/123447,4d517a78dd53222d874678de72bd3da8e117d8e1110150b3c57a5a0ebf3f8d9c references,issue,9674,issue,1369,medium,issue.body,"hat implicit 0s do not participate, and it would allow us to have a sparse gradient tensor. We can call this operator s_log1p or nnz_log1p (#1369 has similar discussions). However, we are not sure if they are really what users want, because these operators have different behav...",https://github.com/pytorch/pytorch/issues/9674,b82b646c3a1f67dee9c7234e315944af6dc99d958c82e6a975fb871a2a0a8bf1 references,issue,51039,issue,47993,medium,issue.comments[7].body,"port torch; print(torch._C._GLIBCXX_USE_CXX11_ABI)""` with the conda version also yields `False` I tried compiling from source but am facing #47993. Any feedback or tips would be appreciated.",https://github.com/pytorch/pytorch/issues/51039,18fd22059eb5a2b39f9dd6ae6cbea04897b9b0d8619f3998e2480517b7342979 references,issue,151579,issue,24422,medium,issue.comments[1].body,"Please remove me kit1980 from #24422 , I can't edit that.",https://github.com/pytorch/pytorch/issues/151579,912f062fedc8e8fb65ea0cc6a5eaceb1834f0dc96a395803c05bcece368a2c3c references,issue,117350,issue,101314,medium,issue.body,"nv, but fails with Bazel. I also created a tool there that patches the Torch wheel. The patched wheel works in Bazel for us. Related Issues #101314 #92096 Versions $ python collect_env.py /pay/tmp/venv-torch21/lib/python3.8/site-packages/torch/nn/modules/transformer.py:20: Use...",https://github.com/pytorch/pytorch/issues/117350,85f520ecf23c304a893c74aaa20740e4dad1a91f76c45f9231421614eaa91114 references,issue,106624,issue,54389,medium,issue.body,"forms.Normalize.html) Being able to do int32_tensor.mul_(float_constant) or uint8_tensor.mul_(float_constant) is also useful and related to #54389. It could certainly support rounding, similar to torch.div(..., rounding_mode = 'trunc') e.g. by doing int32_tensor.mul_(numerator...",https://github.com/pytorch/pytorch/issues/106624,57a45bc5cec745e4e9fdfb4b7187b743ea496d222d7fbd45e1e331b07e8cc17a references,issue,106624,issue,104849,medium,issue.body,", dtype=...) (torch.add(a, b, alpha = ...) exists but also doesn't support the needed casting). Related: addcmul generalization discussion: #104849 (although I realize, this might be somwhat niche - also sth similar exists in torchvision: https://pytorch.org/vision/stable/gene...",https://github.com/pytorch/pytorch/issues/106624,15ba782af2858b7ab3c9f3f23153717ca9385d57b1c9563a9b68687f2623bf43 references,issue,106624,issue,42959,medium,issue.comments[1].body,"Related: #42959 NumPy does provide a dtype= argument for np.divide, but seems only for upcasting: #42959 (comment)",https://github.com/pytorch/pytorch/issues/106624,37ffc67bedfa388270d2bb28338104dca524dff5aeb53d81dcd6b73df4d7b737 references,issue,57947,issue,56356,medium,issue.comments[0].body,"We should be more consistent about our interpolation UX, but we probably don't want to extend type promotion support at this time. See #56356",https://github.com/pytorch/pytorch/issues/57947,ad7299b4bcba67cd5f221d16b548f554e98a7ce471daa412496b444a5a442cf6 references,issue,151960,issue,150943,medium,issue.body,w/ a /user/ prefix. This will make it easier for more advanced users to instantiate PGs directly as well as for users to use things like in #150943 dist.store_from_uri(...) and Store.uri() create persistent URIs that can be exchanged to rehydrate stores and share access across...,https://github.com/pytorch/pytorch/issues/151960,381d6eb008edf0376c46de9536feada8da5d916c996d15df1d5cd936163cdc75 competes with,issue,145693,issue,140787,medium,issue.body,". Adding support for unified memory on ROCm for that particular APU would allow for zero-copy operations. The motivation is similar that of #140787, but for ROCm instead of MPS. Given that this APU is targeted for the most demanding HPC ML workloads, there is a great interest...",https://github.com/pytorch/pytorch/issues/145693,578dde8e66995fedb7c4742732d70a1a002f1dcf9ddb2ecede9bdb2b5f3ec875 references,issue,148048,issue,108158,medium,issue.body,"ing script to notify agent, and workarounded by modified agent. A less related issue about global restarting count for external management: #108158. cc @H-Huang @awgu @kwen2501 @wanchaol @fegin @fduwjj @wz337 @wconstab @d4l3k @c-p-i-o",https://github.com/pytorch/pytorch/issues/148048,b1d7e117ea668b65e1b72b7a2d61035ba551ab20c09e08f0b38343f9b5458519 references,issue,148048,issue,133877,medium,issue.body,"e case in #136312, looking for a method of handling agent node restart on hardware-failure. Another discussion aboud non-retriable error in #133877, which seeks interface for training script to notify agent, and workarounded by modified agent. A less related issue about global...",https://github.com/pytorch/pytorch/issues/148048,3b163d49479c91050153fff7662261193db0d9d1ec9d91eb00aeab704965359a references,issue,148048,issue,136312,medium,issue.body,"ausing large-scale rescheduling (or even re-queued for long time) Additional context There is a issue aboud k8s fault-tolerance use case in #136312, looking for a method of handling agent node restart on hardware-failure. Another discussion aboud non-retriable error in #133877...",https://github.com/pytorch/pytorch/issues/148048,2d709048e68e81be54baaf04066720aa06f4f7c28cb392ecfb736f1246f560a9 references,issue,147141,issue,88,medium,issue.body,EvalFrameDefault from ??:0 #85 _PyFunction_Vectorcall from ??:0 #86 _PyEval_EvalFrameDefault from ??:0 #87 _PyFunction_Vectorcall from ??:0 #88 _PyEval_EvalFrameDefault from ??:0 #89 method_vectorcall from :0 #90 _PyEval_EvalFrameDefault from ??:0 #91 _PyFunction_Vectorcall fr...,https://github.com/pytorch/pytorch/issues/147141,d92b6432dcffb64ecbf7b1c98171b1a0cc120068b9d06a7b5b56fe85c1470e61 references,issue,147513,issue,129553,medium,issue.body,"ted, which aim to enhance PyTorch's support for RISC-V architecture and the RVV (RISC-V Vector Extension). Current Progress As mentioned in #129553, we have completed the key areas of focus, as follows: Kernel Optimization Enhance DepthwiseConvKernel to take advantage of RVV....",https://github.com/pytorch/pytorch/issues/147513,6ef499d614c0abeaf9919da28104990f1fcd54672e3807d14685be7e1b80351f references,issue,147513,issue,141550,medium,issue.body,"rchitecture, thereby accomplishing the compilation and build verification of the code. Please refer to #143979 for details. Please refer to #141550 for further discussion In order to enable cross-compilation of PyTorch, support for cross-compiling the third-party library SLEEF...",https://github.com/pytorch/pytorch/issues/147513,b9bf19badc8a6a0dccd9ac55435660cf450a26d9e60c862f6998585f2b0533b4 references,issue,151561,issue,151558,medium,issue.body,"same for all ensemble members. if that worked I'd be happy too. but SDPA's validation of masks' batch dimensions is broken when under vmap (#151558). from typing import Optional import torch from torch import BoolTensor, FloatTensor, Tensor, inference_mode from torch.nn import...",https://github.com/pytorch/pytorch/issues/151561,719342622f8ab42d5baf79d7b58742e516d33ab098bb84c8091f0e8376cbad68 references,issue,144206,issue,70608,medium,issue.comments[1].body,"This (#144206), #70608, and #104113 seem to be duplicates.",https://github.com/pytorch/pytorch/issues/144206,436222181727b474c31e95484e373f855b8908701c5ea99d4fb73a109295abe2 references,issue,144206,issue,104113,medium,issue.comments[1].body,"This (#144206), #70608, and #104113 seem to be duplicates.",https://github.com/pytorch/pytorch/issues/144206,b43b55e9f4e3e80c28b029df617f4dfd5379517c135dc605fa99e900c16d1b04 references,issue,151558,issue,151559,medium,issue.body,🐛 Describe the bug note: I also ran into an SDPA docs bug #151559 whilst producing this repro. note: I also ran into a vmap dispatch bug #151561 whilst producing this repro. SDPA gives us RuntimeError: att,https://github.com/pytorch/pytorch/issues/151558,984d13b81adc4a39d9069ad3fde4b11affba06714132aeaf02c4aa003a05e500 references,issue,151558,issue,151561,medium,issue.body,🐛 Describe the bug note: I also ran into an SDPA docs bug #151559 whilst producing this repro. note: I also ran into a vmap dispatch bug #151561 whilst producing this repro. SDPA gives us RuntimeError: attn_bias: wrong shape (batch dimension) when we vmap a mask over it. the s...,https://github.com/pytorch/pytorch/issues/151558,f297e8b0f03b9f1995010191960da832b84c3fd5911191bfc686000568fa0052 references,issue,101699,issue,13246,medium,issue.body,useful to avoid copies related to copy-on-write (actually copy-on-read because of python's finicky ref-counters) problems with DataLoader: #13246. A typical application: list of file names or file paths in a dataset (avoiding creating hundreds of thousands/millions of python s...,https://github.com/pytorch/pytorch/issues/101699,e083445a3415e675859a99cde048fd1886d6e39ebcc7fd084a626d26ffa6705c references,issue,101699,issue,33041,medium,issue.body,"oncept can be ""string lists"" that allow appends (with some exponential storage reallocation): #64359 Related issues on ""zero-copy"": #43949, #33041, #34651 (about getting a bytes view over a sub-tensor - can be useful as an ascii string substitute, and in general for zero-copy...",https://github.com/pytorch/pytorch/issues/101699,b7ca356a54ca3154b6e8ef6c002936d3537fa5ee31bd2bc45eb79daa701c74d2 references,issue,101699,issue,34651,medium,issue.body,"an be ""string lists"" that allow appends (with some exponential storage reallocation): #64359 Related issues on ""zero-copy"": #43949, #33041, #34651 (about getting a bytes view over a sub-tensor - can be useful as an ascii string substitute, and in general for zero-copy pytorch...",https://github.com/pytorch/pytorch/issues/101699,a5bcc8c9be55ad66006e616719413f589345fb76b7bb9cbf364d7f3a980207a8 references,issue,101699,issue,43949,medium,issue.body,"useful concept can be ""string lists"" that allow appends (with some exponential storage reallocation): #64359 Related issues on ""zero-copy"": #43949, #33041, #34651 (about getting a bytes view over a sub-tensor - can be useful as an ascii string substitute, and in general for ze...",https://github.com/pytorch/pytorch/issues/101699,a8da7064f630432a4acaec1dc2a12a424196f116b93250bcde39c22415859d06 references,issue,101699,issue,64359,medium,issue.body,"tion / keys hash computation. Another useful concept can be ""string lists"" that allow appends (with some exponential storage reallocation): #64359 Related issues on ""zero-copy"": #43949, #33041, #34651 (about getting a bytes view over a sub-tensor - can be useful as an ascii st...",https://github.com/pytorch/pytorch/issues/101699,a80a91d3579548ba4585c90b9fc08d67fdfbb3c5bcd2bba63070cdf97429eb0e references,issue,148743,issue,130217,medium,issue.body,guarantee nothing is missed just specify the len of that dimension from the input. Hopefully I made this part clear. NOTE: I disagree with: #130217 I find the simple 1d case understandable and usable. But with an additional dimension(s) it isn't that it couldn't be understanda...,https://github.com/pytorch/pytorch/issues/148743,006983a145d189644492aa3019234d960514763725cd9ed7f26d4fbe7128c661 references,issue,138696,issue,148335,medium,issue.body,"indows.yml Getting started page nightly instructions Validation Framework in builder - pytorch/builder#2036 #148339 #145872 #145890 #148110 #148335 #148336 #148340 #148338 #148341 #148342 Phase 2 CI Deprecation (WIP Scection): Domains CI: text, vision, audio, tensorrt, data Re...",https://github.com/pytorch/pytorch/issues/138696,1491b1838b5777b0427fb7f029096d7c2620e6fabb8b208f113365263af544ac references,issue,138696,issue,148336,medium,issue.body,"ml Getting started page nightly instructions Validation Framework in builder - pytorch/builder#2036 #148339 #145872 #145890 #148110 #148335 #148336 #148340 #148338 #148341 #148342 Phase 2 CI Deprecation (WIP Scection): Domains CI: text, vision, audio, tensorrt, data Replace by...",https://github.com/pytorch/pytorch/issues/138696,475d6b5edcba11307229f2b930195233e5b7783b82036908b5ef448678ae4dea references,issue,138696,issue,148338,medium,issue.body,"ed page nightly instructions Validation Framework in builder - pytorch/builder#2036 #148339 #145872 #145890 #148110 #148335 #148336 #148340 #148338 #148341 #148342 Phase 2 CI Deprecation (WIP Scection): Domains CI: text, vision, audio, tensorrt, data Replace by miniforge and c...",https://github.com/pytorch/pytorch/issues/138696,3b074e1d0b5ac620b8245842351a07e5ac518ce82b2cbfeddc4e64c931ee8590 references,issue,138696,issue,148339,medium,issue.body,"/.github/workflows/build-magma-windows.yml Getting started page nightly instructions Validation Framework in builder - pytorch/builder#2036 #148339 #145872 #145890 #148110 #148335 #148336 #148340 #148338 #148341 #148342 Phase 2 CI Deprecation (WIP Scection): Domains CI: text,...",https://github.com/pytorch/pytorch/issues/138696,feef55995e079673a13fa1467f1696d8a40f9b1a9ca01c68e62b97a3dad31120 references,issue,138696,issue,141764,medium,issue.comments[0].body,"e nightly instructions Documentation, Tutorial pytorch/builder#2036 Torchbench Domains CD visio, audio, data Make workflow test-infra no-op #141764 PyTorch Core and Builder CD and tests PyTorch conda Docker images builds",https://github.com/pytorch/pytorch/issues/138696,272769ac03b74c996e46a65a70ca4dbf86aa79038fd0ee6dafefc19be244d7c0 references,issue,150285,issue,103820,medium,issue.body,"🐛 Describe the bug I need to compute very large sparse matrix multiplications in my project, and I encountered the same issue as #103820. It's mentioned that CUDA 12 provide two new algorithms with less memory occupation. I tested them on my matrices and they worked well, so",https://github.com/pytorch/pytorch/issues/150285,96c87ace6e9d333cfec9ad4721833c834f335656134811f5bc968d152d45c796 references,issue,61474,issue,46544,medium,issue.body,case the nan* operators would call the corresponding operator with the correct nan flag. See SciPy's design for nan_policy see also: #21987 #46544 nan* operators to implement #21987 torch.nanmax torch.nanmin torch.nanstd torch.nanvar torch.nanargmin torch.nanargmax torch.nanpr...,https://github.com/pytorch/pytorch/issues/61474,795e508e5b72dddb9ecba8ea27c1a3ee3431bbe261623e5821a4249a1424d42c references,issue,149909,issue,91439,medium,issue.comments[0].body,maybe related: #91439 #123089 #120626,https://github.com/pytorch/pytorch/issues/149909,509601dcd1b0597127a8d4af18cc4eeea5e1b607044589313125003d11ff4f2b references,issue,146201,issue,145710,medium,issue.comments[1].body,"There is currently no plan for adding bounds support to L-BFGS, though there has been at least 1 issue talking about it lately: #145710 We typically are conservative about adding new functions/optimizers to the codebase and would encourage development in third-party repos fi",https://github.com/pytorch/pytorch/issues/146201,427195d85f54f0bdb678c6929040b619597c00b0502b357529faa2da56495a3c references,issue,68978,issue,67760,medium,issue.comments[0].body,"er: scheduler.step(*args, **kwargs, epoch=0) else: scheduler.step(*args, **kwargs) Inheritance and modifying _LRScheduler Given issues like #67760, #68332, and #68979, perhaps we should rewrite ReduceLROnPlateau to inherit from _LRScheduler and modify parts of the base class A...",https://github.com/pytorch/pytorch/issues/68978,696daf6bcb4f95d4d6dcd60660d7c877e443983ad61f68c416d207e71c7e2c6e references,issue,68978,issue,68332,medium,issue.comments[0].body,"duler.step(*args, **kwargs, epoch=0) else: scheduler.step(*args, **kwargs) Inheritance and modifying _LRScheduler Given issues like #67760, #68332, and #68979, perhaps we should rewrite ReduceLROnPlateau to inherit from _LRScheduler and modify parts of the base class API. This...",https://github.com/pytorch/pytorch/issues/68978,4ae64373515cf431585075eb20d6b978b9bcb55171b9c740c043c7dfc7ec9899 references,issue,68978,issue,68979,medium,issue.comments[0].body,"args, **kwargs, epoch=0) else: scheduler.step(*args, **kwargs) Inheritance and modifying _LRScheduler Given issues like #67760, #68332, and #68979, perhaps we should rewrite ReduceLROnPlateau to inherit from _LRScheduler and modify parts of the base class API. This direction w...",https://github.com/pytorch/pytorch/issues/68978,1d51ac3e5fd36dd3d0c7f5f56c22fe8f7faca09472f7a484725bd40544e5b167 references,issue,146647,issue,146414,medium,issue.body,Context Context: #146414 Tracking the issues we encounter as we add MX dtypes to core for fixing later. printing of shell dtypes We should be able to print a tensor,https://github.com/pytorch/pytorch/issues/146647,00a5ecfbcea5ed84e27211c5d1783c604cdcf5b56570f40b10c174d55f5888e8 references,issue,149734,issue,146898,medium,issue.body,"or its triton and halide backends. Alternatives Note: this feature can probably be handled by the more general device abstraction proposal: #146898 but this separate issue was created for visibility, especially for the suggestion to deprecate config.{device}_backend for an alt...",https://github.com/pytorch/pytorch/issues/149734,039f5b6de3772c3437c75fe16d742ee964b7484e54f273e1f4dfcfe73496f4ab references,issue,148682,issue,149982,medium,issue.comments[1].body,splitting the request for a fast dim1 kernel to a dedicated issue: #149982,https://github.com/pytorch/pytorch/issues/148682,34801018a63993065b8c3d45d2adab53d3bfb9e6536b9abbc384d8edba7f276f references,issue,94083,issue,31394,medium,issue.body,"Feel free to add further requests to this issue logsumexp #31394 indices of max and min #80439 #83980 composite reductions in pytorch_scatter (softmax, std, logsumexp)",https://github.com/pytorch/pytorch/issues/94083,cbb466db990adb01def09f2ce26a094f92003063089039dc16e805455ee954e7 references,issue,94083,issue,80439,medium,issue.body,"Feel free to add further requests to this issue logsumexp #31394 indices of max and min #80439 #83980 composite reductions in pytorch_scatter (softmax, std, logsumexp)",https://github.com/pytorch/pytorch/issues/94083,31c81b826a9764a7b0a1a04f56764159a0e0639b1d484efae784560425a506ef references,issue,94083,issue,83980,medium,issue.body,"Feel free to add further requests to this issue logsumexp #31394 indices of max and min #80439 #83980 composite reductions in pytorch_scatter (softmax, std, logsumexp)",https://github.com/pytorch/pytorch/issues/94083,4f27f92a560bfccdf1687a6355cb241a21c320de817f358fcca19cc4833174be references,issue,118421,issue,117748,medium,issue.body,"🚀 The feature, motivation and pitch Context: #117748 All-reduce comms are used in DDP's backward pass and by default the bucket size is set to 25MB via bucket_cap_mb. Documentation about this",https://github.com/pytorch/pytorch/issues/118421,c933c6ff0dcc38ee73675a4711b979af950f0c931df3f4f3dd980b42d35da745 references,issue,135704,issue,101699,medium,issue.body,"ing index arrays: batched varlen arange, or batched slice. This is related to: Standardized ways of representing/storing strings in tensor: #101699 repeat_interleave which can repeat with different lengths and concat batched arange (where start/end vary across the batch and ev...",https://github.com/pytorch/pytorch/issues/135704,2aaff39e51143e91b9884d71e503f89a3326ada5839586afeea418bffdb06636 references,issue,135704,issue,13246,medium,issue.comments[1].body,"f the underlying storage (as it is by default represented as a concatenated contig memory without gaps IIUC). It would also partially solve #13246 (which would be much alleviated if a standard way of representing array of strings as a tensor is added to core) In general, I thi...",https://github.com/pytorch/pytorch/issues/135704,1bf976793a50217ae9b6c8aa731a07968a8bed70e3c252a922f176b891aa0218 references,issue,148946,issue,148748,medium,issue.comments[0].body,This feels like a duplicate #148748,https://github.com/pytorch/pytorch/issues/148946,f4787332781bf14f152907163aba0c0dfde6e4c9dc875165b1f02e894245f79d references,issue,72759,issue,39242,medium,issue.body,api_docs/python/tfp/math/softplus_inverse Forum issue: https://discuss.pytorch.org/t/branching-for-numerical-stability/15763 Related issue: #39242 Alternatives import torch from torch.nn.functional import softplus from matplotlib import pyplot as plt from torch.autograd import...,https://github.com/pytorch/pytorch/issues/72759,ae8351c1eb29b4bb760030df1d1dba1735be260743d9ca0a5692686a3b6a01a1 references,issue,78065,issue,76324,medium,issue.body,operators. Enjoy! One of a five-part series of special functions issues: Gamma and Related Functions (#78065) Bessel and Related Functions (#76324) Orthogonal Polynomials (#80152) Elliptic Functions and Integrals (#80157) Operators barnes_g beta binomial_coefficient double_fac...,https://github.com/pytorch/pytorch/issues/78065,0f1a784dbecd3b56abf7be867b6988285acba17741217f298720c51a74963ab2 references,issue,78065,issue,80152,medium,issue.body,part series of special functions issues: Gamma and Related Functions (#78065) Bessel and Related Functions (#76324) Orthogonal Polynomials (#80152) Elliptic Functions and Integrals (#80157) Operators barnes_g beta binomial_coefficient double_factorial factorial falling_factori...,https://github.com/pytorch/pytorch/issues/78065,456be9a89c2a39028c8ee77c330fa2425451cb8adc1dfd1275cfaf7729be2c78 references,issue,78065,issue,80157,medium,issue.body,amma and Related Functions (#78065) Bessel and Related Functions (#76324) Orthogonal Polynomials (#80152) Elliptic Functions and Integrals (#80157) Operators barnes_g beta binomial_coefficient double_factorial factorial falling_factorial gamma log_barnes_g log_beta log_binomia...,https://github.com/pytorch/pytorch/issues/78065,4a5f45648c97a665ea456fb29254907b1ef473d45d020a6d7244e3812a2c8ded references,issue,141014,issue,138460,medium,issue.comments[0].body,#138460 seems to be a related issue cc @albanD,https://github.com/pytorch/pytorch/issues/141014,e54a30134eb0fc2d769667125a162b52fb93e6550404a7cb04b831fa96644036 references,issue,147628,issue,52743,medium,issue.comments[0].body,"al value to mean dynamic detection would be better... My old related issues on better native support for compact, non-int64 indices: #61819 #52743 (Also, https://pytorch.org/docs/stable/generated/torch.argsort.html maybe could support dtype argument similar to https://pytorch....",https://github.com/pytorch/pytorch/issues/147628,da52f7108eeeb82859cab7010c9183989ea5b7fc8de4c35bf1ff36d698d31f75 references,issue,147628,issue,61819,medium,issue.comments[0].body,"a special value to mean dynamic detection would be better... My old related issues on better native support for compact, non-int64 indices: #61819 #52743 (Also, https://pytorch.org/docs/stable/generated/torch.argsort.html maybe could support dtype argument similar to https://p...",https://github.com/pytorch/pytorch/issues/147628,6d76b84dce58a8640ac418d863da3974ccf6ac643af45b89c359706a986eaf64 references,issue,81104,issue,80549,medium,issue.comments[0].body,"resizing sparse dimensions and a solution I like this solution, especially considering the ongoing discussion about parameterized layouts (#80549). It seems that the idea to consider blocksize as part of the layout specification is accepted as a good idea (still questions abou...",https://github.com/pytorch/pytorch/issues/81104,1b9cdc4d073d00805fe9d0e8566db8d3309d2d49a4735006968ab5ef1236c3c8 references,issue,143670,issue,128071,medium,issue.body,"🐛 Describe the bug I encountered this while investigating recompilations in #128071. The relevant model code is here. Repro import torch @torch.compile(backend=""eager"") def f(x, int_dict, n): if n in int_dict: return x + 1",https://github.com/pytorch/pytorch/issues/143670,7e6b2d30ff3739896a5d874180e6cb33d26e13272e595c3f22d538d169bb787e references,issue,140437,issue,138786,medium,issue.body,"would take to get allow_in_graph to work in more situations. Here is an accounting of some situations which we thought of: NN module input #138786 Local functions that close over local variables Methods that take in some object, potentially temporary, as a self argument In gen...",https://github.com/pytorch/pytorch/issues/140437,d2de2bf7e573d9611cb2864aa82f6e003c15d18387be3aa35be5b06c35507881 references,issue,133230,issue,117394,medium,issue.comments[1].body,"There is a open RFC in #117394, my hand is a bit tight with other projects right now so this rfc is making very slow progress lately.",https://github.com/pytorch/pytorch/issues/133230,4e20118eb8981f0f2f3d8a8176fb48ab300415af2e6c297ac8eee86849fc8e31 references,issue,147692,issue,145949,medium,issue.comments[0].body,#145949,https://github.com/pytorch/pytorch/issues/147692,860aa3a89fad3787097acbdc3b7a027d1a81a5f2cbdcb4742b7871fde25a9313 references,issue,131667,issue,131650,medium,issue.body,ts about internal failures when MIG is used on Ampere+ devices if the device visibility is not used correctly e.g. in #130181 (comment) and #131650. In #92315 (comment) we already checked all error messages and should follow up on improving this error reporting. Especially rai...,https://github.com/pytorch/pytorch/issues/131667,550aa7fc318dbf3b06d9c5e74464757c5cbaff3141879511818959deb6a0cb8c references,issue,146980,issue,146414,medium,issue.comments[0].body,"e in core rather than extension that implements this one via TensorSubclassing. If it must be part of PyTorch, will it be a shell type (see #146414 that outlines what shell types are) or a fully functional one. If later, then it would be nice to better. understand what hardwar...",https://github.com/pytorch/pytorch/issues/146980,8472b9d03fdaa7af924721de7fe4bf90733f018f218bd0877a1b883f062d8c2d references,issue,146980,issue,146414,medium,issue.comments[1].body,"e in core rather than extension that implements this one via TensorSubclassing. If it must be part of PyTorch, will it be a shell type (see #146414 that outlines what shell types are) or a fully functional one. If later, then it would be nice to better. understand what hardwar...",https://github.com/pytorch/pytorch/issues/146980,28449b7bbd7c876ef6080e07e3a7d5dd2094ca820231670fa721323234ec4295 references,issue,107909,issue,71404,medium,issue.comments[0].body,Related: #71404 cc @mikaylagawarecki,https://github.com/pytorch/pytorch/issues/107909,739e3bf1a05e8efdd2973a38460ccde3db5fcc19cc785184454021eabadf9158 references,issue,147048,issue,13811,medium,issue.body,"hat exposes an .rsample() method is available in this repo. Moreover I can see that in the original Von Mises distribution feature request (#13811), comments are still stating that the distribution only works for 3D, if that's the case, then we really need an N-D implementatio...",https://github.com/pytorch/pytorch/issues/147048,4028792bace7c032d133beb44cf8c6c8ddc6d1ae99e2c9f62bb2744167733a4a references,issue,146951,issue,141258,medium,issue.comments[1].body,"doesn't work for np.bool_. And unfortunately np.bool_ hits this code path because NumpyNdarrayVariable is a subclass of TensorVariable, and #141258 (comment). pytorch/torch/_dynamo/variables/builtin.py Lines 2051 to 2061 in 0de27ee if op in [operator.is_, operator.is_not]: is_...",https://github.com/pytorch/pytorch/issues/146951,ec713c0b45d2dc9a9e2a3073da8410b67711eb8648c48e888aa1e4ab7a433779 references,issue,144012,issue,115852,medium,issue.comments[1].body,"Maybe related - for better UX / expressivity in frontend: #115852 UPD: seems here a different, unrelated concept of ""group"" is discussed",https://github.com/pytorch/pytorch/issues/144012,c46b528fd648ff50db6138f645a02b13f9ab074018f9be01d38ca46ee81cf810 references,issue,104702,issue,69364,medium,issue.body,"feima) For onednn it is: nChw16c nChw8c @jerryzh168 @vkuzo @ngimel this is probably also related to int8 gemm code paths as we discussed in #69364 (256 bit AVX2's equivalent is float32x8, and for AVX512 it's float32x16",https://github.com/pytorch/pytorch/issues/104702,c9ff598e4ceb5bcb504aadc58deffaebcda3542be00093a97919bbd5c2e911c4 references,issue,145687,issue,145709,medium,issue.comments[1].body,mers/models/llama/modeling_llama.py#L351-L352 The source of this confusion is the templated kernel naming scheme that I'll fix in #145711 & #145709.,https://github.com/pytorch/pytorch/issues/145687,fb9881c84d24774a91e7933567f5962da48ffdd1e7fcfcca3741e7f27bb36757 references,issue,145687,issue,145711,medium,issue.comments[1].body,c/transformers/models/llama/modeling_llama.py#L351-L352 The source of this confusion is the templated kernel naming scheme that I'll fix in #145711 & #145709.,https://github.com/pytorch/pytorch/issues/145687,49c183cba23af484a11dec0fc0d45f549e43be3459b44b88266155e4fd1f5e93 references,issue,48559,issue,44380,medium,issue.comments[0].body,Can you look whether Pytorch-uvm 1.7.1 supports your suggestion? #44380,https://github.com/pytorch/pytorch/issues/48559,e7c95138c67a9ff7c6bf51209fe1267cd882192afcbdf91a69a35515898421ca references,issue,144563,issue,126523,medium,issue.body,🐛 Describe the bug I noticed this while working on #126523 Basically the test suite runner run_test.py runs each test file separately or in parallel. It boils down to e.g. executing: python -bb dist,https://github.com/pytorch/pytorch/issues/144563,744c95f739ac2c2b958fd7c4973135676c6788d6740891bd747e8fee8a5576ff references,issue,134323,issue,70073,medium,issue.body,processing. This is a non-exhaustive compilation of issues: pytorch/audio/issues/427 pytorch/audio/issues/500 #62323 #81428 #91309 #118507 #70073 The importance of those issues is also reflected by the number of re-implementations of iSTFT in audio signal processing repositori...,https://github.com/pytorch/pytorch/issues/134323,8b87c4f03ad29243264b933247971421946aae72959d7f6d691a48a78c84913f references,issue,134323,issue,81428,medium,issue.body,ivated in audio signal processing. This is a non-exhaustive compilation of issues: pytorch/audio/issues/427 pytorch/audio/issues/500 #62323 #81428 #91309 #118507 #70073 The importance of those issues is also reflected by the number of re-implementations of iSTFT in audio signa...,https://github.com/pytorch/pytorch/issues/134323,b6a3145c8cb87641fb1c15f11ae8c9f280a3dd9327dbcb434d7c4fb8f625307d references,issue,134323,issue,91309,medium,issue.body,in audio signal processing. This is a non-exhaustive compilation of issues: pytorch/audio/issues/427 pytorch/audio/issues/500 #62323 #81428 #91309 #118507 #70073 The importance of those issues is also reflected by the number of re-implementations of iSTFT in audio signal proce...,https://github.com/pytorch/pytorch/issues/134323,c41df58a9f24e99c7521b2f98211131807e416f26f464ab36d0e3ab088ef38bd references,issue,134323,issue,118507,medium,issue.body,o signal processing. This is a non-exhaustive compilation of issues: pytorch/audio/issues/427 pytorch/audio/issues/500 #62323 #81428 #91309 #118507 #70073 The importance of those issues is also reflected by the number of re-implementations of iSTFT in audio signal processing r...,https://github.com/pytorch/pytorch/issues/134323,c2a6cdae79e0e2bc9f1c50dc63f03aa55d56429c73e6fac9a19063dcfe6114a5 references,issue,124950,issue,104506,medium,issue.comments[1].body,A related issue here: #104506,https://github.com/pytorch/pytorch/issues/124950,9465470a5e34d2f6d05a1b3f14e1acf6671b3864a6acbf15c4608d738c40600f references,issue,145624,issue,124505,medium,issue.comments[0].body,Maybe related to: #124505,https://github.com/pytorch/pytorch/issues/145624,693ea5058f6e2935c91adbb262e4d4ca241e817495fff38c8ec5f1ae30ef18dd references,issue,145233,issue,90760,medium,issue.body,".all() print(time.time() - start) with code commented: > 6 seconds with code uncommented: <0.1 seconds. I initially thought it is linked to #90760 , but that issue mentions only matmul and cholecky (also, because of that, I was initially searching for those operations), nothin...",https://github.com/pytorch/pytorch/issues/145233,2538ad46abeca1b217c5c17f709cf8798378257258c55f5ed71b0e57001baf48 references,issue,145014,pr,132135,medium,issue.body,"FX graph during tracing. If this step is missed, the pull request introducing the native function may encounter confusing CI failures(e.g. #132135). For example, a PR author intending to add a native function for eager mode might see numerous compile-related test failures, whi...",https://github.com/pytorch/pytorch/issues/145014,081f96921cae531641dbba8285f05c6005c91cb000be4ed8da33e7241347134e references,issue,129889,issue,50122,medium,issue.comments[0].body,aybe in general an STE helper could be introduced to torch.higher_order/torch.func ops or as some sort of pseudo-function (like proposed in #50122 - if possible at all to have an API like this - maybe this sort of pseudo-function could attach some hook instead? or a new kind o...,https://github.com/pytorch/pytorch/issues/129889,dc2a9395a99a30ad5d8b9c25af9a17f38fda9cbc98799b6f9c5fe08aced9c3ed references,issue,68559,issue,32976,medium,issue.body,"thon PyTorch. Environment PyTorch version 1.10.0+cu111 Additional context Potentially related issues: #17869 (pretty sure this blocks that) #32976 (not entirely sure it's related, but similar error) #61464",https://github.com/pytorch/pytorch/issues/68559,3e979ead50944875676cbcea2e49612b8a8636b5034cdfd76700339ca938f647 references,issue,143913,issue,142703,medium,issue.body,"sion greater than 15), which causes a slower bfloat16 gemv kernel to be used. We should have test coverage for CPU bfloat16 support on Mac (#142703) -- clang 16 purports to be able to build it, but is buggy and we actually need 17+. Alternatives do nothing until Apple gets aro...",https://github.com/pytorch/pytorch/issues/143913,408c41847b9425cfd4bad0e503034674ccfe0c17c058340a2dba39f32b4d2905 references,issue,141832,issue,116490,medium,issue.body,"rue) x = torch.full((1, 1), 1.0) y = 2 * x y_pred = model(x) loss = criterion(y_pred, y) loss.backward() optimizer.step() 📝 Related content #116490 https://discuss.pytorch.org/t/differentiable-optimizer-not-working-for-simple-example/207483 🔖 Versions System 1: PyTorch version...",https://github.com/pytorch/pytorch/issues/141832,b73561fcc983601425a14e68554a06530696c27dc6b0bd4124bf02772f78d5ff references,issue,144633,issue,110737,medium,issue.comments[0].body,Possibly related issue (with no resolution): #110737,https://github.com/pytorch/pytorch/issues/144633,759b666b691ca70bccf00f6ffce8d4918b78cc12f94f5fcefa408d01ac820bda references,issue,144673,issue,120819,medium,issue.comments[1].body,Possible duplicate of #120819?,https://github.com/pytorch/pytorch/issues/144673,24a095d95483c5b9f9659e31e011280df03b0fe7156379feb1bcc36757961dfb references,issue,102911,issue,94691,medium,issue.body,"kernel. This issue is related to Issue 96416 (Loss.backward() error when using MPS on M1 #96416), Issue 94691 (Nan is output by GRU on mps #94691), Issue 97552 (PackedSequence failure with MPS #97552), and PR 96601 ([MPS] LSTM grad_y missing fix #96601). To faithfully reproduc...",https://github.com/pytorch/pytorch/issues/102911,f24a6664eda10e8aa519391dce96848f5528cd5049a109b09368b8ad19ebb902 references,issue,102852,issue,77764,medium,issue.body,"y implemented for the MPS device. If you want this op to be added in priority during the prototype phase of this feature, please comment on #77764. As a temporary fix, you can set the environment variable PYTORCH_ENABLE_MPS_FALLBACK=1 to use the CPU as a fallback for this op....",https://github.com/pytorch/pytorch/issues/102852,cc149db853b31f3cb5dcd4cc26de1559d0b1ead30ac565a3f369018625a01a47 references,issue,111173,issue,97310,medium,issue.comments[0].body,Related performance problem with the same implementation: #97310,https://github.com/pytorch/pytorch/issues/111173,2c7147eff6cfd52d359c86902bbf7e3f0c6bef620020869ad3805bab48c7e9ec references,issue,113062,issue,77764,medium,issue.comments[0].body,"@albanD, not sure if you want to track these requests for MPS ops separately or all in #77764. I see that @AmineAndam04 already commented there.",https://github.com/pytorch/pytorch/issues/113062,a0bcca74f937dbc3c3a8cc737cc0ab0779338eebd8258273067a2f13e7f9b210 references,issue,84489,issue,77764,medium,issue.body,"t implemented for the MPS device. If you want this op to be added in priority during the prototype phase of this feature, please comment on #77764. As a temporary fix, you can set the environment variable PYTORCH_ENABLE_MPS_FALLBACK=1 to use the CPU as a fallback for this op....",https://github.com/pytorch/pytorch/issues/84489,c4245a30f7cda7625282fe14e563ea34fe47984d73e31ed86f55c229ca79c625 references,issue,136692,issue,77764,medium,issue.body,"y implemented for the MPS device. If you want this op to be added in priority during the prototype phase of this feature, please comment on #77764. As a temporary fix, you can set the environment variable PYTORCH_ENABLE_MPS_FALLBACK=1 to use the CPU as a fallback for this op....",https://github.com/pytorch/pytorch/issues/136692,dd8cdce432ef241945f306bc065a68e1befd74253d82e0b697d0017c6b2b6dd8 references,issue,144247,issue,139124,medium,issue.comments[1].body,"Tentatively grabbing for myself, as it somewhat related to #139124",https://github.com/pytorch/pytorch/issues/144247,d38e410be084e870daad362e7a6bc48abf5400f8777095d00fa25b291dad2549 references,issue,143658,pr,139524,medium,issue.comments[0].body,print the actual tensor data by default (unless there's an easy way to do it in a sync-free way?) as one thing we may want in the future is #139524. It also sounds like the shape would've been enough to debug the scenario motivating this issue.,https://github.com/pytorch/pytorch/issues/143658,27e590febd0678672147d89f0696cd6e7853b5967b067b737b429e05a06ada94 references,issue,138460,issue,131312,medium,issue.body,"s not contain the appropriate entries for the newly added nvjitlink library. This is the root cause of several user issues like #134929 and #131312 as far as I can tell. And I can also observe it locally, where running ldd on the libtorch_cuda.so that is shipped with the PyTor...",https://github.com/pytorch/pytorch/issues/138460,04a1f5bf09f18d629ec8256aec7ef5a94012be45c0b9e39fc4b7b990a26aba52 references,issue,132133,pr,132135,medium,issue.comments[0].body,You can use the feature with this command before #132135 is merged. docker pull voidbag/pytorch:2.4.0-cuda12.1-cudnn9-devel-wip-betainc-with-bwd-voidbag-v0.1.0 docker pull voidbag/pytorch:2.6.0a0-,https://github.com/pytorch/pytorch/issues/132133,591a670a5610837d0376ef4ce5e6df2bb540f69cac1b06b171ab44f13f1c4342 references,issue,141822,issue,141177,medium,issue.comments[0].body,Potentially related? I thought the issue is with the class indices implementation but perhaps it's not #141177,https://github.com/pytorch/pytorch/issues/141822,38d19ca4d9c7758ed3ce6f407b21ac5c66d054d4755efbc0a221e5f5d68d8edc references,issue,85773,issue,73487,medium,issue.body,"ppen when running Pytorch on the Windows that's hosting this WSL container. On Windows, everything works fine. This is a different bug from #73487. . They report Pytorch not finding Cuda, while what I've found is Pytorch crashing on a call to, e.g., a forward pass of a Conv2d...",https://github.com/pytorch/pytorch/issues/85773,f3dbf93971214290fe942d6a64a0e15b05df9bd185e8276193dbbddea6966bef references,issue,113750,issue,51455,medium,issue.body,"📚 The doc issue E.g. for various Norm layers (as discussed in #51455). I think this already done now for MHA/SDPA. Ideally, these code snippets would be auto-copied from the refs/decomps code. And if not, the",https://github.com/pytorch/pytorch/issues/113750,e76c3a54db0d6086a992d3b89ebcff2231fc7c4c822dbb947e7a365dfad370f5 references,issue,143588,issue,135826,medium,issue.comments[0].body,Related: #135826,https://github.com/pytorch/pytorch/issues/143588,9dcc57ae30c7d6b6d524cf3302144c092505e4314f70c8786fc231f499a8d2d8 references,issue,91524,issue,89491,medium,issue.comments[0].body,This report duplicates #89491.,https://github.com/pytorch/pytorch/issues/91524,1227c306304bf4a9aa2b6f303f8451df4b7d16aeb6851115bafb95bbb791d6b1 references,issue,139429,issue,126921,medium,issue.body,"ed if to() is invoked by the containing Module. Using the cast works fine for eager execution, but introduces issues with ONNX export ( see #126921) Additional context No response cc @albanD @mruberry @jbschlosser @walterddr @mikaylagawarecki",https://github.com/pytorch/pytorch/issues/139429,6070cdab18a9b614493b0f9a496869f0785634efe8f72abfbdac60d8d8d0792f references,issue,142392,issue,134323,medium,issue.comments[0].body,Is this potentially the same as #134323 ?,https://github.com/pytorch/pytorch/issues/142392,8006db47f873c18123b722727283260e4c3002fbf9c6be73469d5b264136a473 references,issue,143258,issue,75002,medium,issue.body,"ality). Of course that is pretty bad, because we are forced to lie to the type checker, sweeping bugs under the carpet. Possibly related to #75002 cc @EikanWang @jgong5 @wenzhe-nrv @sanchitintel",https://github.com/pytorch/pytorch/issues/143258,c63339f1bcec428aea368542b0e46c5ae72b31c5d88aee4cf136d7e202259a30 references,issue,89884,issue,72138,medium,issue.body,Provide a detailed API design for high-level PyTorch Tensor Parallelism API design. This is an evolvement of PyTorch Sharding introduced in #72138 and is directly built on top of DTensor proposed in #88838. We want users to only focus on how their modules to be distributed and...,https://github.com/pytorch/pytorch/issues/89884,030d745bf75d33417192fbb453f0a844b6cd3a802d006248a863d36c985ddf26 references,issue,89884,issue,88838,medium,issue.body,"Parallelism API design. This is an evolvement of PyTorch Sharding introduced in #72138 and is directly built on top of DTensor proposed in #88838. We want users to only focus on how their modules to be distributed and hide all other details. (Caveat, for now, we only support l...",https://github.com/pytorch/pytorch/issues/89884,1cfb164bdfee6baa032a93806b0f19baf8b4aad7480a65937f5dd2a6ee0de42f references,issue,139908,issue,141265,medium,issue.body,": 1.42, 1.68, 2.32 geglu: 1.00, 1.00, 1.00 jsd: 1.18, 3.38, 3.67 kl_div: 1.18, 1.21, 1.29 rms_norm: 0.90, 0.99, 1.29 rope: 0.41, 0.48, 0.58 #141265 swiglu: 0.99, 1.00, 1.00 fwd_bwd fp32 peak gpu memory usage: cross_entropy: 0.75, 0.75, 0.75 embedding: 0.77, 0.84, 0.93 fused_li...",https://github.com/pytorch/pytorch/issues/139908,2a9c31a6e53be87b92a6ddace753ff23bbe5759b83d85cbeb08126b4fc869975 references,issue,139908,issue,141916,medium,issue.body,"3.58 fused_linear_jsd: 2.90, 3.63, 5.18 geglu: 0.99, 1.00, 1.02 jsd: 2.96, 15.97, 17.37 kl_div: 0.90, 0.92, 0.97 rms_norm: 0.91, 0.96, 0.98 #141916 rope: 0.80, 0.90, 0.95 swiglu: 1.00, 1.00, 1.00 fwd fp32 peak gpu memory usage: cross_entropy: 1.00, 1.00, 1.00 embedding: 1.00,...",https://github.com/pytorch/pytorch/issues/139908,5737183d18af5ca1be6421ebe97ca71994f39fcff524f947248a7bf1c01dcb35 references,issue,139908,issue,142250,medium,issue.body,"1.01, 1.01 rope: 0.53, 0.54, 0.59 swiglu: 0.55, 0.74, 0.91 fwd fp32 latency: cross_entropy: 1.71, 2.28, 2.28 embedding, (0.83, 0.95, 1.00) #142250 fused_linear_cross_entropy: 3.07, 3.25, 3.58 fused_linear_jsd: 2.90, 3.63, 5.18 geglu: 0.99, 1.00, 1.02 jsd: 2.96, 15.97, 17.37 kl...",https://github.com/pytorch/pytorch/issues/139908,79ec6d3573b9a9ffeb6275c550b09ac6c7adbbce2979b2f57e900657fdc8b9c8 references,issue,114153,issue,132041,medium,issue.body,etermine whether a fix is needed or not. The following lists summarize the status for these individual issues. Needs Investigation: #132031 #132041 #132045 #132047 Needs Fix: #142466 - behavior with out_channels=0 should be consistent across devices Resolved: #120875 #114081 #...,https://github.com/pytorch/pytorch/issues/114153,427edb0bde2d20ae2b49e926b5dd962b85c15b5a0d5c4a7da48c1b140f60f241 references,issue,114153,issue,132045,medium,issue.body,whether a fix is needed or not. The following lists summarize the status for these individual issues. Needs Investigation: #132031 #132041 #132045 #132047 Needs Fix: #142466 - behavior with out_channels=0 should be consistent across devices Resolved: #120875 #114081 #114085 #1...,https://github.com/pytorch/pytorch/issues/114153,55b6a879d41eab6301016f36a08d339ee8022c9d5224413217768db3d13730d5 references,issue,114153,issue,132047,medium,issue.body,a fix is needed or not. The following lists summarize the status for these individual issues. Needs Investigation: #132031 #132041 #132045 #132047 Needs Fix: #142466 - behavior with out_channels=0 should be consistent across devices Resolved: #120875 #114081 #114085 #114052 #1...,https://github.com/pytorch/pytorch/issues/114153,045e03987855fb140eb4ff153d84eaa34de7280b85a4f89eaf98a8c4385daa11 references,issue,140222,issue,140252,medium,issue.comments[1].body,@mingfeima in suspect it's similar to #140252 i.e. CUDA probably does computation/accumulation on full precision floats while CPU keeps it in bf16. Perhaps good first step would be to u,https://github.com/pytorch/pytorch/issues/140222,17ab4006acbfa6953e5d78816e964fa8add9f4edd9479132ba857558d06f95af references,issue,127521,issue,69364,medium,issue.comments[0].body,Related: #69364,https://github.com/pytorch/pytorch/issues/127521,696a5eba86e29db8b185de083eb321b45b5e6201f1a2dc04a11cd5e8ec451414 references,issue,139643,issue,141079,medium,issue.comments[1].body,"for reference: iirc, this is the same as #141079",https://github.com/pytorch/pytorch/issues/139643,9ecace5c0914f217a342e2da0a3cd249cd03c46c007d132b76ce8ad0b59faca1 references,issue,141764,issue,138696,medium,issue.body,Subtask of #138696: [CI/CD] Deprecating PyTorch’s official Anaconda channel. cc @seemethere @malfet @pytorch/pytorch-dev-infra,https://github.com/pytorch/pytorch/issues/141764,9873014ace1ef835bebdc42c8cd25c1df0f7d834a11dfd985e60749a358249ba references,issue,141831,issue,129118,medium,issue.body,"1].abs().sum()) tensor(0., device='cuda:0', dtype=torch.bfloat16) Workaround Make the input tensor contiguous, as in #81665. Related issues #129118 #81665 Additional considerations The fact that the failure is silent makes this hard to identify, track and debug. The CPU implem...",https://github.com/pytorch/pytorch/issues/141831,92a32641158eeeb13120b247378b3b651bc4b2fcbd252e2e77552d4192693059 references,issue,78109,issue,69353,medium,issue.body,"ttention, it will never be called. Which may cause some abnormal behaviors when use tools implemented with hook, e.g. torch.nn.utils.prune (#69353). And this is because when you forward a MultiheadAttention, it directly call the out_proj.weight without forwarding it in torch._...",https://github.com/pytorch/pytorch/issues/78109,73b0d470cb286870cb326307e5c4560e13c6b0866e2d5e13a2ffdb9ddea7aaf2 references,issue,124931,issue,32590,medium,issue.body,"📚 The doc issue nn.TransformerDecoder seems to support 3D (N*nhead, T, S) memory_mask (since #32590 (comment)), and the description of nn.MultiheadAttention has been updated correctly. But the description of nn.Transformer argument memory_",https://github.com/pytorch/pytorch/issues/124931,5f60209a1b56e554d9e1a1c72fd06e9c303842d1ea362840b8949818f54b21c9 references,issue,113713,issue,58414,medium,issue.body,"ed by the yet-to-be finalized cuDNN v1.0 frontend. It also differs from previous cuDNN frontend integrations, such as for convolution (see ##58414), in that the flash-attention implementation requires some (lightweight) runtime compilation. Of course, these will be cached afte...",https://github.com/pytorch/pytorch/issues/113713,fd691aa525e89cb892e7dc5f87c003993d7ea660065370ba2bbe440fec6fc2ac references,issue,141219,issue,141218,medium,issue.body,"aled_dot_product_flash_attention_for_cpu triggered a crash. By analyzing the results of ASAN, I think it may be different from the cause of #141218 import torch query = torch.full((7,9,0,7,), 0, dtype=torch.half) key = torch.full((7,7,0,7,), 0, dtype=torch.half) value = torch....",https://github.com/pytorch/pytorch/issues/141219,3d80f503413330c6d40a0f89a07895fbf92644c43ecf4019a0e38fadf38b98b8 references,issue,110681,issue,108108,medium,issue.body,"PyTorch to upgrade its FlashAttention kernel to the newest Implementation. There is currently some limitations regarding BC concerns, see: #108108. And user motivation for this change: #110144 Proposed Implementation: The foundational changes to the C++ core are outlined in th...",https://github.com/pytorch/pytorch/issues/110681,ab74fc7a1120a54cfdbc12ac379fad20d7fe254cd441edb6255bbe02c99bf4d7 references,issue,110681,issue,110702,medium,issue.body,e structured and if the proposed API is versatile enough to encompass them. Addendum A related issue to AttnBias variant can be found here: #110702 Which tensor subclass? When attempting to pass a dispatch tensor subclass as an attn_bias to torch.scaled_dot_product_attention()...,https://github.com/pytorch/pytorch/issues/110681,77fd833dbe386217c8e3943b47973dac171b05249f3677d59d3cbc9b1aaf6533 references,issue,122660,issue,94428,medium,issue.body,"2 Definition of attn_mask is flipped compared to SDPA/implementation in other frameworks #120668 #121193 Performance related issues #116175 #94428 Bugs related to the ""fast path"" #116546 #107084 We also recognize that if these modules are to be deprecated, we would need to pro...",https://github.com/pytorch/pytorch/issues/122660,8c296d825b72b32269ddaf4b2a4868ee3aeae620ea89c767704c9932f6b52473 references,issue,122660,issue,116175,medium,issue.body,"s #118972 Definition of attn_mask is flipped compared to SDPA/implementation in other frameworks #120668 #121193 Performance related issues #116175 #94428 Bugs related to the ""fast path"" #116546 #107084 We also recognize that if these modules are to be deprecated, we would nee...",https://github.com/pytorch/pytorch/issues/122660,05a4df7ddba11de8262d7f435691237c644e9593f002d00b0f9cbf0f6ab1a585 references,issue,122660,issue,118972,medium,issue.body,have also been several issues surrounding these modules. Examples of issues around these modules include Confusion around various arguments #118972 Definition of attn_mask is flipped compared to SDPA/implementation in other frameworks #120668 #121193 Performance related issues...,https://github.com/pytorch/pytorch/issues/122660,2c32dd985965e0d97823dda16ab9decd8faf3738de1233f3d8e6a31879f0801f references,issue,122660,issue,121193,medium,issue.body,"e Confusion around various arguments #118972 Definition of attn_mask is flipped compared to SDPA/implementation in other frameworks #120668 #121193 Performance related issues #116175 #94428 Bugs related to the ""fast path"" #116546 #107084 We also recognize that if these modules...",https://github.com/pytorch/pytorch/issues/122660,e90c89b7e7112bbc8bacaf0fb2b6d8ef56982f7cff12f5b052aa55611fb0818e references,issue,74092,issue,24930,medium,issue.comments[1].body,Is this a duplicate of #24930?,https://github.com/pytorch/pytorch/issues/74092,6d0154197b2afaa66dd2398293073216671605b7c49809e511111cdd4a91039c references,issue,50804,issue,21018,medium,issue.comments[1].body,I just cannot reproduce on my side for v1.7.1. Not sure if this one helps or not. #21018,https://github.com/pytorch/pytorch/issues/50804,11ba3a1d696de140ceb2a6dde4551b1c9a9e92b7dd5bcda680460bfb3e7359b5 references,issue,40932,issue,41508,medium,issue.comments[0].body,Related to #41508. I believe this will also be solved by usage of safe softmax within MHA.,https://github.com/pytorch/pytorch/issues/40932,fadf7670dafc05452e8b8057b802f89dac15a711e99d784dc1501d01fddcb1c0 references,issue,40932,issue,41508,medium,issue.comments[1].body,"Hi @HelenaHlz, as directed by @jbschlosser please refer to the documentation given for this issue at this link which is located in #41508.",https://github.com/pytorch/pytorch/issues/40932,d21c78792bc2cac367822361ffc5a6b6fee71a1f807cf9179fe451fcaaa098c8 references,issue,34573,issue,32590,medium,issue.body,"dings or attention mechanisms without having to recode the rest. Motivation Addresses the issue of decomposing the function as mentioned in #32590. It also moves forward on including more support for attention mechanisms. Pitch Currently, the mutli_head_attention_forward funct...",https://github.com/pytorch/pytorch/issues/34573,ac05840940d10b11b163cee7c8133c1680716bff914303b88235488322481042 references,issue,32590,issue,21876,medium,issue.body,"lity to change the norm func in TransformerEncoderLayer and TransformerDecoderLayer #26342. The order of layer normalization in Transformer #21876, #24930. Users has the flexibility to pass custom encoder/decoder layers to the transformer modules. Update the doc to explain it...",https://github.com/pytorch/pytorch/issues/32590,39695f25cffcce3c570d70383af3ce25412d5a7ff17103d5c36dd423f64dcc2b references,issue,32590,issue,24826,medium,issue.body,"on layers (e.g. AugmentedConv, StandAloneAttention) - need research #32529 Add embedding layer and positional encoding layer in Transformer #24826",https://github.com/pytorch/pytorch/issues/32590,05fe20cde2124870e4d10c8076a5c4f0289e1c603d9a5262602825c569e8a5cc references,issue,32590,issue,24930,medium,issue.body,"change the norm func in TransformerEncoderLayer and TransformerDecoderLayer #26342. The order of layer normalization in Transformer #21876, #24930. Users has the flexibility to pass custom encoder/decoder layers to the transformer modules. Update the doc to explain it #32373 F...",https://github.com/pytorch/pytorch/issues/32590,a73d8f188924052c49e6f0d1c878b701972dfe8b9a78a03563873bf837b203c5 competes with,issue,32590,issue,25132,medium,issue.body,"nches. pytorch/text#720 Fix the dimension convention for masks in Transformer (Propose to use (N, S, E), instead of (S, N, E). Requested by #25132, #25100. PR #37597 Resolved issues Support ByteTensor for key_padding_mask and attn_mask in nn.MultiheadAttention. Requested by Fa...",https://github.com/pytorch/pytorch/issues/32590,89ef99441bbc2cf112671d2fdd6d148105f2a35cd9ec3db062d69e1fcfff61bc references,issue,32590,issue,28657,medium,issue.body,"ard functional - requested by Fairseq Move the func generate_square_subsequent_mask out of the transformer module #25538 Doc update #28719, #28657 Additional attention layers (e.g. AugmentedConv, StandAloneAttention) - need research #32529 Add embedding layer and positional en...",https://github.com/pytorch/pytorch/issues/32590,080d0d8ccdcd15b329ab974a8fae9821393663c646dfc5c8e5e35985c126b68e references,issue,32590,issue,32529,medium,issue.body,"e transformer module #25538 Doc update #28719, #28657 Additional attention layers (e.g. AugmentedConv, StandAloneAttention) - need research #32529 Add embedding layer and positional encoding layer in Transformer #24826",https://github.com/pytorch/pytorch/issues/32590,7e156314aea298c23953fbd1dbbfa7e209f22d11d848604c6c7499c985e9b48a references,issue,138399,issue,138396,medium,issue.body,", but failed with meta inputs Except for mean (which is an actual bug), all the other operations present the same behavior as identified by #138396. geqrf mean nanmean Dynamic Shape Outputs Similar to #138396, this operation outputs tensors of dynamic shape. Thus, there's no w...",https://github.com/pytorch/pytorch/issues/138399,99457362c351fcfd5ea351448dc2c044e75f7bf9189a35a7ab9720d861983f16 references,issue,138217,issue,117394,medium,issue.comments[1].body,Might be related: #117394,https://github.com/pytorch/pytorch/issues/138217,e02316e773a7fcd200ac3ee1ccd4c717b54fb942c3cb73e77631f3ee770f78af references,issue,120003,issue,113045,medium,issue.body,"vice initialization / _apply() methods Support initial meta-device initialization using swap_tensors path Remove manual padding logic after #113045 @wz337 Outcome: Once we have DTensor manage the padded storage, then FSDP only needs to maintain a reference to the DTensor, not...",https://github.com/pytorch/pytorch/issues/120003,5351af78646a594849e678126b9ba3e2e031baf07c3117b74a1685c357e9fc67 references,issue,120003,issue,114299,medium,issue.body,tensions for float8_experimental @awgu (dynamic scaling eager done) Add pre/post all-gather extensions for QLoRA @weifengpy References RFC: #114299,https://github.com/pytorch/pytorch/issues/120003,3afcf7bc113193a3dd432c9e9fc4c7a558552bf4183e3df70e853758ab93a43a references,issue,140879,issue,140563,medium,issue.body,"Tried to analyze the output from #140563 and ran into a number of issues interestingly i had a fr_trace in /usr/local/bin/fr_trace on my devserver. not sure how it got there, but i",https://github.com/pytorch/pytorch/issues/140879,0a878f6fb7388e1abaa83ab4d746e38af85fe32e29ec311c51b49e07b0b37b2e references,issue,140706,issue,136003,medium,issue.comments[1].body,"@malfet What's the proper way to do it? Do you have a code example? I'd love to know, since it could be relevant for #136003 and #124850. I'd like to start working on those again soon and accurate benchmarks would be very helpful while debugging.",https://github.com/pytorch/pytorch/issues/140706,0a3de811eaa924b8307e3e4c73fffc16d6f156f4e84b297612d727cdda347c0c references,issue,140845,issue,91439,medium,issue.comments[0].body,Related issues: #91439 #115092,https://github.com/pytorch/pytorch/issues/140845,4edc424c2bc008ad420e7fc1a141a30ccb6c5554029dbe826d83a0e9661ec94f references,issue,139629,issue,139521,medium,issue.body,"y more (AOTDispatcher handles both of these). Some of these might be relatively expensive (from an overhead perspective), especially due to #139521 and registrations from Python torch.library.custom_op. cc @ezyang @chauhang @penguinwu @voznesenskym @EikanWang @jgong5 @Guobing-...",https://github.com/pytorch/pytorch/issues/139629,fe61dc7572507a8050ee07006c48ddc60ef40ee3babcf611415c972bfb5f3784 references,issue,100741,issue,36577,medium,issue.body,"ch::jit::Unpickler::readInstruction() + 12824 (0x294f31cf0 in libtorch_cpu.dylib) frame #3: torch: AFAICT this is known see #67902, #37213, #36577 and #95362 model.state_dict() returns an OrderedDict and Libtorch's bundled unpickler doesn't support reading from it. Calling dic...",https://github.com/pytorch/pytorch/issues/100741,523d690c692a4faf35a88dbcde0c3dd83c9aa410dde3dcb11cb093ca02c816d0 references,issue,100741,issue,67902,medium,issue.body,"b) frame #2: torch::jit::Unpickler::readInstruction() + 12824 (0x294f31cf0 in libtorch_cpu.dylib) frame #3: torch: AFAICT this is known see #67902, #37213, #36577 and #95362 model.state_dict() returns an OrderedDict and Libtorch's bundled unpickler doesn't support reading from...",https://github.com/pytorch/pytorch/issues/100741,93a5691f961658be2f6dc7375c5aa10ce7d9645fd9ff578106246011a0693eae references,issue,140090,issue,125837,medium,issue.comments[0].body,"This sounds very similar to #125837 And I have a similar question as asked in that issue: is there an example of an Android wheel published on PiPY right now? Also, can you pl",https://github.com/pytorch/pytorch/issues/140090,38ac8f8ddaf1b4510ff8ad6386fbbd3d67197d3de590a0d88ddfe1e315055c43 references,issue,96726,issue,35600,medium,issue.comments[0].body,"Originally reported in #35600, but this issue's description supplements the debugging info present in related issues, so not marking it as a duplicate, but simply linkin",https://github.com/pytorch/pytorch/issues/96726,56f47ff5e2224a9785371ee9281d5278181d95924d4471f9e2430309b99af69a references,issue,106455,issue,40208,medium,issue.body,"orld (by which I mean communities like numerical methods for solving partial differential equations). Some potentially related open issues: #40208 pytorch/functorch#767 We might be able to have some one take a stab at this next FY, and hopefully submit a PR. I would like to se...",https://github.com/pytorch/pytorch/issues/106455,47e2e7cab01220d08c898a084190528226e8a6ba119564d5c94dfd7570b63b12 references,issue,35527,issue,31779,medium,issue.body,"but expected one of: * (Tensor input, torch.Generator generator, Tensor out) * (Tensor input, float p, torch.Generator generator) Related: #31779 cc @vincentqb @fritzo @neerajprad @alicanb @vishwakftw",https://github.com/pytorch/pytorch/issues/35527,77ca64ba52d9ecc66d5802e1f51f0a85ed96fd86d8fedf3bed05ec960ce069de references,issue,78262,issue,40761,medium,issue.comments[0].body,"discussion post here as well. @nirzaa there are currently three open issues on sparse convolutions, which I assume is what's missing here. #40761 #47915 #64544 Are any of these relevant to your case? Further, can you detail your use case a bit more and the model you want to us...",https://github.com/pytorch/pytorch/issues/78262,460da37cbd75ca9d7e8339022ef0fb03c61d3f1d076b19db7cac8e69d6359120 references,issue,78262,issue,47915,medium,issue.comments[0].body,"sion post here as well. @nirzaa there are currently three open issues on sparse convolutions, which I assume is what's missing here. #40761 #47915 #64544 Are any of these relevant to your case? Further, can you detail your use case a bit more and the model you want to use? Tha...",https://github.com/pytorch/pytorch/issues/78262,a146635da0e553e83746b5849ea2f4d7444503f5515f90b430e5bbb00a5e17be references,issue,78262,issue,64544,medium,issue.comments[0].body,"st here as well. @nirzaa there are currently three open issues on sparse convolutions, which I assume is what's missing here. #40761 #47915 #64544 Are any of these relevant to your case? Further, can you detail your use case a bit more and the model you want to use? Thanks, Ch...",https://github.com/pytorch/pytorch/issues/78262,a5ae3e5c696f3f182a6c9f51dcc021f0199df8485b673bb73ac2d09847a8f6f6 references,issue,138221,issue,138220,medium,issue.body,lacing. The main thing to do to get auto_functionalized_v2 to work with export is to develop a BC serialization scheme for it. We should do #138220 as a part of that work so that we don't need to break BC in the future when we implement that. cc @ezyang @chauhang @penguinwu @a...,https://github.com/pytorch/pytorch/issues/138221,6a2fee1b829afd68e4ed153c9f6af1be68b338e29f4e1f11eae8e11865ea07f7 references,issue,95108,issue,88658,medium,issue.body,"crash when using torch.bfloat16 dtype, but the same codes run well with torch.float32. I think this issue may have a similar root cause to #88658. In addition, both torch.nn.Linear and torch.nn.LazyLinear have other interesting errors in Pytorch 1.10.0 and 1.11.0. When they ar...",https://github.com/pytorch/pytorch/issues/95108,f6168ff4e2a07c9022cec626162e2ea8ee7797e7f641736ca53526b11f640cef references,issue,111003,issue,94652,medium,issue.comments[1].body,Not sure if this is related: #94652,https://github.com/pytorch/pytorch/issues/111003,60f1d644e7c245ab291d23e0711f8b9e94e66d91aecce23e97b05f0a1299043b references,issue,137482,issue,42246,medium,issue.comments[0].body,"I don't think #42246 was fixed, but sure, if you have a fix in mind, please do not hesitate to propose a PR.",https://github.com/pytorch/pytorch/issues/137482,d532a97dad25be616bb7bebd286a5ab540e3ca3ad2a0fc9b330075903da2a4c9 references,issue,137367,issue,137373,medium,issue.comments[1].body,See my comment here: #137373 (comment),https://github.com/pytorch/pytorch/issues/137367,5d24531fdc7eb7ed4047ccec6cec4837f572bfbabb4c059696ed285a0f8d99f5 references,issue,137564,issue,40770,medium,issue.body,"especially one that leverages GPU acceleration. I think this feature would benefit many users as I also found this existing issue about it: #40770 Pitch: Memory efficiency: Allows processing of datasets larger than available RAM, addressing a common limitation in data analysis...",https://github.com/pytorch/pytorch/issues/137564,ca250126d55266eeaf3f7c3e110ef565df493a906d7e96ac5180289cc3169ba6 references,issue,133823,issue,73008,medium,issue.body,"an abbreviated blame history of MKL_ROOT usage in cmake/public/mkl.cmake: Originally introduced in commit 7b0d577 (#89359), mentions issue #73008 pytorch/cmake/public/mkl.cmake Lines 13 to 17 in 7b0d577 # TODO: This is a hack, it will not pick up architecture dependent # MKL l...",https://github.com/pytorch/pytorch/issues/133823,90bc0b8613da0546a931d4188fbc3d44674eb608222a94e9924052379eea2b4c references,issue,131452,issue,137252,medium,issue.comments[0].body,same issue: #137252,https://github.com/pytorch/pytorch/issues/131452,a4be45186f7bc0401b163cf1dac58d767465554aa8e14675966eb44e977a4113 references,issue,135927,issue,124245,medium,issue.body,"🐛 Describe the bug I have enabled Windows inductor on main branch and release/2.5 branch now. We also get good quality in models passrate: #124245 (comment) But we still have a problem that, we can't enable Windows inductor UTs, due to it always timeout by 210 minutes. Actuall...",https://github.com/pytorch/pytorch/issues/135927,65b0108b3900d406442050d8bb21c4dad20190dd06b3f5fa79355f9a4c0e883d references,issue,107298,issue,75097,medium,issue.comments[0].body,Possibly related: #75097,https://github.com/pytorch/pytorch/issues/107298,f6d54bddcf224522a2bbd833cba3558d38e4a345aa75c88c877ca31436b3e62c references,issue,107298,issue,75097,medium,issue.comments[1].body,"Hi, @awgu. I don't think it is related to #75097 . My program is not hanging.",https://github.com/pytorch/pytorch/issues/107298,aa86a75956f7acb26a3376fd4bc10989668d9dc7018f0929f434f921823bdfc5 references,issue,135724,issue,124505,medium,issue.comments[1].body,"Also, might be nice to be able to have a natural way for controlling in a local way Inductor properties: #124505",https://github.com/pytorch/pytorch/issues/135724,d123a51d31c58daee65df083ea029eab180b441bb337ce97e064854f64a85bfa references,issue,50341,issue,5565,medium,issue.body,"functional.Bilinear, #48389 add more modes to torch.nn.functional.grid_sample, #25039 review torch.nn.functional.grid_sample documentation, #5565 Bugs: torch.nn.functional.grid_sample downsampling doesn't always match torch.nn.functional.interpolate downsampling, #21457 torch....",https://github.com/pytorch/pytorch/issues/50341,d2294855fda05a6fc2cb785854aa3bd768c0e99deb2b01084080605f3160f52d references,issue,50341,issue,15386,medium,issue.body,"ampling, #21457 torch.nn.functional.interpolate behaves unexpectedly with output size 1, #30565 nearest interpolation is misaligned, #34808 #15386 cc @mruberry @rgommers @heitorschueroff @ppwwyyxx",https://github.com/pytorch/pytorch/issues/50341,e99fa79a38d6fea0e62d6cfdf716924d85978f3fd30357085c0bda47b4ff83d8 references,issue,50341,issue,21457,medium,issue.body,"documentation, #5565 Bugs: torch.nn.functional.grid_sample downsampling doesn't always match torch.nn.functional.interpolate downsampling, #21457 torch.nn.functional.interpolate behaves unexpectedly with output size 1, #30565 nearest interpolation is misaligned, #34808 #15386...",https://github.com/pytorch/pytorch/issues/50341,26c6c6ffa641dd681b6e5e9efd6f1dd0be0ceb838ba2eef020458c5d8389dedb references,issue,50341,issue,24870,medium,issue.body,"be reviewed; we would not accept PRs implementing these requests at this time. Operator improvements and extensions improved grid sampling, #24870 ""clean-fid"", https://github.com/GaParmar/clean-fid Operator requests: #50334 #50335 #50340 #50337 #50338 #50339 #50336 #61528 Addi...",https://github.com/pytorch/pytorch/issues/50341,a7dd56cc896ab7704a7dfbc8e4dfb06816bb331c0ee769dd4a679cd034e18d2a references,issue,50341,issue,25039,medium,issue.body,"Additional mode requests: add half_pixel_centers to torch.nn.functional.Bilinear, #48389 add more modes to torch.nn.functional.grid_sample, #25039 review torch.nn.functional.grid_sample documentation, #5565 Bugs: torch.nn.functional.grid_sample downsampling doesn't always matc...",https://github.com/pytorch/pytorch/issues/50341,4d481d48c5a7367fe39aa177d43a2f84183955fb6fc74141b57dc602fe15905a references,issue,50341,issue,30565,medium,issue.body,"always match torch.nn.functional.interpolate downsampling, #21457 torch.nn.functional.interpolate behaves unexpectedly with output size 1, #30565 nearest interpolation is misaligned, #34808 #15386 cc @mruberry @rgommers @heitorschueroff @ppwwyyxx",https://github.com/pytorch/pytorch/issues/50341,5cde6950fea54e0f8dc07ebffd68c9ca6fe7056fa04307a6f5e991b3a1ea7637 references,issue,50341,issue,48389,medium,issue.body,": #50334 #50335 #50340 #50337 #50338 #50339 #50336 #61528 Additional mode requests: add half_pixel_centers to torch.nn.functional.Bilinear, #48389 add more modes to torch.nn.functional.grid_sample, #25039 review torch.nn.functional.grid_sample documentation, #5565 Bugs: torch....",https://github.com/pytorch/pytorch/issues/50341,4d0494278045c39ef5e6e5a668bcb38549c7e43d137c4501e0763debbf5af7df references,issue,50341,issue,50334,medium,issue.body,". Operator improvements and extensions improved grid sampling, #24870 ""clean-fid"", https://github.com/GaParmar/clean-fid Operator requests: #50334 #50335 #50340 #50337 #50338 #50339 #50336 #61528 Additional mode requests: add half_pixel_centers to torch.nn.functional.Bilinear,...",https://github.com/pytorch/pytorch/issues/50341,ec4ab4c9b9862f8f54352fdcaf55ea861b3a5fcf0c6dbd4c73d399a3dc9e0c7f references,issue,50341,issue,50335,medium,issue.body,"tor improvements and extensions improved grid sampling, #24870 ""clean-fid"", https://github.com/GaParmar/clean-fid Operator requests: #50334 #50335 #50340 #50337 #50338 #50339 #50336 #61528 Additional mode requests: add half_pixel_centers to torch.nn.functional.Bilinear, #48389...",https://github.com/pytorch/pytorch/issues/50341,985c33a02d9f45eb10e4ff47eff73d7bf821040e7b42c2ef81fda254552eb5af references,issue,50341,issue,50336,medium,issue.body,"roved grid sampling, #24870 ""clean-fid"", https://github.com/GaParmar/clean-fid Operator requests: #50334 #50335 #50340 #50337 #50338 #50339 #50336 #61528 Additional mode requests: add half_pixel_centers to torch.nn.functional.Bilinear, #48389 add more modes to torch.nn.functio...",https://github.com/pytorch/pytorch/issues/50341,b7e46b7fb0e12524b6d0f062d27f37aa1d4cfb83366ad62287df2568c530c5ea references,issue,50341,issue,50337,medium,issue.body,"ts and extensions improved grid sampling, #24870 ""clean-fid"", https://github.com/GaParmar/clean-fid Operator requests: #50334 #50335 #50340 #50337 #50338 #50339 #50336 #61528 Additional mode requests: add half_pixel_centers to torch.nn.functional.Bilinear, #48389 add more mode...",https://github.com/pytorch/pytorch/issues/50341,3a42f3cc1efdc6e27b92d36b4fc36a55c0800c946043aec9bdd13197c4c643b5 references,issue,50341,issue,50338,medium,issue.body,"extensions improved grid sampling, #24870 ""clean-fid"", https://github.com/GaParmar/clean-fid Operator requests: #50334 #50335 #50340 #50337 #50338 #50339 #50336 #61528 Additional mode requests: add half_pixel_centers to torch.nn.functional.Bilinear, #48389 add more modes to to...",https://github.com/pytorch/pytorch/issues/50341,8e8027a4f66d5e0c8e3ea5dde12a72a3ba5666ba53cd6d77822c43f19ac2f2bf references,issue,50341,issue,50339,medium,issue.body,"ons improved grid sampling, #24870 ""clean-fid"", https://github.com/GaParmar/clean-fid Operator requests: #50334 #50335 #50340 #50337 #50338 #50339 #50336 #61528 Additional mode requests: add half_pixel_centers to torch.nn.functional.Bilinear, #48389 add more modes to torch.nn....",https://github.com/pytorch/pytorch/issues/50341,aee4d0f4d101dc24aa5661a971b1dd039dbe436727420ab6b04f51a70a29d61f references,issue,50341,issue,50340,medium,issue.body,"rovements and extensions improved grid sampling, #24870 ""clean-fid"", https://github.com/GaParmar/clean-fid Operator requests: #50334 #50335 #50340 #50337 #50338 #50339 #50336 #61528 Additional mode requests: add half_pixel_centers to torch.nn.functional.Bilinear, #48389 add mo...",https://github.com/pytorch/pytorch/issues/50341,5b8783ba1d64ec745e6c4f1fa5ff9519b141c5bb6f5747a7863691b4fcf8dc43 references,issue,50341,issue,61528,medium,issue.body,"rid sampling, #24870 ""clean-fid"", https://github.com/GaParmar/clean-fid Operator requests: #50334 #50335 #50340 #50337 #50338 #50339 #50336 #61528 Additional mode requests: add half_pixel_centers to torch.nn.functional.Bilinear, #48389 add more modes to torch.nn.functional.gri...",https://github.com/pytorch/pytorch/issues/50341,7febb6b98d0ec8751131b818c1eff14cab9ccbecb6c54ae3353c8bdfcfe3d057 references,issue,98675,issue,14489,medium,issue.comments[1].body,"#14489 is related, but seems to be about COO format.",https://github.com/pytorch/pytorch/issues/98675,ca064a98e839c12711bb4ede272e71a5642bd66bab20db3785d2725b10bcba0b references,issue,135826,issue,55279,medium,issue.body,"ell-factored code, tests or documentation. Anyone willing to help to add this feature, starting from my code? This is partially related to: #55279, #80553 Alternatives I considered using pytorch-minimize, but BFGS is not efficiently implemented and Hager-Zhang line-search is n...",https://github.com/pytorch/pytorch/issues/135826,6e0a244d7bfe8a47ae1e8a12543f85cf6c5df90e30b39a9e870e8d9be208a64b references,issue,135826,issue,80553,medium,issue.body,"ored code, tests or documentation. Anyone willing to help to add this feature, starting from my code? This is partially related to: #55279, #80553 Alternatives I considered using pytorch-minimize, but BFGS is not efficiently implemented and Hager-Zhang line-search is not avail...",https://github.com/pytorch/pytorch/issues/135826,33e5838966852e3435b41d2d97d3f6c29a20f7e12e197c362bcf5fbc536942b4 references,issue,118798,issue,118492,medium,issue.body,"🐛 Describe the bug Looking at #118492 and #118213 it occurs to me that if we are generating hundreds of guards involving a single symbolic variable, we probably should just give",https://github.com/pytorch/pytorch/issues/118798,63bf3ad530b2121432a2d2b571faab8c497fb38021458631cc8082c2050cd657 references,issue,135359,issue,129657,medium,issue.comments[0].body,"Also, maybe regardless of problems in this issue, having a stable log(1-sigmoid(x)) would also be nice (for log(1-softmax(x)) discussed in #129657) ...",https://github.com/pytorch/pytorch/issues/135359,39ffe33bfc512d09ba19cc7d1d3938b2a4734b2ab0b3bfdd6672f1051be9e916 references,issue,135437,issue,82627,medium,issue.comments[0].body,This relates to #82627 (comment),https://github.com/pytorch/pytorch/issues/135437,9c241ba525313cb56fb56ea562db1baaba11c6d1183e85476b81263d906e887e references,issue,134885,issue,128830,medium,issue.body,"🐛 Describe the bug (Might be a duplicate of #128830) When running a repro.py generated with TORCHDYNAMO_REPRO_AFTER=""aot"" TORCHDYNAMO_REPRO_LEVEL=4 on nightly (2.5.0.dev20240830+cu124), I run",https://github.com/pytorch/pytorch/issues/134885,88a571152f11a64c1e39f64e7085735edc237e681919c7cdcc7a5b522c399a60 references,issue,43369,issue,16797,medium,issue.comments[1].body,#16797 solution This solution works for me! Thank you man!,https://github.com/pytorch/pytorch/issues/43369,cb9c798855998e7702f10e4f8deecd4086e04ae2a86ac02379c11ff19386c66d references,issue,50092,issue,38642,medium,issue.comments[1].body,@mingzhe09088 @mrshenli Is it similar to this issue?#38642,https://github.com/pytorch/pytorch/issues/50092,e243566acd0977fdc8fda079abcc811f166edcc5e184dcc3d2800e66c7aafad2 references,issue,108968,issue,105641,medium,issue.body,"tensor([2, 3, 5]) out = torch.repeat_interleave(A, num_repeats.cuda(), dim=0) Indexing with a scalar tensor performs a synchronization. See #105641 for more details. torch.normal also incurs a sync on std: https://github.com/pytorch/pytorch/blob/main/aten/src/ATen/native/Distr...",https://github.com/pytorch/pytorch/issues/108968,fd9285e3129ea44974c4e036f59a015880275b61a72081df056425febf7680bf references,issue,108968,issue,128396,medium,issue.body,ates.h#L222 nanmedian incurs a sync: https://github.com/pytorch/pytorch/blob/main/aten/src/ATen/native/cuda/Sorting.cpp#L149 prod_backward: #128396 Alternatives No response Additional context No response cc @ptrblck,https://github.com/pytorch/pytorch/issues/108968,6d6661db577d26412bca1504edb9b54d823293626df6bc6f3bd14594e2568289 references,issue,108968,issue,73175,medium,issue.comments[0].body,"@Chillee maybe also that's why part of why repeat_interleave is slow: #31980, also a bit related: #73175",https://github.com/pytorch/pytorch/issues/108968,0e7e8dd6804ff780b603fc9d3f85a951351de98e8852e20487d79a6dfe53ed05 references,issue,122641,issue,72775,medium,issue.body,build torch from source with vulkan backend enabled but when I tried to run some minimal examples with vulkan backed I got same errors as: #72775. Setup: import torch cpu_device = torch.device('cpu') vulkan_device = torch.device('vulkan:0' if torch.is_vulkan_available() else '...,https://github.com/pytorch/pytorch/issues/122641,accc786b3373d7d5c164eb474bdc6be1d08392bad2c7ec9bd31e32189fce4af6 references,issue,117736,issue,95024,medium,issue.comments[0].body,"Very much related to #95024 although that issue refers to Conv3D. But I can confirm that similarly, the conv gets dispatched to slow_conv2d_forward_cuda.",https://github.com/pytorch/pytorch/issues/117736,766c18b31f7a1b0b67c1e032c5f748cb6593c21ac90d363050a985358a8bd926 references,issue,133566,issue,97575,medium,issue.comments[0].body,Related #97575,https://github.com/pytorch/pytorch/issues/133566,041c4cfa47b122e738c9aa872f35ff4a4ad102eeff533d557353967b1e2d846b references,issue,30968,issue,11389,medium,issue.comments[0].body,This appears to be an especially important instance of #11389. This refactoring will also be necessary for #18906,https://github.com/pytorch/pytorch/issues/30968,507796d224a59612ddef78ef3931d3cfd9fd35c0c94bde4af0f812e77101b4cf references,issue,30968,issue,18906,medium,issue.comments[0].body,This appears to be an especially important instance of #11389. This refactoring will also be necessary for #18906,https://github.com/pytorch/pytorch/issues/30968,77c1a05694fa6064abfe195b393c269f55db8011af47e8af26a8e01f2331b6fe references,issue,134570,issue,69519,medium,issue.comments[0].body,on for GPU-accelerated histogram calculations. It claims to be faster than NumPy on CPU and benefits greatly from CUDA capabilities[5]. [1] #69519 [2] https://dev-discuss.pytorch.org/t/cuda-histogram-feature-proposal-help/888 [3] https://hippocampus-garden.com/color_histogram/...,https://github.com/pytorch/pytorch/issues/134570,c50e275fdd04048ac8862565de12d7bc7c101820ee2117e5c77736ba1d77584a references,issue,74537,issue,50097,medium,issue.body,is the operator coverage: Fundamental Ops (expected 1.12) copy_ (to support casting to and from complex32) (PR: #73847) print support (Ref: #50097) storage support #73502 type promotion NOTE: List will be updated if supporting above list requires supporting another operator Hi...,https://github.com/pytorch/pytorch/issues/74537,1a77905aeff5aff084e3268583836aa8241042933ddc17735a50a8fc559f71d8 references,issue,74537,issue,67324,medium,issue.body,rfft fft.rfft2 fft.rfftn fft.ifft fft.ifft2 fft.ifftn fft.irfft fft.irfft2 fft.irfftn fft.ihfft fft.ihfft2 fft.ihfftn Support for stft (see #67324) (expected 1.12) fill_ pad Support for convolution (see also #71108) High-Priority Ops (expected 1.12) add sub mul div sum mean eq...,https://github.com/pytorch/pytorch/issues/74537,452ecee52203d0f39d448a03184fe44c4fe348635c394a855d7ee7e469563268 references,issue,112459,issue,105982,medium,issue.body,"ut.shape) Alternatives The normal GRU class, or any class that inherits RNNBase, doesn't work with vmap either, this has been documented in #105982 Additional context No response cc @zou3519 @Chillee @samdow @kshitij12345 @janeyx99",https://github.com/pytorch/pytorch/issues/112459,73f5b2971cd7cf1689dea5763611df2d03288bd98ff59b3c934850e659cbddc4 references,issue,112459,issue,134606,medium,issue.comments[0].body,Hello! I'm facing a similar issue and reported in #134606. Are there any updates on this?,https://github.com/pytorch/pytorch/issues/112459,b887e2fd794d3ee1d95c20f55d3650872f139255d428c588f47836248df5d315 references,issue,134363,issue,133250,medium,issue.body,"should try to change it to be more automatic and determine if those passes are safe to run (or add invariants if this is too hard, see also #133250). This is non-trivial: For DCE we can check if there are any mutable ops. If there are, then it's probably a bad idea to DCE For...",https://github.com/pytorch/pytorch/issues/134363,cffcdbb75e31c0fa680ae2b8289e93d986f5f7c5fa1df078da376d62c400c4eb references,issue,82443,issue,82139,medium,issue.comments[0].body,because the test uses nightly build to generate models). This issue currently blocking enabling DynamicQuantModule in iOS simulator tests: #82139 cc @pytorch/pytorch-dev-infra,https://github.com/pytorch/pytorch/issues/82443,561ec9f1ec4df964e8f089eaa5ce150db956c0d26b85d78052f02b3ad584f74e references,issue,134117,issue,123835,medium,issue.body,"(see below). Could we get the build process to include builds for conda? Related tickets #126174 #1856 , N8-CIR-Bede/documentation#199 and #123835 (and likely others). The following wheel file is available which has aarch64 cuda support (note that it seems this is only true fo...",https://github.com/pytorch/pytorch/issues/134117,c7b6054640bd70b1fbbfa5a0f01645d18b0c2a42ee86531f3d381835eac45623 references,issue,128046,issue,68332,medium,issue.comments[0].body,This is a bit related to #68332 and #110636. In the sense that there is a more general problem that a pytorch function currently doesn't work nicely with Python scalars. I,https://github.com/pytorch/pytorch/issues/128046,466e33b0f77bdcd73d87a4c3e85dc6d87e81c8fc57f7b349a4f8c2fdc844070e references,issue,133493,issue,129845,medium,issue.body,Related: #129845 cc @ezyang @chauhang @penguinwu @Chillee @samdow @kshitij12345 @janeyx99 @voznesenskym @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhu,https://github.com/pytorch/pytorch/issues/133493,f06b10fe849ca24d097c871a9e5d368f5168904e9b2c18b2ae2fef0ee9f9847f references,issue,108977,issue,24185,medium,issue.body,"free to close this issue if duplicate to the very related #58828 🚀 The feature, motivation and pitch Orignally requested and discussed at: #24185 #24185 (comment) In practice, the top needed things is actually extremal eigval/eigvec solver, as in #24185 Available functions bin...",https://github.com/pytorch/pytorch/issues/108977,80493093752f32faaff9e78f162eb313ca35eafdd8910f448bbfd895fa0337d3 references,issue,108977,issue,58828,medium,issue.body,"UPD: please feel free to close this issue if duplicate to the very related #58828 🚀 The feature, motivation and pitch Orignally requested and discussed at: #24185 #24185 (comment) In practice, the top needed things is act",https://github.com/pytorch/pytorch/issues/108977,b1d7008616b585e54c9f83b6ff842feaf75811ae4aedfd691fdb556d2af535a6 references,issue,108977,issue,87358,medium,issue.comments[1].body,"Related issues: #69538 to have some basic impl of https://docs.scipy.org/doc/scipy/reference/generated/scipy.sparse.linalg.spsolve.html #87358 Also, scipy.sparse.linalg namespace can be a reference for various sparse system solvers to consider: https://docs.scipy.org/doc/scipy...",https://github.com/pytorch/pytorch/issues/108977,7ae476caac2e7280724d264c2b6e1f3d87f9cce3a9cb20036f6f0f79c2ed7982 references,issue,26288,issue,23756,medium,issue.body,"Following the discussion in #23756, a simple way to enable users implementing inplace-activated batchnorm: provide inplace mode for BatchNorm and batch_norm expose backward m",https://github.com/pytorch/pytorch/issues/26288,7d4bb7819de539a9fc02e6264af355d7846011a7ea5160ddc3da6d80294b659e references,issue,114455,issue,68114,medium,issue.body,"contains torch.no_grad() , gc, and del model. But still during embeddings generation 250-300 mb memory getting stuck in ram. Similar issue: #68114 Is there any way to release cpu memory after every api call ? Versions 2.1.0+cu121'",https://github.com/pytorch/pytorch/issues/114455,4fae8760f0e352f0f493780adb4fc70e59716054b8bba535c6c167cd7ad9b1f1 references,issue,120814,issue,117394,medium,issue.comments[1].body,For 3 I am working on a RFC to generate a full graph that includes backward #117394,https://github.com/pytorch/pytorch/issues/120814,eac00b161829290d73b92b9b71966899dfcd90636acca54c39263128cebf459d references,issue,60294,issue,28214,medium,issue.body,ng width is less than the input's width #57911 - circular padding can only wrap once #46240 - we don't support symmetric reflection padding #28214 and #29863 - we only support a subset of input dimensionality Pitch Most of the numpy.pad interface can probably be copied. Howeve...,https://github.com/pytorch/pytorch/issues/60294,f068ff43c6f820d851fb17c83a75b206e31b1d3c08dc0becf568ecdc739f0e60 references,issue,60294,issue,29863,medium,issue.body,"less than the input's width #57911 - circular padding can only wrap once #46240 - we don't support symmetric reflection padding #28214 and #29863 - we only support a subset of input dimensionality Pitch Most of the numpy.pad interface can probably be copied. However, we could...",https://github.com/pytorch/pytorch/issues/60294,221563a6ba173bf5e605b2a753e561547a072c6748a0642d86c7c8cf79d89cf6 references,issue,60294,issue,46240,medium,issue.body,#52205 - reflection padding is only supported if padding width is less than the input's width #57911 - circular padding can only wrap once #46240 - we don't support symmetric reflection padding #28214 and #29863 - we only support a subset of input dimensionality Pitch Most of...,https://github.com/pytorch/pytorch/issues/60294,4d686ad4647bb23f73f447173e71646bd8863a620e5b8cbf7b0592c01063fbcb references,issue,60294,issue,52205,medium,issue.body,"nction, based on numpy.pad Motivation NumPy compatability. Plus, this will offer a solution to several issues with torch.nn.functional.pad: #52205 - reflection padding is only supported if padding width is less than the input's width #57911 - circular padding can only wrap onc...",https://github.com/pytorch/pytorch/issues/60294,84d8fc83eb1f3493b15698f0aa18046a7f3e81f55075905e07d4de15b166a31e references,issue,60294,issue,57911,medium,issue.body,several issues with torch.nn.functional.pad: #52205 - reflection padding is only supported if padding width is less than the input's width #57911 - circular padding can only wrap once #46240 - we don't support symmetric reflection padding #28214 and #29863 - we only support a...,https://github.com/pytorch/pytorch/issues/60294,173e15c566feffacb81432a5e56bfab7799eda23d739f44ae30a9048eb0d1b0b references,issue,105465,issue,58734,medium,issue.body,"result are not well defined in C++, so at least for bit manipulations being able to clearly express uint32 might be useful. Existing issue: #58734 Related: pytorch/ao#292 on supporting BitTensor natively (and especially as outcome for boolean ops like torch.ge) Versions N/A cc...",https://github.com/pytorch/pytorch/issues/105465,be268cdca3750eb08cf31e3d3ac7c102f5fe92e78fb6223d24a0addf336d2050 references,issue,132212,issue,118107,medium,issue.comments[0].body,d actively working on addressing operator coverage gaps. I'll suggest calling out individual ops that are most relevant to your use case in #118107 to help us prioritize. Thanks!,https://github.com/pytorch/pytorch/issues/132212,19b72d11c6d6668e039db8ce142e8bfd5d7eb6b9e7c01d38a8c80c31158a89d6 references,issue,126963,issue,88,medium,issue.comments[0].body,"allable=0x7ffeed539d50, args=, nargs=, keywords=0x0) at /usr/local/src/conda/python-3.11.9/Objects/call.c:214 #88 0x00000000005116e7 in _PyEval_EvalFrameDefault (tstate=, frame=, throwflag=) at /usr/loc...",https://github.com/pytorch/pytorch/issues/126963,283ec2ea3c576f7fd7fa3072fc62acdacc7a82018929ff3955a8b965be3f3f6b references,issue,132559,issue,109017,medium,issue.comments[0].body,I would propose .numpy be extended for sparse tensors: #109017 But for all other tensors without overridden .numpy() I would suggest .numpy() to auto-convert first to dense tensor. This might be useful,https://github.com/pytorch/pytorch/issues/132559,f9c48ca8abdb4ee40afd8a63af534289519dc2532910bed0d08230865cfbed50 references,issue,132559,issue,109017,medium,issue.comments[1].body,e name them as .scipy()/.from_scipy() or at least under some torch.utils.* namespace for zero-copy conversions to<>from scipy sparse arrays #109017 But for all other tensors without overridden .numpy() I would suggest .numpy() to auto-convert first to dense tensor. This might...,https://github.com/pytorch/pytorch/issues/132559,4bfc064712e275da986b69e9b51a92bd86f663272a9ec1f24e29eab010df2176 references,issue,132020,issue,41226,medium,issue.body,"at).is_contiguous(memory_format=torch.contiguous_format)=False Similar bugs have already been reported a few times (#62027, #86558, #85613, #41226), but I haven't seen updates since 2022. The discussion in #62027 suggests that the memory_format kwarg to torch.Tensor.to is not...",https://github.com/pytorch/pytorch/issues/132020,3740a96ac6de6df0a2e4ff55c81f2722ef7741835c0c1b4aee4709d5d970cb28 references,issue,132020,issue,62027,medium,issue.body,"at=torch.contiguous_format).is_contiguous(memory_format=torch.contiguous_format)=False Similar bugs have already been reported a few times (#62027, #86558, #85613, #41226), but I haven't seen updates since 2022. The discussion in #62027 suggests that the memory_format kwarg to...",https://github.com/pytorch/pytorch/issues/132020,f0175d88bb6436aa144c234dedb5a735dfa43f472f9363a338049fe4aeda49e1 references,issue,132020,issue,86558,medium,issue.body,".contiguous_format).is_contiguous(memory_format=torch.contiguous_format)=False Similar bugs have already been reported a few times (#62027, #86558, #85613, #41226), but I haven't seen updates since 2022. The discussion in #62027 suggests that the memory_format kwarg to torch.T...",https://github.com/pytorch/pytorch/issues/132020,33ddc808b092ff8ec91d7ea4d962db6038e6d1bc4c25341699f1b3741b102ebd references,issue,105982,issue,103161,medium,issue.body,"dnn.is_acceptable(fw.data)): RuntimeError: accessing `data` under vmap transform is not allowed Superficially, this seems quite similar to: #103161, minus the jacrev stuff. I'll add a couple additional observations in a reply to this issue to avoid making this post excessively...",https://github.com/pytorch/pytorch/issues/105982,f1f1d45418584302105d4ffa880f8dff22a5a649ca7ce30cdaf4c292eed4f744 references,issue,105982,issue,103161,medium,issue.comments[0].body,wargs) RuntimeError: Cannot access data pointer of Tensor that doesn't have storage The above looks even more similar to what's going on in #103161. Any ideas? PS Tysm to all the PyTorch team. Have absolutely loved this library since switching from TensorFlow and have never lo...,https://github.com/pytorch/pytorch/issues/105982,f28f399ffd247d9d5bcccfe0131c7c4c0ae2e84beb4c81e61380a24d09923fba references,issue,115819,issue,64208,medium,issue.comments[0].body,Related: #64208,https://github.com/pytorch/pytorch/issues/115819,911486d4630825a31ebc070ee4e8a7aa60b97ec716a838be3ad91df90e452272 references,issue,130430,issue,77764,medium,issue.body,"y implemented for the MPS device. If you want this op to be added in priority during the prototype phase of this feature, please comment on #77764. As a temporary fix, you can set the environment variable PYTORCH_ENABLE_MPS_FALLBACK=1 to use the CPU as a fallback for this op....",https://github.com/pytorch/pytorch/issues/130430,0a8714d2f524f03c68357f124acdd3c1baf4cfad272e81b900ec6683ca5414b8 references,issue,130430,issue,77764,medium,issue.comments[0].body,Could you comment on #77764 as well?,https://github.com/pytorch/pytorch/issues/130430,274ff5316a669fe88299ca874c5919bcf524d2123afa2d03165fc2e8ab77a302 references,issue,83818,issue,51720,medium,issue.body,it down to a torch::linalg::eigh call. ** On entry to SSYEVD parameter number 8 had an illegal value See also? Possibly related to #68291 / #51720? Versions 1.12 $ wget https://raw.githubusercontent.com/pytorch/pytorch/master/torch/utils/collect_env.py --2022-08-21 09:45:24--...,https://github.com/pytorch/pytorch/issues/83818,9ba39ecd40197bead27d20cca1131a41a33184614a2c1bad9f79debc5c32f9a6 references,issue,83818,issue,68291,medium,issue.body,arrowing it down to a torch::linalg::eigh call. ** On entry to SSYEVD parameter number 8 had an illegal value See also? Possibly related to #68291 / #51720? Versions 1.12 $ wget https://raw.githubusercontent.com/pytorch/pytorch/master/torch/utils/collect_env.py --2022-08-21 09...,https://github.com/pytorch/pytorch/issues/83818,ba1f7c32aa11edb1057425fd77cb4d09d8ee455c3a7c08a7b8e5f927d9211e65 references,issue,130529,issue,30702,medium,issue.comments[0].body,Related: #9410 #94233 #30702,https://github.com/pytorch/pytorch/issues/130529,4c112809ac04d84cfbfa6d8a9c8ce279dd002e0670d42e85afbba56f1d75e828 references,issue,130529,issue,94233,medium,issue.comments[0].body,Related: #9410 #94233 #30702,https://github.com/pytorch/pytorch/issues/130529,c9b90801b684517879df707219ef9d0b49be959b7ef5b63fcfb785f6fbc047bb references,issue,130055,issue,130058,medium,issue.body,constructor which allows users to dictate how the args and kwargs will be passed into the model's forward. Also solves issues mentioned in #130058 because we can pass in schedule state arguments. Plan of changes Update PipelineStage constructor such that it takes input_args an...,https://github.com/pytorch/pytorch/issues/130055,78b73bb889fcddf12eb2a1e7f2c8f96c743b457102b7d2bf6753aa8b0d53a890 references,issue,129442,issue,129749,medium,issue.comments[0].body,Is this identical #129749 ?,https://github.com/pytorch/pytorch/issues/129442,fd2fed3dcf68ad8ca72ef26a925abfcae851a07398b8006b479ca8bf8f750f42 references,issue,126472,pr,124624,medium,issue.body,"by setting: import torch._dynamo torch._dynamo.config.suppress_errors = True Failure 2: Based on the failure, I tried with @soulitzer's PR #124624 patched on top: Details Traceback (most recent call last): File ""/home/dberard/local/scripts/nt_2.py"", line 38, in torch....",https://github.com/pytorch/pytorch/issues/126472,d88ade39342b4c60868d7f45d9d326e310dc525715355ee9f8f6271760b46df4 references,issue,126472,pr,124624,medium,issue.comments[0].body,"ix to is_concrete_int() to return true for nested ints #126563 Avoid problem of getting multiple nested ints for the same dimension I tried #124624 here, but that didn't work for me. I think there's another unbacked SymInt issue, but I also don't really understand why we need...",https://github.com/pytorch/pytorch/issues/126472,5a2488432844cba2b7ef17dff272f4eb4d953990af8042380d190b7ce3359ef8 references,issue,129272,issue,128641,medium,issue.comments[0].body,Most probably related to #128641,https://github.com/pytorch/pytorch/issues/129272,dd334e75ff1e92d78c3beb2623f75df71eae6c5ed801880f28e49bef803de288 references,issue,50122,issue,35666,medium,issue.body,"Created per @albanD request. Originally started in #31829 (comment) and in #35666 (comment) Motivations: these gradient hacks are quite common and not completely trivial, should be nice to have them in core as idioms In a",https://github.com/pytorch/pytorch/issues/50122,51bd3ca49edd3d3cb2fab14e658c87421480844841d6e9d6122edc24432c610d references,issue,39007,issue,18095,medium,issue.body,Originally in #18095 (comment): @vadimkantorov: Works: https://pytorch.org/docs/stable/torch.html#torch.flip Breaks: https://pytorch.org/docs/master/torch.html#,https://github.com/pytorch/pytorch/issues/39007,3d6284f5af35566bcef91457eb3d2eb67fbcd1c7b864750f9405ffd13e2d23b1 references,issue,128778,issue,128564,medium,issue.comments[0].body,It looks similar to #128564. Can you edit c:\Users\georg\Documents\University\FIT\Research\KAN Network\Quadcopter.venv\Lib\site-packages\kan\spline.py and print the si,https://github.com/pytorch/pytorch/issues/128778,d6eb77908ca71060d513893594de0df8ed1b03622f1675ea3f6c79a22124847b references,issue,128694,issue,128693,medium,issue.comments[1].body,"uant called at::_fused_moving_avg_obs_fq_helper, and the error was triggered in torch.fused_moving_avg_obs_fake_quant using the same input. #128693",https://github.com/pytorch/pytorch/issues/128694,ceb9a98b4ce78b6da344d7b59012acbd9b14fb0d67a662d5f6aaec7fa0467118 references,issue,88148,issue,87960,medium,issue.body,"🐛 Describe the bug Does not seem like a real user issues, but detected using some sort of API fuzzing: #87960 #87961 #87963 #87964 #94594 Sanitizer detected issues: #88724 #88940 #88939 Versions None",https://github.com/pytorch/pytorch/issues/88148,506802ff46aa9a3a8447c43cc57771d491a9007d9ed9ebc31e89ca85df073a0e references,issue,88148,issue,87961,medium,issue.body,"🐛 Describe the bug Does not seem like a real user issues, but detected using some sort of API fuzzing: #87960 #87961 #87963 #87964 #94594 Sanitizer detected issues: #88724 #88940 #88939 Versions None",https://github.com/pytorch/pytorch/issues/88148,1effade9ee7ca5abe3f688df3c267e5fb0da0c3cb98c1d459788e4a077907827 references,issue,88148,issue,94594,medium,issue.body,"🐛 Describe the bug Does not seem like a real user issues, but detected using some sort of API fuzzing: #87960 #87961 #87963 #87964 #94594 Sanitizer detected issues: #88724 #88940 #88939 Versions None",https://github.com/pytorch/pytorch/issues/88148,24d10973a8f245050952b8805a443e686b804c64070ea7f98c767ddda99c85ee references,issue,45685,issue,45682,medium,issue.comments[1].body,"ithub.com/huggingface/transformers/blob/master/src/transformers/modeling_bert.py for example, we added torch.Assert to replace assert Maybe #45682 is related?",https://github.com/pytorch/pytorch/issues/45685,5633cafbdefd6aebdf4a13935f23e399302238902c42a7fffbca675071959a0a references,issue,100626,issue,35749,medium,issue.body,"icular, the following issues persist: JIT is incompatible with custom backwards. The only way around this is writing custom C++ extensions. #35749 Implementing custom backward hooks for a whole nn.Module is extremely counter-intuitive, consider this example from Chapter 4: def...",https://github.com/pytorch/pytorch/issues/100626,32470f1aa238a69d843de6061576d9037fab16e33de8077f832113f4ed54e3c7 references,issue,52289,issue,52260,medium,issue.comments[0].body,Maybe related: #52260,https://github.com/pytorch/pytorch/issues/52289,3c414e24621031ff3cdd43f1272b937023f1448ab23769409e406149185490f5 references,issue,46948,issue,40373,medium,issue.body,Motivated in #40373 (comment). Recording assertions assert X.shape[1] == Y.shape[2] and X.shape[2] == 64 and Z.ndim == 3 (or directly via torch.assert_shapes(.,https://github.com/pytorch/pytorch/issues/46948,518a5e55d0319bcdb93bfc59c323e16edbf94f759253c47fa1d0431ed088c14e references,issue,121207,issue,101314,medium,issue.comments[0].body,#101314 reported similar bug,https://github.com/pytorch/pytorch/issues/121207,0ccb4786bfbc15deda03e97e0f0dfb989f106e167fe15600fe0530d31bf676ed references,issue,38019,pr,127702,medium,issue.comments[1].body,"= torch.cuda.nccl.init_rank(nRanks, id, rank) causes SystemError: PY_SSIZE_T_CLEAN macro must be defined for '#' formats. I've fixed it in #127702 along with #38019 (comment).",https://github.com/pytorch/pytorch/issues/38019,6d7ff80bc1de6891236b0348a2dfef61a8a68dcc89e4e4fe761f661be8901467 references,issue,127062,issue,69364,medium,issue.comments[0].body,Somewhat related on dynamic quantization support - but on CUDA: #69364,https://github.com/pytorch/pytorch/issues/127062,f2568465063d4afd87fc4ac7c2a2f6c5eb155c35978c9914d73e30a30450c4bd references,issue,126714,issue,24422,medium,issue.comments[1].body,otifications I will remove them from the list. I haven't updated the labels I'm waiting for this issue to be resolved :P. More seriously -- #24422 gets edited 6 times a month. What's the frequency for which we'd make a cutoff for features like this?,https://github.com/pytorch/pytorch/issues/126714,a7ae44a41c18a0db7b19aed6b52302ff5fdfce3e03cefc5a1bc28eb167a8c909 references,issue,127164,issue,79703,medium,issue.comments[0].body,#79703 looks similar to this,https://github.com/pytorch/pytorch/issues/127164,a134e968889b2554808152913c5ccd5cb6ca29ec3c3f46dcd02001ef0cce5ed6 references,issue,22281,issue,6564,medium,issue.comments[0].body,Superset of #6564,https://github.com/pytorch/pytorch/issues/22281,2709fcc3d4e13a5d5e41bb041b313bed3a152de86d9447eff4283b2556a9f38e references,issue,127075,issue,82510,medium,issue.comments[0].body,Sounds like #82510,https://github.com/pytorch/pytorch/issues/127075,5d6e3e7ea8d937a19a23cc4074ac1323091b536bd1ec87cb037e6a5e7a963998 references,issue,52915,issue,49252,medium,issue.body,oesn't work # both cases work with torch.linalg.solve Additional context Memory inefficiency of the actual implementation is discussed here #49252. cc @jianyuh @nikitaved @pearu @mruberry @heitorschueroff @walterddr @IvanYashchuk,https://github.com/pytorch/pytorch/issues/52915,1df893d74e5dd998e21b3cdc0b3b48b0bf988bd5fafc04ffa16132282e40fa3a references,issue,86162,issue,15457,medium,issue.comments[0].body,Related: #15457,https://github.com/pytorch/pytorch/issues/86162,412b47c6b18c1c0f810808f746a39cdc5ee14ec3676769fbf5a273e219301362 references,issue,126276,issue,38034,medium,issue.body,"pe[1], -1).permute(0, 2, 1) How to slove it ? Versions torch 1.10.0+cu113 torchaudio 0.10.0+cu113 torchvision 0.11.0+cu113 OTH-similar case #38034 #34302 #29236 https://discuss.pytorch.org/t/torchscript-does-not-support-assigning-output-of-module-to-element-of-tensor/108371 cc...",https://github.com/pytorch/pytorch/issues/126276,92005f3f31e40d8834c69f05ce40fdcf121a7ba4aa7f4cfc768f6496e64d0b23 references,issue,106485,issue,106469,medium,issue.comments[1].body,Can this be because of CUDA driver or GPU specific bug? #106469 also report slowdown on A100.,https://github.com/pytorch/pytorch/issues/106485,e8520e86256c6cb5a0daa937d1f647ce6e12302d480cf74cd2fc49e5d74f6d8c references,issue,82139,issue,82443,medium,issue.comments[0].body,@linbinyu isolated the issue here: #82443,https://github.com/pytorch/pytorch/issues/82139,a320479cfd33ea804e5e695ca7c24214dc785f5ef57d60542651d3ffbed206f8 references,issue,82139,issue,82443,medium,issue.comments[1].body,Dependent on #82443,https://github.com/pytorch/pytorch/issues/82139,d8b9f283157544eeffb7c1c8ae609d75411cd105475d207fe06b1195b5fd1029 references,issue,126164,issue,72831,medium,issue.body,ous mode. Such a context manager could later be extended to allow for example toggling the mode for individual autograd backward functions (#72831). Alternatives No response Additional context No response cc @mruberry @kurtamohler,https://github.com/pytorch/pytorch/issues/126164,944e7c5f3e93d942ff64b3ec1673b379574f6c8880a4ab644acacf7559886a58 references,issue,126039,issue,126037,medium,issue.comments[0].body,same as #126037 (comment) said,https://github.com/pytorch/pytorch/issues/126039,be80b710766a7473caed5fb4d2c12d08dce1e4bedd34f4f9af503cd6454625ba references,issue,125861,issue,18182,medium,issue.comments[0].body,There were also some related discussions on passing in keras-like weight init options to non-lazy modules as well: #18182,https://github.com/pytorch/pytorch/issues/125861,7b07adc43e73d5406f523bc0c8b9686f49beadead0f6aefb6cddcfc941b82161 references,issue,125531,issue,55135,medium,issue.body,"ReduceLROnPlateau. Unittests for SequentialLR expect all schedulers to use close-form implementation, whereas the ChainedScheduler is not. #55135 #67586 #116776 #67958 Allow param_groups to be passed to lr_scheduler class #72146 probable issues: param_groups must be an instanc...",https://github.com/pytorch/pytorch/issues/125531,b9d29a332b83d3267607192b26b44cab4e324f4283cc1474a115a72c0e5e94ba references,issue,125531,issue,67586,medium,issue.body,"LROnPlateau. Unittests for SequentialLR expect all schedulers to use close-form implementation, whereas the ChainedScheduler is not. #55135 #67586 #116776 #67958 Allow param_groups to be passed to lr_scheduler class #72146 probable issues: param_groups must be an instance of d...",https://github.com/pytorch/pytorch/issues/125531,d922443cadfbc2c034633f4dd8942509d0f4036f5ae6c6e2bf357e7e6b7424ea references,issue,125531,issue,67760,medium,issue.body,"🚀 The feature, motivation and pitch Feature Proposal to solve issue mainly raised from issue #67760, and some other issues related to lr_scheduler module. Refactor lr_scheduler module make lr_scheduler class to support all hyperparameter s",https://github.com/pytorch/pytorch/issues/125531,1c4b5804a44b673053d2c9ad42487d5556e8f683417a01b9278e09a2a2958ef0 references,issue,125531,issue,67958,medium,issue.body,"ittests for SequentialLR expect all schedulers to use close-form implementation, whereas the ChainedScheduler is not. #55135 #67586 #116776 #67958 Allow param_groups to be passed to lr_scheduler class #72146 probable issues: param_groups must be an instance of dict right from...",https://github.com/pytorch/pytorch/issues/125531,9b4095704c0eecd1fde0f6920bb04356e7c3ec08a183b8d09f8d252bdf60daa7 references,issue,125531,issue,68978,medium,issue.body,"given optimizer (#83159 ) refactor SequentialLR & ChainedScheduler to support ReduceLROnPlateau with arbitrary step kwargs inputs. #110761 #68978 Deprecate step(epoch) warnings and remove supports of epoch, and probably change to closed-form implementation as well. conflicts:...",https://github.com/pytorch/pytorch/issues/125531,5ad1e3b9ee18ce81dc7e1069eb1d5baf044093dd266381b1dbe6e1fe81015ba0 references,issue,125531,issue,72146,medium,issue.body,"orm implementation, whereas the ChainedScheduler is not. #55135 #67586 #116776 #67958 Allow param_groups to be passed to lr_scheduler class #72146 probable issues: param_groups must be an instance of dict right from the beginning, or else any additional keys added by the optim...",https://github.com/pytorch/pytorch/issues/125531,290774d0062c128b79eceab3fa8791b1cf3e7ef800173f966a3a7f77bb6b9ea5 references,issue,125531,issue,83159,medium,issue.body,to lr_scheduler module. Refactor lr_scheduler module make lr_scheduler class to support all hyperparameter settings in the given optimizer (#83159 ) refactor SequentialLR & ChainedScheduler to support ReduceLROnPlateau with arbitrary step kwargs inputs. #110761 #68978 Deprecat...,https://github.com/pytorch/pytorch/issues/125531,cbecffd539a73bb65b41a4cdbc09c86a59171281a0b58c94d142b51f99b8936e references,issue,125531,issue,110761,medium,issue.body,"s in the given optimizer (#83159 ) refactor SequentialLR & ChainedScheduler to support ReduceLROnPlateau with arbitrary step kwargs inputs. #110761 #68978 Deprecate step(epoch) warnings and remove supports of epoch, and probably change to closed-form implementation as well. co...",https://github.com/pytorch/pytorch/issues/125531,89fb6f1ec072792e1303c6d39165c91ed90ba269108f61fd3f86cded23d59a22 references,issue,125531,issue,68332,medium,issue.comments[0].body,"Related discussion: #68332 The main motivation is that object-oriented API design (statefulness, chaining, reloading from state dict, re-warming, level of complexity,",https://github.com/pytorch/pytorch/issues/125531,34f035956d7abc6b979bb138079c57710f0170f1315e73cd12c3f52eec550689 references,issue,125626,issue,124016,medium,issue.comments[0].body,This is basically the same as #124016,https://github.com/pytorch/pytorch/issues/125626,4fb96cada08171a6aceba0708bff75133bf3ffbba2f945f0c279cdf9dc5d285e references,issue,59168,issue,28619,medium,issue.body,"vices, and improve performance on 3D model training and inference, e.g. U-Net 3D. This task is similar to the channels_last (2d) support in #28619, but would be much less complicated. This is because most of the C++ infrastructures already have some kind of support with at::Me...",https://github.com/pytorch/pytorch/issues/59168,a4d6e29e78558acdfde8d4d18f562d7b2981d1d953caa5c28495c6709805df7f references,issue,105582,issue,49444,medium,issue.body,"her features include low precision. Since PyTorch 1.12, this API has been added in TorchScript JIT fuser path showing promising performance #49444. Integrating the oneDNN Graph Compiler with Inductor C++ backend offers further performance enhancements. Additionally, adopting t...",https://github.com/pytorch/pytorch/issues/105582,6c27f5586fa64fb5185b61d69d21fde40501c27bdb5cf42861a23c0b184bd17b references,issue,114835,issue,114723,medium,issue.body,"Background We are upstreaming for Intel GPU ([RFC] Intel GPU Upstreaming · Issue #114723 · pytorch/pytorch (github.com)). For the first step, targeting recent popular and typical DL workloads, we plan to enable Intel GPU backend",https://github.com/pytorch/pytorch/issues/114835,8aaf332b63891635970f7452b863eb515a7113e4651e5f605de663afea389d8d references,issue,125239,issue,110295,medium,issue.comments[0].body,ong indicator of a deterministically flaky test due to test order. Maybe we don't need the expensive second step. There are some context on #110295 to figure out a way to reduce the effect of global states on how a test is run. Another source of deterministically flaky test is...,https://github.com/pytorch/pytorch/issues/125239,7d92b09b685a43785b599fdc56016d9734b9b679b4f7bf77c1334f1daacdcb00 references,issue,124788,issue,84864,medium,issue.comments[1].body,This is most likely related to #84864,https://github.com/pytorch/pytorch/issues/124788,b7bd285f91735ff4bcdffa302f04324798c2b8ac7f720ac5f0cff2d53dbaa5e9 references,issue,124262,issue,94294,medium,issue.body,"🐛 Describe the bug I met a problem similar to #94294 when using torch.multiprocessing RuntimeError: CUDA error: CUBLAS_STATUS_EXECUTION_FAILED when calling `cublasGemmEx( handle, opa, opb, m,",https://github.com/pytorch/pytorch/issues/124262,6886f0b222e9dc3af8920acb94e71b093d1fa479b5d5782c3f799fb24bd712a5 references,issue,116506,issue,115852,medium,issue.comments[0].body,A related issue on uniting kernels for pointwise convs and matmul: #115852 - feature request for support of groups for nn.Linear / F.linear / torch.matmul / etc,https://github.com/pytorch/pytorch/issues/116506,0d739e7b70dde9886c23718ff4eb424cfd2f8329d12b74d39ea9da122dd9218a references,issue,124572,issue,41625,medium,issue.comments[0].body,"If you do this, might also be good to make public the classical slice method in general (narrow does not support the step argument) :) #41625",https://github.com/pytorch/pytorch/issues/124572,5315f0ac2f6dd9d44237d5f86113064a2bd0055ba15c8d9ac7854a00c972badb references,issue,124222,issue,124016,medium,issue.body,"currently 400 issues with high pri label, most (250+) from 3+ months ago and oldest of which is from 2017 This should help with a subset of #124016 cc @ZainRizvi @kit1980 @huydhn",https://github.com/pytorch/pytorch/issues/124222,fd02c4a17b94d856d9b71f6c7ada49713ee6c8fcb11b0339c62f0809390013cb references,issue,34279,issue,40570,medium,issue.comments[1].body,Relates to #40570,https://github.com/pytorch/pytorch/issues/34279,0893d27a25d51d3b84bf5f56d036cebb15e4c4156fabe198024f86158d1c78b9 references,issue,124461,issue,52312,medium,issue.body,"s way seems a bit redundant. So when encountering global variables, can you treat them as constants during torch.jit.script? Related issue: #52312 Alternatives No response Additional context No response cc @EikanWang @jgong5 @wenzhe-nrv @sanchitintel",https://github.com/pytorch/pytorch/issues/124461,af647312f05685645de31189fedbf56e20a0882324deab2291c586fb3399eca6 references,issue,124291,issue,115347,medium,issue.body,"y2 = torch.tensor([0, 2]) result2 = batched_index(x, y2) print(result2) Expected result2 to return torch.tensor([0, 2]) This is related to #115347 but not exactly the same. Versions Versions of relevant libraries: [pip3] numpy==1.26.4 [pip3] optree==0.11.0 [pip3] torch==2.4.0a...",https://github.com/pytorch/pytorch/issues/124291,ba315280592feb39648cc22bb237737b280e16d9c205e8bcc6f279862073c73d references,issue,110758,issue,106951,medium,issue.comments[1].body,I suspect the size and strides don't match. Could you print v.size() and v.strides() too? This may be related: #106951,https://github.com/pytorch/pytorch/issues/110758,380ec13fb68cb2ed351f618c7699f90f614e3b94fbb213df368b521059d96d84 references,issue,123421,issue,88,medium,issue.body,000000005106ed in ?? () #85 0x0000000000528d21 in _PyFunction_Vectorcall () #86 0x000000000053bdb1 in ?? () #87 0x00000000005302b5 in ?? () #88 0x00000000005f702d in PyObject_CallMethod () #89 0x00007fb6119eb0d3 in torch::dispatch_on_mode (torch_function_name_str=0x7fb611b8794...,https://github.com/pytorch/pytorch/issues/123421,04ccb726c237f8f2286a32aff921f799b0883b309089a0a76a6738bceab05aec references,issue,37488,issue,34606,medium,issue.body,ome/cloudhan/workspaces/pytorch/build/lib/libtorch_cpu.so: undefined reference to `pthreadpool_compute_4d_tiled' similar issue mentioned in #34606 (comment) To Reproduce Steps to reproduce the behavior: build the code Expected behavior Environment PyTorch version: N/A Is debug...,https://github.com/pytorch/pytorch/issues/37488,5c9cabbd2eb1f0f1ecbefac4af27d2ff0c98603626587d7a2ec1d12d5226710c references,issue,117617,issue,116567,medium,issue.comments[1].body,#116567 related nan sorting issue,https://github.com/pytorch/pytorch/issues/117617,94b149c10217d12f053aee432dbc7727c66b2b5737e6134f6ec6a5b53726f9bf references,issue,105943,issue,102517,medium,issue.comments[0].body,Maybe related #102517,https://github.com/pytorch/pytorch/issues/105943,f208af55f370aaeb794d20d4217ee28e337e6762bf0007cacf57b20c58d4722e references,issue,66707,issue,35666,medium,issue.comments[0].body,"Also there used to be problems with too small default epsilon for fp16, the result was division by zero and nans, one of related issues: #35666. Oh, maybe I remembered this incorrectly: #41527 (comment)",https://github.com/pytorch/pytorch/issues/66707,b10368a17d31ea94ac3ed28ef0eadd288e4b96a164ab94f288317f8ec5717325 references,issue,123157,issue,52439,medium,issue.comments[1].body,"y of setting these options for a given gemm call - either with with statement, or maybe even better also allowing some explicit hints= arg: #52439",https://github.com/pytorch/pytorch/issues/123157,0419c159648b69ae4449e8e430c412bf87c0df842509f6389734c68a5ede6986 references,issue,120139,issue,47055,medium,issue.body,"n this could solve the nasty https://github.com/pytorch/pytorch/blob/v2.2.1/torch/utils/data/sampler.py#L70-L95. On that, see also #122410, #47055. Related Discussions, Issues and Commits #19228, #47055, #23587. Thoughts Thoughts? :) cc @VitalyFedyunin @ejguan @dzhulgakov @ssnl",https://github.com/pytorch/pytorch/issues/120139,1c39f1974f8cd6f87f67526d41ed71cee3201787ce5588f173a9581b8cb8dff6 references,issue,120139,issue,122410,medium,issue.body,"hinking on this could solve the nasty https://github.com/pytorch/pytorch/blob/v2.2.1/torch/utils/data/sampler.py#L70-L95. On that, see also #122410, #47055. Related Discussions, Issues and Commits #19228, #47055, #23587. Thoughts Thoughts? :) cc @VitalyFedyunin @ejguan @dzhulg...",https://github.com/pytorch/pytorch/issues/120139,46e1eba8e107d73c1e0f2a80f70d7acba7b392fa7216fffa34320c2f775af44d references,issue,122410,issue,47055,medium,issue.comments[0].body,Didn't see but #47055 is related.,https://github.com/pytorch/pytorch/issues/122410,7dee24a3a10345a8530fbb8a17164118b76bfad9e6f97e3161ab14e8145d95a6 references,issue,123082,issue,106614,medium,issue.comments[0].body,"rically can produce nonzero-values. So you should probably just reset the diag to zero as postproc. You can also hope to use torch.compile (#106614) or PyKeOps As a feature request enabling a fix for this, I would propose that either: cdist to implement a true pdist if only a...",https://github.com/pytorch/pytorch/issues/123082,a9884fec89bd2067acb7324f7f4a5e5dc67de26a9d305eff4429061d75e02af5 references,issue,65156,issue,1512,medium,issue.body,"This is a frequent primitive for collation: #1512 import torch def stack_jagged(tensors, fill_value = 0): # does not use F.pad to not allocate, could as well use F.pad's arg names instead s",https://github.com/pytorch/pytorch/issues/65156,58900a0378df89b23e243414b498fbb44d61cac5735ab5f55ea2fc8185a30b27 references,issue,67760,issue,67761,medium,issue.comments[1].body,"Maybe actually a contradicting usecase: #67761 (different schedulers for different param groups), but IMO schedulers need a redesign :(",https://github.com/pytorch/pytorch/issues/67760,1ef00df6236a599f8592415b158d1ac86002218d9c5dcf4232fb515d0a02e100 references,issue,72106,issue,71911,medium,issue.body,ly it is linked statically) #72106 (comment). Additional context Another issue that requires attention is due to the upgrade to MKL 2022.0: #71911. cc @seemethere @malfet @pytorch/pytorch-dev-infra,https://github.com/pytorch/pytorch/issues/72106,bbdcdbb71b4d1876882799ca28a072883353aa23a396bad97b139f8ef8f55b11 references,issue,122287,issue,116626,medium,issue.comments[1].body,"his thus far is that most modules tested in test_modules.py are only tested for torch.float32 and torch.float64, which is on my list to fix #116626 (comment) So far I have only enabled float16 testing for MPS. Let me know if you are keen to help with this :)",https://github.com/pytorch/pytorch/issues/122287,f13dda5fa4350db46de03434a569bc79e351775fcac8c5344d2e98859e3e2f02 references,issue,106991,issue,106667,medium,issue.comments[0].body,Mildly related on better support of NVRTC in PyTorch (which could also enable faster compilation of inline CUDA-code extensions): #106667,https://github.com/pytorch/pytorch/issues/106991,efe67baa8aba2c89ffee059af90dc4e25162a8364ae919b4acf248050e5f9343 references,issue,121596,issue,26950,medium,issue.body,"m2col, as 512 * 2048^2 * 3^2 * 2 = 36G and very close the increased memory usage. Similiar thoughts are also mentioned in pytorch forum and #26950. We have figured out a workaround to compute results patch-by-patch (@lmxyy will soon share the code), but utlimatley, we hope thi...",https://github.com/pytorch/pytorch/issues/121596,661027f011acece84e2608648be697d8bd147863a0618d5d65b9a9801ea650fb references,issue,121596,issue,52439,medium,issue.comments[0].body,Might be related: #52439,https://github.com/pytorch/pytorch/issues/121596,be389e9a32219fd65a5111c78acc16d5911c62b648b41c0e50d29741b09d82c5 references,issue,108739,issue,73656,medium,issue.comments[1].body,I think this can be considered a duplicate of #73656. See also #81691 for a solution.,https://github.com/pytorch/pytorch/issues/108739,444b7b71e40302fef1845a605070fc29982484d049d3f323bb1952ec15947a91 references,issue,120930,issue,71631,medium,issue.body,"assert_close should fail. Interestingly, the assert_close passes if we change the autocast dtype to torch.float16. Here is a similar issue: #71631 To fix the bug locally, I copied make_graphed_callables and disabled autocast when calling autograd.grad, and the repro script pas...",https://github.com/pytorch/pytorch/issues/120930,1f11e8dda9d307ab02a2f97e7c818e304ca464d8712d733c1713a1820b476710 references,issue,63929,issue,63812,medium,issue.body,"ocal gradient penalty before running the backwards pass. However, DDP doesn't have great support for higher-order gradients as evidenced by #63812. A couple of issues: .grad field of the gradient is not synchronized across processes Calling .backward() twice is not well suppor...",https://github.com/pytorch/pytorch/issues/63929,303a13c4a95ed8154400282dcc9a646fd6b9d3a48411f7ac6614c65d556e70c3 references,issue,63929,issue,63740,medium,issue.comments[1].body,"re working on higher order derivative in DDP. I've been investigating computing Hessian vector product in DDP for a while. See my use case: #63740. (Related algorithms: algo in this paper and its related works On top of #63812 , can we also add double backward support for dyna...",https://github.com/pytorch/pytorch/issues/63929,7dab938a09d81136ff6c389bc6b248d002637ba145481f25650278f015cd2fe7 references,issue,63929,issue,63812,medium,issue.comments[1].body,"Hessian vector product in DDP for a while. See my use case: #63740. (Related algorithms: algo in this paper and its related works On top of #63812 , can we also add double backward support for dynamic graph? Basically make it works like autograd on single GPU.",https://github.com/pytorch/pytorch/issues/63929,532ac24e404754826a95eb8f4d21430684bd616ab14f54821eff052675af19b0 references,issue,118736,issue,87371,medium,issue.body,ike this can be fixed by either removing the __new__ annotation or replacing the return type with {typing/typing_extensions}.Self. related: #87371 Versions Details Collecting environment information... PyTorch version: 2.2.0+cu118 Is debug build: False CUDA used to build PyTor...,https://github.com/pytorch/pytorch/issues/118736,764a518bcdc868db5190af9b43e3eac9c2de05e00be861baa1d3ede148ba14f1 references,issue,46169,issue,11202,medium,issue.body,"only enough used to be implemented in scipy's pdist as ""sqeuclidean"". For instance, it can used to easily compute the cosine distance - see #11202 (comment). Also, anyone who only cares about the ordering of the 2-norm distances to the other pairs don't actually need the expen...",https://github.com/pytorch/pytorch/issues/46169,c34303d191dc349870c4245ce39478f854a7fb795f948f9a27e7592da7edfaa9 references,issue,46169,issue,28119,medium,issue.comments[0].body,Related #28119,https://github.com/pytorch/pytorch/issues/46169,868e8211a63c201e8bdee6ea13b203c3ed27938c0e61b21f075129d8de2e2545 references,issue,118463,issue,118462,medium,issue.body,"🐛 Describe the bug torch.argmin output is diffrent between eager mode and JIT-compiled mode. Similar with #118462 Allthough the model output is irrelevant from torch.special.erfc, deleting the torch.special.erfc from the code would prevent error from ha",https://github.com/pytorch/pytorch/issues/118463,e8945fd7a8553a5075c2e3caf8873a3df8495a11846e410f09f7c6e875dd2a5e references,issue,97772,issue,118092,medium,issue.comments[0].body,".py implementation doesn't support protocol 3/4/5. It is used when calling torch.load with weights_only=True. I submitted a separate issue: #118092 Once #118092 is solved, it will simple to bump up to the version.",https://github.com/pytorch/pytorch/issues/97772,09e6a5d568dd7bb2ffa2d4c62722e777b4c9e72f1543fc9e8bb23e1052ef7bf3 references,issue,118093,issue,56300,medium,issue.body,"oat (excluding special cases for float16 in speed/sparsity applications) then this would make sense, IMHO. Possible similar issue raised at #56300 Versions Collecting environment information... PyTorch version: 2.1.0+cu121 Is debug build: False CUDA used to build PyTorch: 12.1...",https://github.com/pytorch/pytorch/issues/118093,8c862630131345215a55de0821c1f638d38b7efe4edd58d11bc82d93c1422a71 references,issue,4392,issue,3667,medium,issue.body,it's very important to not accidentally state-dependence (which will lead to hard to debug performance regressions and problems.) Related: #3667 cc @csarofeen @ptrblck,https://github.com/pytorch/pytorch/issues/4392,fe918c1af2adfc11f19206e63103dc7c8eef5165da0b7f9551ff2ee4747b91bb references,issue,43501,issue,26889,medium,issue.comments[1].body,I think it might be best to merge this and your comments into #26889. Part of the problem could be that you get funny things with dynamic patterns (eg padding in conv).,https://github.com/pytorch/pytorch/issues/43501,26d33634046a71ed6d18476fd6ccbf6923b31e02d96c66c0507a13fac6d242cc references,issue,68407,issue,66504,medium,issue.comments[7].body,maybe related to #66504,https://github.com/pytorch/pytorch/issues/68407,d2215cb21e5ef3022f3ad5f63ee6593ac23573a7469472bb0378378147629a20 references,issue,34329,issue,25888,medium,issue.body,") tensor(12) Eager result tensor(42) Script result tensor(2) Hooks are also currently not supported in the PyTorch C++ API for nn::Modules (#25888). The the TorchScript C++ API script::Module methods can only be run by fetching them explicitly, e.g. my_module.forward(the_input...",https://github.com/pytorch/pytorch/issues/34329,841961d272d155de661d29c90d696d61aed6d19ed859a5228fe72faad5803a22 references,issue,112509,issue,112398,medium,issue.body,"to open an according PR if that is desired, maybe this feature can also be included for the promotion of nested Tensors to the beta stage (#112398) import torch from time import time # create a random dataset of variable sequence lengths (10k elements) sizes = torch.randint(32...",https://github.com/pytorch/pytorch/issues/112509,734f037b446954ea146a75157c89474822ca164391473e6652971d2e88fde711 references,issue,114465,issue,30896,medium,issue.body,"🐛 Describe the bug Error is similar to #30896, but errors not quite the same, and the problem is not caused by lack of GCC libraries in LD_LIBRARY_PATH Following source installation ins",https://github.com/pytorch/pytorch/issues/114465,dc556b5e55de86001d315e96a8a5d0732a571c882313250ecce2493a8f630064 references,issue,116461,issue,69359,medium,issue.comments[0].body,Related: #69359,https://github.com/pytorch/pytorch/issues/116461,e4c11bf58462d4187460afb79cc84c7bea0f7ab9dadfe44e1d371f031714254e references,issue,116461,issue,69359,medium,issue.comments[1].body,@vadimkantorov should this issue be closed as duplicate of #69359?,https://github.com/pytorch/pytorch/issues/116461,ac97cf4f0bfd2c80adf5cc7a906a779401ee90fdd86d0874023296cc770a94a1 references,issue,115821,issue,5405,medium,issue.comments[0].body,Maybe related about scatter arg dtype: #5405,https://github.com/pytorch/pytorch/issues/115821,3dd4a295e881cfa699ab56106cf4c4b0902231be580fe1f554e2d93a43766f2a references,issue,105319,issue,104193,medium,issue.body,"a (b, m, n) sparse tensor in the CSR format, unless we could provide masks as a list consisting of b (m, n) tensors. It might be blocked by #104193 though. Alternatives No response Additional context No response cc @alexsamardzic @nikitaved @pearu @cpuhrsch @amjames @bhosmer",https://github.com/pytorch/pytorch/issues/105319,883df171ba2c4adee7742c373ddbc38081e0879c4f3a2d5603f2e039e0418725 references,issue,105255,issue,68332,medium,issue.body,"as stateless schedulers used by detectron2: ""A stateless, scale-invariant hyperparameter scheduler: see its API doc"". It is also related to #68332, so maybe can be this design can be taken as base for schedulers enhancement / redesign in core (it has float iterations and uses...",https://github.com/pytorch/pytorch/issues/105255,43706f6cc96dda364b4247ff717a43ca55b97efb558fd9b948915867a432a01e references,issue,74420,issue,74236,medium,issue.comments[0].body,A resolution to this would probably also address #74236.,https://github.com/pytorch/pytorch/issues/74420,64114901fe81d92f62fe921285aab168fbc75d98287cdbe2997a1e0692089270 references,issue,62466,issue,13447,medium,issue.body,"🚀 Feature Similar to #13447, the current implementation of gather function does not support string objects. However, when we need metadata (e.g., the name of the file)",https://github.com/pytorch/pytorch/issues/62466,064101c32b4e6db5e18f745a046d6b83a0bfcd4de981ccb3721ad2ff4cb72514 references,issue,70173,issue,62448,medium,issue.comments[1].body,"lated to this issue. Since this issue is updated on Jan 12 2022, do we have already have the DTR ready in PyTorch now? I also noticed issue #62448 is also not closed. @MarisaKirisame Recently, oneflow has a paper related to using sliding window to evict tensor with lowest cost...",https://github.com/pytorch/pytorch/issues/70173,fb4aa8606a750707f8905abb84c947f453bd92f003f9615b9608e68a8ca4d441 references,issue,115434,issue,75599,medium,issue.body,"(input, k, dim): sorted, _ = torch.sort(input, dim=dim) return torch.take_along_dim(sorted_matrix, ks, dim=dim) As far as I understand from #75599, this approach might even be faster. Alternatives No response Additional context One use case of this is when you have multi label...",https://github.com/pytorch/pytorch/issues/115434,9562d0dd9f839a637eba221b37ba37520f83014295d8a01b68e2552284871778 competes with,issue,47513,issue,38622,medium,issue.body,ogous for the case mode='max' and threshold_mode == 'rel' return a - best > self.threshold * abs(best) This behavior was already noticed in #38622 which believes this to be an error in the documentation rather than the code. cc @jlin27 @mruberry @vincentqb,https://github.com/pytorch/pytorch/issues/47513,1465ecadabb600e8915060b0aa99046082de91219ba174555f4d65cfb3d3d4d6 references,issue,114967,issue,111676,medium,issue.comments[0].body,duplicate of #111676,https://github.com/pytorch/pytorch/issues/114967,f49aba65c915ac3fbc5cf031167053a14ccad4224b678455506cce03a88404e1 references,issue,114951,issue,103352,medium,issue.comments[1].body,ter if the underlying tensor is contig enough) But it would still be nice to have a faster item call. A related discussion for Python side: #103352,https://github.com/pytorch/pytorch/issues/114951,83e214e779f99e4186c198b68e4d16f9ab61e925680b7865ad0b92a2a04abe73 references,issue,103475,issue,46948,medium,issue.comments[1].body,less_than(torch.uint32) (this latter can be even more expressive for shorter dtypes assert tensor.is_numel_less_than(torch.uint8)) Related: #46948,https://github.com/pytorch/pytorch/issues/103475,7252ad18b68251937c2cc8a444950ef71ba9ef90ff38f63456cf2cf18ba55c03 references,issue,95604,issue,70954,medium,issue.body,"🐛 Describe the bug This issue should be related to #70954. The convolution performance problem of Conv2d has been mentioned in 70954 will actually affect multiple convolution operations, including",https://github.com/pytorch/pytorch/issues/95604,0fcc22c8eaf03eaff77892d3b63a597f007ff3d68af4255335b9aafaa8d5cbdc competes with,issue,62619,issue,62614,medium,issue.body,🚀 Feature Opening this issue to track requests: intermediate tensor reuse instead of re-computation #62614 fp16 support #62440 dynamic shapes #62675 opaque operators with implicit shape lowering rule for missing operations #62431 aten fallback on,https://github.com/pytorch/pytorch/issues/62619,b122d66d3cabac925f74d7a7d55374153593ee617e81a772af8df86267f115e1 competes with,issue,62619,issue,62616,medium,issue.body,shapes #62675 opaque operators with implicit shape lowering rule for missing operations #62431 aten fallback on CUDA device instead of CPU #62616 initialize cuda context #65073,https://github.com/pytorch/pytorch/issues/62619,3cd01f860dfffe6beccfb1913214b10d34bf9947b21d3d5b8e842d8dc3aa4cbd competes with,issue,62619,issue,62675,medium,issue.body,Feature Opening this issue to track requests: intermediate tensor reuse instead of re-computation #62614 fp16 support #62440 dynamic shapes #62675 opaque operators with implicit shape lowering rule for missing operations #62431 aten fallback on CUDA device instead of CPU #6261...,https://github.com/pytorch/pytorch/issues/62619,8f998c4b3fe3e14d86f904239dea73b8f70521ae67bb19081a8149b200498385 references,issue,20433,issue,13246,medium,issue.comments[0].body,Sounds like the cause is #13246 (comment),https://github.com/pytorch/pytorch/issues/20433,bf6a668ac01cb1c5e0a79892d5f191f802d6f5a007028a951bf4109239b0c315 references,issue,84290,issue,36524,medium,issue.comments[0].body,"one, scale=1.0, zero_point=0): super().__init__(dim = dim) Also, what's situation about dim=None for nn.Softmax? In #83163 (comment) and in #36524 is discussed that implicit dim for F.softmax is deprecated. What's the behavior for dim=None? Is then softmax computed across all...",https://github.com/pytorch/pytorch/issues/84290,200504dddc3e993dba226dd401f378b8b730855a7f9d77052fa9fffc76860134 references,issue,84290,issue,71357,medium,issue.comments[0].body,ted issue about multidim for softmax (then global dim=None or dim=() can be converted to global operation across all dims as in other ops): #71357,https://github.com/pytorch/pytorch/issues/84290,f648cd239c646549a8988558e51b6c5612dafcfed3966e12a58f53a510d3ce1a references,issue,84290,issue,83163,medium,issue.comments[0].body,"def __init__(self, dim=None, scale=1.0, zero_point=0): super().__init__(dim = dim) Also, what's situation about dim=None for nn.Softmax? In #83163 (comment) and in #36524 is discussed that implicit dim for F.softmax is deprecated. What's the behavior for dim=None? Is then soft...",https://github.com/pytorch/pytorch/issues/84290,0f52457702774d1cfd21dfaf551f146b79ea5c0ee7679fb91cf166fcec6d20e6 references,issue,42779,issue,40316,medium,issue.body,"ng5 @Xia-Weiwen @leslie-fang-intel @dzhulgakov @VitalyFedyunin @ppwwyyxx @ngimel @z-a-f, @vkuzo. 🐛 Bug As part of the refactoring effort in #40316, I ran some benchmarks on adaptive_avg_pool(2/3)d to see what is the difference in execution time for MemoryFormat::Contiguous vs...",https://github.com/pytorch/pytorch/issues/42779,3206bad0307cc10d7ae9729e74b123db6bb481c26ddeba830185b2e319e8c104 references,issue,92073,issue,71470,medium,issue.comments[0].body,"Same as #71470 ... definitely an issue, and looks like easy fix",https://github.com/pytorch/pytorch/issues/92073,2360c8be664f1c8c5c39400ed45dbbb57ee6cc260b57ec55b8953632bc87f0b1 references,issue,109108,issue,43949,medium,issue.body,"bug['values'], size=bug['shape']) # TypeError: only integer tensors of a single element can be converted to an index Same error appears in #43949 in a completely different context. This error is extremely mysterious and weird. Versions 2.1.0.dev20230802+cpu cc @alexsamardzic @...",https://github.com/pytorch/pytorch/issues/109108,6cf231aab13e7d0d31cfb6ef7c28735c5392d7ebef4f8009c1dadf671da4e968 references,issue,109108,issue,109017,medium,issue.comments[1].body,Related on args inconsistencies compared to scipy: #109017,https://github.com/pytorch/pytorch/issues/109108,770d54ef1403cb64f35cdc104559a6539238a5981335a1c692c822f6f76f2d65 references,issue,92987,issue,92910,medium,issue.body,the multiple-byte data without swapping on a big-endian machine while pickle serialization store data as mostly little-endian (as shown in #92910). Versions PyTorch version: 2.0.0a0+git215f4fc Is debug build: True CUDA used to build PyTorch: None ROCM used to build PyTorch: N/...,https://github.com/pytorch/pytorch/issues/92987,7d8fb9522eabea0d2918248abcf87867251c66edd8568c787f9d735657feaac0 references,issue,81868,issue,24870,medium,issue.comments[0].body,1.2069940567016602e-06 Baseline impl in pix units - Residual with align_corners=False: 0.0 This issue may thus hinge upon #36107 (See also #24870) and may add further justification for allowing grid_sample to work in absolute coordinates (similar to scipy.ndimage.map_coordinates),https://github.com/pytorch/pytorch/issues/81868,829901b7544749ceefc10a82fff19fcf9d33a443548935466fc3694e9e309ff9 references,issue,81868,issue,36107,medium,issue.comments[0].body,gn_corners=False: 1.2069940567016602e-06 Baseline impl in pix units - Residual with align_corners=False: 0.0 This issue may thus hinge upon #36107 (See also #24870) and may add further justification for allowing grid_sample to work in absolute coordinates (similar to scipy.ndi...,https://github.com/pytorch/pytorch/issues/81868,790556383caec20c912b1afe51b5527ef4a1833c1f8c9f5d2954c2f7e5e051c6 references,issue,78618,issue,38208,medium,issue.comments[1].body,"ar = Bar(123) print(bar.x) # prints 123 bar = Bar(nn.Linear(1, 2)) print(bar.x) # prints None dir(bar).count('x') # prints 2 Edit: See also #38208",https://github.com/pytorch/pytorch/issues/78618,08b36cbf886dd0e4c9a30f643dba7ab83293da10982f3774f63136818a25f990 references,issue,96060,issue,75336,medium,issue.body,"torchlib is enough to have the issue. (using official supported cmake procedure) When OpenGL is NOT used, eveything works. A similar issue (#75336) has been reported as well, probably the root cause is the same. If the application is run on a native machine and using a real Nv...",https://github.com/pytorch/pytorch/issues/96060,29751495a66b555fa84de31d45bba1398f40944264d267d2302b1dacd4c4de0d references,issue,92226,issue,91570,medium,issue.body,"rted this to facebook's bugbounty program (as your security policy says), however it seems facebook no longer is responsible here (see also #91570). Facebook's security team came to the conclusion that the issue is not severe. Nevertheless I now have registered these package n...",https://github.com/pytorch/pytorch/issues/92226,a7502848cbb1c289670b5a6d5610527e98f406d57489080209b76d64e618d125 references,issue,94233,issue,30702,medium,issue.body,"🚀 The feature, motivation and pitch Originally discussed in #30702 (comment). One usecase is that it could be used for (implementation details of) broadcasting right by inserting unitary dimensions lambda x",https://github.com/pytorch/pytorch/issues/94233,7c07e6ad05949e17b93f6a799a4ce16e3812e43bf65430201ce411beb5565405 references,issue,112352,issue,112342,medium,issue.body,"fn(): with using_pytree_namespace(""torch.vmap.Size""): _, _ = pytree.tree_flatten(..) Handling Errors Errors should be namespace aware. See #112342 (comment) Inspecting Namespaces print(pytree.show_namespace(""torch.vmap.Size"")) # { 'registrations': # [ # {""cls"": torch.Size, ""fl...",https://github.com/pytorch/pytorch/issues/112352,9088ca78da00b4ead57513f4bd679ef4367e2985f6cf20413973197ca6a9cdb0 references,issue,112352,issue,112342,medium,issue.comments[1].body,See discussion here: #112342 (comment) Summary: Avoid polluting global registrations with custom registrations Persist custom registration info into TreeSpec - This avo,https://github.com/pytorch/pytorch/issues/112352,89e4b22ac29dd7cdc18a148b19fb4079a4da3186134712dafc120dbdf967db83 references,issue,90544,issue,90848,medium,issue.body,"8330 (comment) Figure out why NCCL_DESYNC_DEBUG environment variable is getting set in our tests Fix destruction fiasco in ProcessGroupNCCL #90848 Feature related follow up tasks: Update ProcessGroup.hpp to call Ops.cpp directly, remove the pybinded function definitions for Pr...",https://github.com/pytorch/pytorch/issues/90544,762efe5b67406ec431684a44aad4adbc0364baccbc07144d8de7b345ecf0319c references,issue,86537,issue,35600,medium,issue.comments[0].body,Probably a duplicate of #35600.,https://github.com/pytorch/pytorch/issues/86537,56f8bcad69fb320370e3133a48a7186a43b0c85d1b54cb7ce04cc0e8c91d49fb references,issue,86537,issue,35600,medium,issue.comments[1].body,"riments again, and found that just calling jit trace for a same model could trigger this issue. I've corrected the statement above. I think #35600 is a simpler case of mine because it just compiles the same model each time. It's worth to note that after the simpler one being f...",https://github.com/pytorch/pytorch/issues/86537,a43b90cfe52ef91bca9708f1eb7b9eac5e057ac8c90f2d9dfa5c24b63e2cc525 references,issue,103444,issue,68332,medium,issue.comments[0].body,"Related: #68332 (if we have functional forms, more of these problems with state-passing can at least be worked around in a cleaner way)",https://github.com/pytorch/pytorch/issues/103444,e5fce6f2a107d71bc306ac53cb4e7d0fefc9a11dbb62153aad6196f463eaa4f2 references,issue,99719,issue,14095,medium,issue.comments[1].body,"for an image. I tried a few approaches which each had some limitations: support for this feature has been voiced but not addressed: #30443 #14095 unlike histogram, histc does not support user-specified bins histogram only runs on CPU (see #69519) histc can't be easily parallel...",https://github.com/pytorch/pytorch/issues/99719,d97ad25e5fddff981c4574be88ef8e4e1bf45a276f8f511349bbfdd45fe19ee9 references,issue,99719,issue,30443,medium,issue.comments[1].body,"channel for an image. I tried a few approaches which each had some limitations: support for this feature has been voiced but not addressed: #30443 #14095 unlike histogram, histc does not support user-specified bins histogram only runs on CPU (see #69519) histc can't be easily...",https://github.com/pytorch/pytorch/issues/99719,172ddadc105e8106bf224fc488aab11a5c94e97350ff02b580b51e2d32810c19 references,issue,99719,issue,69519,medium,issue.comments[1].body,"s been voiced but not addressed: #30443 #14095 unlike histogram, histc does not support user-specified bins histogram only runs on CPU (see #69519) histc can't be easily parallelised and vmap even seems to degrade a bit the performance with respect to a for loop A workaround u...",https://github.com/pytorch/pytorch/issues/99719,d96e23bc29d60e08a74208f51904e9a7450d711c4d796170a78c22a2c6674073 competes with,issue,110605,issue,109819,medium,issue.body,"🐛 Describe the bug When using math.isfinite with a tensor a wrong ValueError is issued instead of a TypeError. This creates bugs (see #109819). Fixing this issue will fix the other one as well. >>> import math, torch >>> a = torch.randn(5) >>> math.isfinite(a) ----------------...",https://github.com/pytorch/pytorch/issues/110605,3eea8fdf2e1af658b3542c0d202af134f7ff7b08fcfe906d167b92cbf6eb52ee references,issue,29843,issue,27504,medium,issue.comments[1].body,Dup of #27504,https://github.com/pytorch/pytorch/issues/29843,3d9cde92afad6a75b4e633dd762771a48aa6715eadbc32b7633b806986d9f368 references,issue,50292,issue,49726,medium,issue.body,"that arises if a user employs the @property decorator in their nn.Module subclass. Lots of detail on the problem can be found in #33934 and #49726, as well as here, here, and here. TL;DR: __getattr__ and @property don't play well together, resulting in confusing error messages...",https://github.com/pytorch/pytorch/issues/50292,ba0e90b1e3a3facfc12d20874f7f8f32b8f9a49807d34568d1ab160ff1a1aa5a references,issue,95116,issue,88464,medium,issue.body,"🐛 Describe the bug It is similar to #88464, torch.nn.AvgPool1d can also get negative shape with specific input shape when ceil_mode=True in Pytorch 1.9.0\1.10.0\1.11.0\1.12.0\1.13.0.",https://github.com/pytorch/pytorch/issues/95116,81a0be55e0eb28da1e51abe53bbc7b5ab6674bb92916217fc759f1133cf7b330 references,issue,71409,issue,42502,medium,issue.comments[0].body,Related request: #42502,https://github.com/pytorch/pytorch/issues/71409,4d21f52310cf03a7d2a944881b89ee90f32fabb7a31993d6ba1ff4f580fc9741 references,issue,106596,issue,46948,medium,issue.comments[0].body,The related issue is #46948 on using assert statements to give shape prop hints to the compiler. I propose to use more explicit compiler hints (for torch.compile) and,https://github.com/pytorch/pytorch/issues/106596,88f58a395c4a4a5d8c2ae9aa75d3c779fe611d84683c110d2a6b05e2582b7ca9 references,issue,20102,issue,18182,medium,issue.comments[1].body,Related to #18182,https://github.com/pytorch/pytorch/issues/20102,b1d84938144f4860f8837333e1ca0f54f69ca3bbe246d480ae71a9e50bb692ae references,issue,109946,issue,105319,medium,issue.comments[0].body,"Might be related: #105319 (comment) - usecase of ""masked matmul"" - torch.sparse.sampled_addbmm with CSR format. But I would also support a high-level strided tensor",https://github.com/pytorch/pytorch/issues/109946,17bcc711c6df4d462125d34612d5335d221646838efdadaa2cba49cd48162bd1 references,issue,24870,issue,2732,medium,issue.body,"o single precision floats). Perhaps there should be a conditional cast only if the input grid has less than single precision. New Features [#2732] Support broadcasting in grid_sample and affine_grid. Specifically, for affine_grid, allow it to accept a single 2x3 2D affine matr...",https://github.com/pytorch/pytorch/issues/24870,928d38b508f7b78b8e5be4e54cc5adf5de0f797a3c0069ffc03b2be418089667 references,issue,24870,issue,4085,medium,issue.body,"lly #36107) Implementation: Combinations of these options (residual and/or pixel units) could be added to grid_sample(), or as suggested in #4085, put in a separate flow_sample() high-level function (which should probably use the same kernels as the existing grid_sample()). Ei...",https://github.com/pytorch/pytorch/issues/24870,538c63a7b810cd16d26e53f7859b990bf40593f49b43803c1b116cd7e80c8a25 references,issue,24870,issue,5565,medium,issue.body,"or messages: Produce a more useful error message when the affine matrix has the wrong shape or dtype (#12362), or when it is not a tensor. [#5565] bilinear mode for 3D grid_sample should really be trilinear. Personally, I would prefer to have just a single mode name for linear...",https://github.com/pytorch/pytorch/issues/24870,561ac248ae8d9648afdcbeb07201f3133b2523e12f46b6215846443b3ae266f1 references,issue,24870,issue,19826,medium,issue.body,"image borders. (originally #23925) [#24823] Fix handling of NaN and inf grid inputs, which causes segfaults and corrupted CUDA memory (also #19826, forum/39837). [Fixed in #24929] affine_grid does currently support generating 3D affine grids. Update the documentation to suppor...",https://github.com/pytorch/pytorch/issues/24870,cc279bb87149c8e853419ab649ca2160767a88bbf475d21c76e836ea20a4a3a1 references,issue,24870,issue,21457,medium,issue.body,"eated/tiled in each dimension [DISCUSSION: #25039] Add new interpolation modes, including bicubic, area, and possibly lanczos. [DISCUSSION: #21457] Improve the ability of grid_sample to downsample while warping (that is, when the grid points are spaced out far apart). This wou...",https://github.com/pytorch/pytorch/issues/24870,8ee4f347911ba5b54b2d9c5591000f2f6f7e7356b38ad4e79c3240b3e957563d references,issue,24870,issue,24470,medium,issue.body,"that. It could probably be moved either up directly into the affine_grid function of nn/functional.py or down into the C++ implementation. [#24470, #25014] Profile CUDA kernels vs. cuDNN kernels, and consider dropping use of cuDNN grid sampler and affine grid generator if they...",https://github.com/pytorch/pytorch/issues/24870,c402666479e91b99948a051b109afffb9e8fda14d8a406fe73198a52b43b79d1 references,issue,24870,issue,24823,medium,issue.body,"ix the border gradient issue, which causes training to occasionally get wild results for grid points at image borders. (originally #23925) [#24823] Fix handling of NaN and inf grid inputs, which causes segfaults and corrupted CUDA memory (also #19826, forum/39837). [Fixed in #...",https://github.com/pytorch/pytorch/issues/24870,5f823920c962b922ff0738c12caf8c5dd27179538f36c60d8d12f18da6c91b53 references,issue,24870,issue,25014,medium,issue.body,"could probably be moved either up directly into the affine_grid function of nn/functional.py or down into the C++ implementation. [#24470, #25014] Profile CUDA kernels vs. cuDNN kernels, and consider dropping use of cuDNN grid sampler and affine grid generator if they don’t gi...",https://github.com/pytorch/pytorch/issues/24870,0352dd4ddd8c3b8fd925e1de7efee3d37b1565745395fec86eed51fe5d799f9d references,issue,24870,issue,25039,medium,issue.body,"if not already a tensor. Add a cyclic/circular padding mode, where the image is treated as if repeated/tiled in each dimension [DISCUSSION: #25039] Add new interpolation modes, including bicubic, area, and possibly lanczos. [DISCUSSION: #21457] Improve the ability of grid_samp...",https://github.com/pytorch/pytorch/issues/24870,ba9fb5530749389743e20dccff305ad6225f732ed92328fdd86bfcf21b31b776 references,issue,24870,issue,36107,medium,issue.body,"coordinates, which, again, is prone to user error, (especially when users don’t know which setting of align_corners to target). (originally #36107) Implementation: Combinations of these options (residual and/or pixel units) could be added to grid_sample(), or as suggested in #...",https://github.com/pytorch/pytorch/issues/24870,0e416b0afd5dd5f51c7aa01715a2232d74c77840f2a78a13bf788ae6ea6855c2 references,issue,24870,issue,36108,medium,issue.body,"e grid_sample resolution agnostic, which is important for predicting optical flow at less-than-full resolution (originally #20785, #23923) [#36108] Allow option for sampling using only the residual displacement/flow. Eliminates the need to constantly add an identity grid to th...",https://github.com/pytorch/pytorch/issues/24870,77ece6cfff163463f03b60397917f0cbe90cf24f8d9816a65878dbfecde3dc0f competes with,issue,24870,issue,24470,medium,issue.comments[0].body,"Fyi I moved one item above to be #24470 instead of ""need discussion"". ;) That should be totally fine as long as the perf gap is not large. ;)",https://github.com/pytorch/pytorch/issues/24870,619a60cf87ee10cba20ed4ea6b4b4b42fd81ef3ed9dc57bb5970e7fa57528003 references,issue,40989,issue,29554,medium,issue.comments[0].body,Tangentially related: #32101 #36990 #29554,https://github.com/pytorch/pytorch/issues/40989,2d444c4b68ada779971bdf1907abe43fff316049a6be51fea805ddac0d0e678f references,issue,40989,issue,36990,medium,issue.comments[0].body,Tangentially related: #32101 #36990 #29554,https://github.com/pytorch/pytorch/issues/40989,52d8ebb52ee4c89c061ac316819fda98bf58e2b527af973e4fb3bc6b4d21d876 references,issue,109923,issue,2575,medium,issue.body,"ssages sent to stderr or stdout. Yet, I'm speculating if this crash might be hiding a Thread-Local Storage (TLS) issue, as documented here: #2575. SEGFAULT BT: Program received signal SIGSEGV, Segmentation fault. 0x00007fffc24d242b in ?? () from /lib/x86_64-linux-gnu/libstdc++...",https://github.com/pytorch/pytorch/issues/109923,6049bd6b8676fabb28a909c61a23646905d1157e8ad37d97e138c321387d67df references,issue,109706,issue,80821,medium,issue.comments[0].body,Related to #80821,https://github.com/pytorch/pytorch/issues/109706,03970f8a1adbcad2879df3fbf2f2368aa53ccc085e45ee7b30eaa48277b337d1 references,issue,103581,issue,101699,medium,issue.comments[0].body,a bi related request on shareable data frames and string arrays: #101699,https://github.com/pytorch/pytorch/issues/103581,aa1b0c722e5c9420bf07b0275f6ad19446841943ad2dec45ec0c6d997db38d9f references,issue,43949,issue,33041,medium,issue.comments[0].body,related: #33041,https://github.com/pytorch/pytorch/issues/43949,2cb3235247524f5577b69cab464e0be76b4fd3136068d8a17cf39f1e08fe5fdd references,issue,105494,issue,124423,medium,issue.body,"loop. Apparently, torch.arange() internally uses .item() and vmap cannot handle that. Is there any workaround? This issue seems similar to #124423, but it is not exactly the same. cc @zou3519 @Chillee @samdow @kshitij12345 @janeyx99",https://github.com/pytorch/pytorch/issues/105494,a5b9ade24b7f8b3fa8fe021ecc70590d91f04b7b6a9e77e011fab53fd65a96c0 competes with,issue,106959,issue,68616,medium,issue.comments[0].body,"tensor.repeat. torch.tile also does not have inplace/out-variants and uses inconsistent arg name dims instead of dim, same as flip/roll :( #68616",https://github.com/pytorch/pytorch/issues/106959,22560d91a0f00b78eed9af7096584b02de1f7b52ec1c9c282b7b023184c001bc references,issue,87358,issue,9222,medium,issue.body,"sparse matrices. I indeed only incidentally found out about torch.triangular_solve supporting BSR inputs when reading about it in an issue #9222 (comment). I then found out about CSR support by ""random"" trial and error. Additionally, the gradient computed using torch.triangula...",https://github.com/pytorch/pytorch/issues/87358,e5f715b581ad4f3578a8bd4f3271e2d17f05c6ca747071c07a2f5394ef6e5894 references,issue,87358,issue,10043,medium,issue.body,o what I suggested here for generic sparse matrices: #69538 (comment) Additional context Related issues include: #53441 #9222 #69538 #28341 #10043 #87085 Rough test script: # Try with the nightly version of PyTorch !pip3 install --pre torch torchvision torchaudio torchtext --e...,https://github.com/pytorch/pytorch/issues/87358,40cf0ef5d4b279d6f2c5758eefb3e51f220602f3148fe1dda067944f52c2bdd9 references,issue,87358,issue,28341,medium,issue.body,milar to what I suggested here for generic sparse matrices: #69538 (comment) Additional context Related issues include: #53441 #9222 #69538 #28341 #10043 #87085 Rough test script: # Try with the nightly version of PyTorch !pip3 install --pre torch torchvision torchaudio torcht...,https://github.com/pytorch/pytorch/issues/87358,ca280d4b39794362751dc6049b6ed848b10137d5a004bc6df291024b410bd30f references,issue,87358,issue,53441,medium,issue.body,ustom backward op similar to what I suggested here for generic sparse matrices: #69538 (comment) Additional context Related issues include: #53441 #9222 #69538 #28341 #10043 #87085 Rough test script: # Try with the nightly version of PyTorch !pip3 install --pre torch torchvisi...,https://github.com/pytorch/pytorch/issues/87358,65e7a839ea393a7c4ef6a9927c996d09a3337d6b5f07ffeace3fcd4c1a095101 references,issue,87358,issue,87085,medium,issue.body,I suggested here for generic sparse matrices: #69538 (comment) Additional context Related issues include: #53441 #9222 #69538 #28341 #10043 #87085 Rough test script: # Try with the nightly version of PyTorch !pip3 install --pre torch torchvision torchaudio torchtext --extra-in...,https://github.com/pytorch/pytorch/issues/87358,b12039cfc8831bb2ac4e01990279c8577327adb4e43900e3c497999ffd11a4a6 references,issue,97210,issue,82886,medium,issue.body,"../c10/cuda/CUDAGraphsC10Utils.h"":73, please report a bug to PyTorch. Unknown CUDA graph CaptureStatus32729 Originally posted by @cdmssa in #82886 (comment) cc @mcarilli",https://github.com/pytorch/pytorch/issues/97210,05214dd785406d81639e0ab4a62a1efeea61b98325a1ec34f8e5bbed03564995 references,issue,108858,issue,44380,medium,issue.body,"an the laptop running Windows as well. So what is the difference that makes it so I can run the model on Windows but not on Ubuntu? I found #44380, which implies that it shouldn't even be running on the 11GB card let alone the 4GB card since falling back to system memory isn't...",https://github.com/pytorch/pytorch/issues/108858,aa92d9d290dcad2cac2adc40bacec7acf2b1beb08355debd8ccb567e116e433f references,issue,99918,issue,93880,medium,issue.body,"🚀 The feature, motivation and pitch RFC: Debug Mode Inspired by issue #93880, this RFC proposes to design a DEBUG mode for PyTorch eager mode. A lot of CUDA kernels fail with device assert but understanding this asse",https://github.com/pytorch/pytorch/issues/99918,c2492ef0ce88d6e5af0a6b641363f5c63dd82d321a71445c75d99fc556d9df26 references,issue,99918,issue,99640,medium,issue.comments[0].body,s that are not at native op boundary. For example pre-condition checking in some autograd formulas here EDIT: and also at higher level like #99640. I think it would be good to: Make sure the API works well with a non-Mode API. I guess global flag-based. This kind of C++ API fe...,https://github.com/pytorch/pytorch/issues/99918,341cf913e2c5c0bb1feff40898da5da44fe81c5d955a49fb96f34ec7c7c7469b references,issue,44989,issue,45009,medium,issue.comments[1].body,Already done :) #45009,https://github.com/pytorch/pytorch/issues/44989,23ea1dd50c402c5310ade10df8fc88a66ae71c01c7d93c19639f55d1de1c53ed competes with,issue,44991,issue,44989,medium,issue.body,Two versions exist for seemingly same functionality: tensor.unfold supports Long tensors (F.unfold does not: #44989) tensor.unfold returns different shape layout tensor.unfold has differnet argument names: size/step instead of kernel_size/stride tensor.un,https://github.com/pytorch/pytorch/issues/44991,f397b3c2bc8ec4a9c6b53cd8ef4bce22945d9fb02b4114d944afe0783cbd39ea references,issue,104193,pr,84843,medium,issue.body,minimum number of zeros Approach 2: allow a variable number of specified elements in batches A prototype of this approach is implemented at #84843 The example tensor y defined above can be represented as a batched CSR tensor uniquely: >>> z = torch.tensor(y).to_sparse_csr() >>...,https://github.com/pytorch/pytorch/issues/104193,701ca80ac05110159c03035837bc7d2954eb00a3acd22c4c4ba883fa841197e6 references,issue,34646,issue,30387,medium,issue.comments[0].body,Can Storage.from_buffer be used now for this purpose? #30387 and I wonder what happens about freeing the memory. pytorch/torch/csrc/generic/StorageMethods.cpp Line 87 in 74ce3a0 static PyObject * THPS,https://github.com/pytorch/pytorch/issues/34646,9b26320d25f59dfbec5c031c28ca9d7707bc55e7fe273f75e618c4abdbea67f7 references,issue,23756,issue,46168,medium,issue.comments[28].body,"e ""level"" at which you implement the feature: either torch.nn or native functions. > inplace gradient computation for sequential container: #46168 As you mentioned above, this can lead to other issues. Depending on how the sequence container works, we can reconsider that later...",https://github.com/pytorch/pytorch/issues/23756,70c6e1695fadd272ca97453539df343c02cc59819c89d8fc40341f93ef51e9fa references,issue,106584,issue,28090,medium,issue.body,🐛 Describe the bug I guess this is just another facet of #28090 flatten could realloc only the needed amount needed to deal with the required dimensions (and keep some unaffected dimensions expanded). In,https://github.com/pytorch/pytorch/issues/106584,712560db44d238c49d37a8c4dff4aacb8fa01014b1b652a89a0349c5ea93ba88 references,issue,106584,issue,106614,medium,issue.comments[1].body,@cpuhrsch So far I'm getting trouble from instability of selected dynamic / static kernels in #106614 (e.g. I want a way to force static shapes and to control which dims are static and which dynamic). Also seems currently not very good for t,https://github.com/pytorch/pytorch/issues/106584,dfe8ddf69531265ffdfd4b69b005d180fe519ee5e5f5f81abbfc49ea6caa621b references,issue,104506,issue,88103,medium,issue.comments[1].body,Seems relevant to #88103,https://github.com/pytorch/pytorch/issues/104506,1b5d4a5576dddc31cbf22133a6f08b444a7ad9b04264e4e0231379b573d1b245 references,issue,106713,issue,106717,medium,issue.body,local GPU development: #106714 Document process for building locally + codespaces Add to contributing guide Add support for MPS based dev: #106717 Add support for AMD based dev: #106718,https://github.com/pytorch/pytorch/issues/106713,29e2cdc467853f32278a5ef358bb0cd7765da3bfbf43b50e67f146e1c1526f36 references,issue,106713,issue,106718,medium,issue.body,t process for building locally + codespaces Add to contributing guide Add support for MPS based dev: #106717 Add support for AMD based dev: #106718,https://github.com/pytorch/pytorch/issues/106713,f20e1699cfbb155c05833e0c8190a5cd1aabf518247a122491b3326d17f09b40 references,issue,103352,issue,29973,medium,issue.body,"Only relevant for large tensors/loops, where materializing a python list first takes too many python objects Related on slow item indexing: #29973 and proposal of tensor.item(i, j, k, ...) fast indexing method returning Python int/float objects without tensor[i, j, k, ...].ite...",https://github.com/pytorch/pytorch/issues/103352,a2775975671b90a2e85004bd06258281ec745c9d7cbdf76a39d75c0980afb189 references,issue,103352,issue,43949,medium,issue.body,"ut tensor[i, j, k, ...].item() first creating an extra tensor object and only then upacking it Related on supporting memoryview on tensors: #43949, then this method could be implemented by simply returning memoryview which supports iteration in python Alternatives No response...",https://github.com/pytorch/pytorch/issues/103352,409270e2a95bf503eac7e13ca6d797c886d964a3f9971c6acacc0ff16aff2048 references,issue,103352,issue,101699,medium,issue.body,"103339 (comment) so this is needed for 1d tensors, although could be useful in the future in other contexts if string arrays are supported: #101699 Only relevant for large tensors/loops, where materializing a python list first takes too many python objects Related on slow item...",https://github.com/pytorch/pytorch/issues/103352,831a4d6289a717e4d1354737aa95724f45a767deb65ade91f2abfd15be0774a7 references,issue,5740,issue,5405,medium,issue.body,"Note that #5659 only fix the bug reported in #5405 (thanks for @zou3519 pointing out, I open a new issue about that). Do we need scalar input to scatter_add_? I think the implementation may",https://github.com/pytorch/pytorch/issues/5740,dead7c94f4fb59ed0005e7104cb5f3e35c62b795ff3f6808a56a4d3f6d3b5e92 references,issue,106129,issue,106128,medium,issue.body,"merous issues with torchscipt optimization and have had to disabled it to avoid breaking the code. I recently debugged one such issue here: #106128 torch::jit::setGraphExecutorOptimize sets a thread-local variable, but the model may be run on a different thread. This was diffi...",https://github.com/pytorch/pytorch/issues/106129,36742d27def3a7ffbc848c94174e639cfed3f777c7775ce028c39b676b4c7c91 references,issue,106297,issue,106206,medium,issue.body,ild/lib/libtorch_cpu.so: undefined reference to `google::protobuf::internal::ThreadSafeArena::thread_cache_' Reason for this is the same as #106206: The generated protobuf files are not correctly linked against the protobuf::libprotobuf target which would have defined PROTOBUF...,https://github.com/pytorch/pytorch/issues/106297,a12ee0916270e7876cad34229823e2ecc4b8c3dd8f2a59c8aa3da5b35ef0ffc6 references,issue,98406,issue,98073,medium,issue.comments[0].body,#98073 Add privateuse1 folder in aten.,https://github.com/pytorch/pytorch/issues/98406,467f2b1afda52ca266d0ad1d387b5b4f8778a11c406268ea3eebb2aa407857ea references,issue,40373,issue,26889,medium,issue.body,"Editor's note: See #26889 for direct type system approaches. This issue is to discuss use of assert to direct type refinement, either in mypy, or also in TorchScript",https://github.com/pytorch/pytorch/issues/40373,3281e64c3f6fddbc194424e1eeffa62e30e629f22e209f7e547daa3337730f3f references,issue,90365,issue,47964,medium,issue.comments[0].body,Likely due to #47964,https://github.com/pytorch/pytorch/issues/90365,4ab2419e534ec74b39abe818290695b4c4f31f00496a9fc2c367faca4b818696 references,issue,103756,issue,102533,medium,issue.comments[0].body,I just stumbled over #102533 and #71283 and found that the stack trace in these older issues seems to be essentially similar with references to __do_global_dtors_aux an,https://github.com/pytorch/pytorch/issues/103756,2f0d50a94f764ecdda23748b60985d27e72ab5c73ba705f011e94a4b16f86258 references,issue,98947,issue,98861,medium,issue.body,se tensors is the ability to torch.cat them together. Currently this is not supported: This issue tracks the failure w/ reproduction steps: #98861 Alternatives No response Additional context No response cc @alexsamardzic @nikitaved @pearu @cpuhrsch @amjames @bhosmer,https://github.com/pytorch/pytorch/issues/98947,9f40841c276201c5f01329533f5c9f41024b546dcaa585e20e19cbe6cfef4e1f references,issue,105457,issue,104623,medium,issue.body,📚 The doc issue Related to #104623 i propose a new top level glassy that defines concepts that a PyTorch developer is interested in - the rest of the documentation could then,https://github.com/pytorch/pytorch/issues/105457,97c459128d8d43cc4965b54247605d5d3e996dbc5e884df0d5a1a79fc93ffa8f references,issue,105457,issue,100904,medium,issue.comments[1].body,This issue is similar to #100904. This difference is that I was hoping to use this issue to introduce a glossary into the documentation that is for the users only so it is,https://github.com/pytorch/pytorch/issues/105457,9d052939a951729b8fd1f1c11e360f49f73275cc7a0ea1859936d939630b4208 references,issue,105329,issue,105319,medium,issue.comments[0].body,"could use bsr_softmax from torch.sparse._triton_ops. Could you please tell more about your use case as well, especially in conjunction with #105319?",https://github.com/pytorch/pytorch/issues/105329,4022aaaa5efbcd057bcb40f4faf88c2c65870544af1ed26ec254ec43d28fb8dc references,issue,25104,issue,92927,medium,issue.comments[0].body,"d_for_backward"". Then it can sometimes be used to avoid saving random tensors for backward and just reproducing them on the fly: e.g. as in #92927",https://github.com/pytorch/pytorch/issues/25104,2cfbeecafb0c6cb84e0ec3cc2e41cc8fda87969b1605a682fd621c9517187068 references,issue,51135,issue,50444,medium,issue.body,"compiled TorchScript custom classes, our compilation (type-checking) surface is smaller, so less likely to break existing codes Related to #50444 @gmagogsfm This one is not one of the trivial fixes, probably need some discussions. cc @gmagogsfm",https://github.com/pytorch/pytorch/issues/51135,f0a84beb4f898bdba6d53da30117223be8366951029421dfdb4e24721871317d references,issue,96085,issue,33181,medium,issue.body,"🚀 The feature, motivation and pitch There seemed to have been some discussions on alternative ways to supply collate function (like in #33181). I also found is slightly inconvenient to add custom way to collate metadata on my dataset. My pitch to solve this would be to add che...",https://github.com/pytorch/pytorch/issues/96085,f54273cc0f7c0bdcc8cd5d1a44d82251c716d3d5465b0ea7202fc6e3ea31b31b references,issue,103588,issue,50688,medium,issue.comments[0].body,Maybe related: #50688 #50688 (comment) (the bit on unrolling multiple time steps into a single kernel),https://github.com/pytorch/pytorch/issues/103588,51ef5e2f9040e744e1b0b4898810026cf496ed40520b001b4241b92fbfd98f48 references,issue,94586,issue,60832,medium,issue.body,"would work for such a fundamental low-level function, esp since named tensors have been around for many years now. Are they actually used? (#60832) Alternatives No response Additional context No response cc @zou3519",https://github.com/pytorch/pytorch/issues/94586,b4c61afa4a63c8aa9b6cf2ebba8e26761d33c23a1c6d8fbf3f5c2a7fb2cbd7b2 references,issue,94869,issue,52439,medium,issue.comments[0].body,And unfortunately no easy way of logging / forcing the algo choice for a given op call :( #52439,https://github.com/pytorch/pytorch/issues/94869,86a2c9954167e6bf307c9226fea22ff6feefce177fca6522247511bd5cf5ff9c references,issue,102269,issue,101850,medium,issue.body,"trying to find what difference in setup/architecture is causing this weird behavior. Thanks in advance for any help Possible duplicate of: #101850, but also breaks normal multiprocessing Versions $ python collect_env.py Collecting environment information... PyTorch version: 2....",https://github.com/pytorch/pytorch/issues/102269,a48519b864f3b33a0b4badf0f67f5b4641c4eb5fc7ad1f422989822a88858c7d references,issue,79337,issue,78681,medium,issue.body,"🐛 Describe the bug A similar error occurred before #33103 Another related issue about pytorch nightly on macOS #78681 Platform: macOS 12.4, Apple M1 Max Conde version 4.13 Conda config: channel_priority: disabled When installing pytorch only the expected ni",https://github.com/pytorch/pytorch/issues/79337,f8ab26ec55ae2a7941c0f5373804c164c0549294af4a77a0df2538d043c61a28 references,issue,104472,issue,58833,medium,issue.body,""", ""_foreach_log2"", ""_foreach_log"", ""_foreach_pow"", ""_foreach_sqrt"", ): value_range = {""low"": 0.5, ""high"": 1.0} else: value_range = {} Rel: #58833 #102409 cc @soulitzer",https://github.com/pytorch/pytorch/issues/104472,241e9794e44cbc79b49ae6f591a79306deb8dc4adc5e022667370ce8a7d587fa references,issue,96766,issue,96764,medium,issue.body,"default, add an API to disable #96866 update docs to mention the early stop feature #96866 update docs/tutorial to feature nested use cases #96764 Missing anything? cc @ezyang @albanD @zou3519 @gqchen @pearu @nikitaved @lezcano @Varal7",https://github.com/pytorch/pytorch/issues/96766,d07646b5e8a02178f9d65656e355b16a27fd9261c2ed32e9ac32891363949682 references,issue,103803,issue,103798,medium,issue.body,"🚀 The feature, motivation and pitch This is a sub-task of #103798. We have functions annotated to accept List[int] but actually get torch.Size passed in, e.g.: pytorch/torch/nn/functional.py Line 2435 in 9",https://github.com/pytorch/pytorch/issues/103803,8cc46e46a997346b37ee4b3b2b5a5b88e1e6486dba016159e0a540d42aaa4142 references,issue,103798,issue,103761,medium,issue.body,🐛 Describe the bug This is a sub-task of #103761. This is not user-facing as this only involves functions prefixed with an underscore. A major problem is JIT cannot understand torch.Size o,https://github.com/pytorch/pytorch/issues/103798,af44556ea33c295ce34bd3c3cf9f95c311c57a2871632773a90afcff227ac047 references,issue,103798,issue,103803,medium,issue.body,"r problem is JIT cannot understand torch.Size or Tuple[int, ...]. This problem is relatively self-contained and tracked in a separate issue #103803. The _list_with_default is used as this, where input.size() is torch.Size: pytorch/torch/nn/functional.py Line 1135 in 918fe51 ou...",https://github.com/pytorch/pytorch/issues/103798,f0f39f574de13d370647f4fbad193f3a4547c4e925d28708b96b36fe6955efa7 references,issue,103787,issue,103761,medium,issue.body,"🚀 The feature, motivation and pitch This is a follow-up of #102918 and a sub-task of #103761. This is not user-facing and only involves changes in the type stub of torch._C._nn which is an internal module. #102918 has introduced ann",https://github.com/pytorch/pytorch/issues/103787,b32b91d6db77d1100a4bc4c7b0adf065b794c79f239289c5cb21a34d92f0391d references,issue,78487,issue,51803,medium,issue.comments[0].body,"Tensor out, torch.dtype dtype, torch.layout layout, torch.device device, bool pin_memory, bool requires_grad) I suspect this is related to #51803 with a potential solution provided by ModelTC/MQBench#82.",https://github.com/pytorch/pytorch/issues/78487,a7ab23dc2ba12f9ad1a11cacb73a9d5f96cf46e56db7f9336590e0953419f535 references,issue,101110,issue,37334,medium,issue.body,"formations. This maybe can also help with this issue, where some users want a model graph without having to provide concrete inputs values: #37334. Alternatives Continue using the status quo with the buggy parsing of torch jit traces Currently, the trace parsing discards Class...",https://github.com/pytorch/pytorch/issues/101110,cc3716f8a00baef3e0d20bde8e6dc239c4ccbd2ae8fb2edc33c8ed0e43457598 references,issue,103393,issue,103375,medium,issue.body,"🐛 Describe the bug This is related to #103375 and #103376, but I assume it's better to split into smaller fixes. Some of the dunder ops are not defined in _C._TensorBase but directly in",https://github.com/pytorch/pytorch/issues/103393,426523751fefc678adcb7b60ea479ae9abcbf761c303a5082258ace619cbe557 references,issue,30386,issue,22281,medium,issue.comments[1].body,"round since 1997, I do not find strong evidence of EG being used in deep learning. @fjanoos @bamos -- Do you have other references? Tagging #22281 for constrained optimization",https://github.com/pytorch/pytorch/issues/30386,965c7a9ea9a1652f9529198fd0d0974fc4094f3ceee06590af521dc3e0cef6f6 references,issue,102977,issue,16797,medium,issue.body,"motivation and pitch Currently when loading an object with a torch CUDA tensor on a CPU-only machine, torch will raise a RuntimeError like #16797 mentioned. A problem in this situation is that torch will spend time (maybe trying to load the tensor bytes) before raise that Runt...",https://github.com/pytorch/pytorch/issues/102977,17c41072b5e46990c51c83785c12a98fc3d01eb6a34ee66aa9462d0ef7273226 references,issue,102730,issue,18182,medium,issue.comments[0].body,Related long-standing issue: #18182,https://github.com/pytorch/pytorch/issues/102730,841cd4f11af2b9db916889f00559cd8ac8b25393b0b6fc31408e3093890fc452 references,issue,102971,issue,94691,medium,issue.body,"02911 stated above is related to: Issue 96416 (Loss.backward() error when using MPS on M1 #96416), Issue 94691 (Nan is output by GRU on mps #94691), Issue 97552 (PackedSequence failure with MPS #97552), and PR 96601 ([MPS] LSTM grad_y missing fix #96601). cc @kulinseth @albanD...",https://github.com/pytorch/pytorch/issues/102971,57e9dad727e18879408bff269fe09a549576bbe64b2416b027bc7d72b3e54f39 references,issue,72948,issue,55655,medium,issue.body,"rograms for nondeterminism or using deprecated operations? How do we make it easier to implement new types of messaging and control it (see #55655)? There are also issues with our current messages and other logging systems, like VLOG (#57118) and GLOG (#14724). This issue prop...",https://github.com/pytorch/pytorch/issues/72948,fd06abf15b917559344343f140a3432a31794313a64508c5b4180d4bab9cc063 references,issue,72948,issue,57118,medium,issue.body,"t new types of messaging and control it (see #55655)? There are also issues with our current messages and other logging systems, like VLOG (#57118) and GLOG (#14724). This issue proposes we add a messaging system to PyTorch that address these issues. It should have the followi...",https://github.com/pytorch/pytorch/issues/72948,68a7464e9ff560b01e264a74a326e2115a98f77d27cb0a458109054049b4c1c7 references,issue,72948,issue,68768,medium,issue.body,"alent for the Python UX? Can libraries extend our messaging system? How can we avoid sending a deluge of unnecessary messages to users (see #68768), or help them consistently upgrade a class of warnings to errors so they better vet their programs for nondeterminism or using de...",https://github.com/pytorch/pytorch/issues/72948,89ebad5871dbe1a113aa47f0a7e5902edb7ac1644415ac6a5c26d006b749a41a references,issue,99982,issue,52675,medium,issue.comments[1].body,n try to run the two setups under strace (or maybe ltrace?) to find any such differences Some related issues / discussions are linked at in #52675. Please chime in and add your usecase there too :) My main suggestion there is that torch.cuda.is_available() and related function...,https://github.com/pytorch/pytorch/issues/99982,9625c0165cd4fcc2ffca3f84f99c2c489cc18f42e0418289ab32d9914d76c50f references,issue,102078,issue,88,medium,issue.body,eType_4.cpp:361 #87 0x7f56d34ee173 in _embedding_bag_dense_backward /home/user/pytorch/torch/csrc/autograd/generated/VariableType_4.cpp:363 #88 0x7f56d3759e45 in operator() /home/user/pytorch/aten/src/ATen/core/boxing/impl/WrapFunctionIntoFunctor.h:13 #89 0x7f56d3759e45 in cal...,https://github.com/pytorch/pytorch/issues/102078,0687d107c59293f4222f0ddf2e3648cee1214d9f7d30d79dc3c7a726b62b3513 references,issue,68340,issue,67955,medium,issue.body,"🐛 Bug Hello. I was having a MAGMA problem and updated my torch version as was recommended in #67955 to 1.10.0. Now, the code that ran in 1.9. is raising this function error: File ""/home/luis/Desktop/cog/vqgan-clip-generator/src/vqgan_clip/",https://github.com/pytorch/pytorch/issues/68340,6ccd7dbde9dfbaaf02e6aff82fe73f086c13b038fcb1cbd8319efc141a7fac04 references,issue,10615,issue,44380,medium,issue.comments[1].body,You might want to save the failed memory in 'shared' or UVM-memory? #44380,https://github.com/pytorch/pytorch/issues/10615,72f045b5c901760e3e9f7c41996aea58ccefa0e37c22c7e1e4772fbcba63f48b references,issue,101359,issue,97068,medium,issue.comments[0].body,"ess this issue in near future, @jiayisunx please take care of this one. Please notice that float16 optimization on CPU device is still WIP: #97068, which means the current status is still not good. Another problem is that we do not have support for F.scaled_dot_product_attenti...",https://github.com/pytorch/pytorch/issues/101359,7716c0d3049d867b2b12f8ce931752a73cd513a9205005179a1458c26efab693 references,issue,81749,issue,46642,medium,issue.body,"takes into consider phase information in speech processing, and would like batchnorm for complex tensor to be possible. Similar to #47052, #46642 Additional context Are there any special points to consider when adding/implementing this feature because of the complex type? cc @...",https://github.com/pytorch/pytorch/issues/81749,e03e6c1c6a90e05102c46fc4a3e1a2c814e130c779755bc36a208c9a7a924d6b references,issue,81749,issue,47052,medium,issue.body,"ork that takes into consider phase information in speech processing, and would like batchnorm for complex tensor to be possible. Similar to #47052, #46642 Additional context Are there any special points to consider when adding/implementing this feature because of the complex t...",https://github.com/pytorch/pytorch/issues/81749,30d759b8fba389219f8f959db7167794f05f410d68fe420a7cdc21737a150572 references,issue,101159,issue,37250,medium,issue.body,"when checking with nvidia-smi it shows 22GB. I have tried many methods, such as upgrading both Torch and Cuda, but the bug still persists. #37250 Versions env OS:unbantu Python:python3.9 Transformers:transformers==4.27.1 PyTorch:torch==1.10.1 / torch==2.0.0 same CUDA Support (...",https://github.com/pytorch/pytorch/issues/101159,e6be8c6fdbd377c83c3dd17cc545934f7fba3490d55337fd04005cfb2a4b56a7 references,issue,101159,issue,37250,medium,issue.comments[0].body,#37250 zai-org/ChatGLM-6B#990,https://github.com/pytorch/pytorch/issues/101159,e50109ddb17772b56625923a2409f368c3eb69c2b158f26ecc85f00706ad5563 references,issue,35633,issue,31252,medium,issue.comments[1].body,Related #31252,https://github.com/pytorch/pytorch/issues/35633,09833b22ec33d3e8c4d0bb1a23dbfd19339e820f02fa292adde60189c2e37aad references,issue,98414,issue,63079,medium,issue.body,"🐛 Describe the bug Similar to #63079 but this is for test_nadam under test_optim instead. I've done some initial investigations on other nodes (8xT4,4xV100,4xA100) and other CU",https://github.com/pytorch/pytorch/issues/98414,3b5a6b35048241d223aa3b0a0737683ba2f504c346d00886ab9ccbf8958045f5 references,issue,98200,issue,65049,medium,issue.comments[1].body,Related requests in #73640 / #65049. We'd still accept a PR doing this :),https://github.com/pytorch/pytorch/issues/98200,308241a83d3b3baabb097c48ccc70869fd8a4d45292458e45859cef441cbc0ee references,issue,98200,issue,73640,medium,issue.comments[1].body,Related requests in #73640 / #65049. We'd still accept a PR doing this :),https://github.com/pytorch/pytorch/issues/98200,a35a9c308bfa7cb4e34de48d8331431e1f7435dccb2e1bf08f86287d16034c62 references,issue,99561,issue,88,medium,issue.body,:44.005 A #87 pc 00000000001fdada /system/framework/framework.jar (android.view.ViewGroup.dispatchTransformedTouchEvent+118) 12:57:44.005 A #88 pc 000000000020a254 /apex/com.android.art/lib64/libart.so (nterp_helper+3924) (BuildId: 28c5aa8a2e8fc5df069f717d6e94f7fe) 12:57:44.00...,https://github.com/pytorch/pytorch/issues/99561,1794af42e8a2a01357e70b385e8ac5ab68379d37de6f62f9e060a05e700ab2ed references,issue,99410,issue,73176,medium,issue.body,🐛 Describe the bug Though #73176 lists a lot of APIs lacks checking out of bound in cuda. It still can be repro in torch.nn.functional.multilabel_margin_loss. import torch,https://github.com/pytorch/pytorch/issues/99410,e44ac75d33d4bf59812a74b051c85fdaaee2c2853137c5e740ac886add823076 references,issue,99351,issue,99265,medium,issue.body,"🚀 The feature, motivation and pitch See #99265 for an example of some of the chaos this causes. Position-only arguments were added in Python 3.8 and we recently dropped support for 3.7 s",https://github.com/pytorch/pytorch/issues/99351,b861775c5590957aa92480380d43ad5d67f0310ea2352c24ae12ebf650feee29 references,issue,99042,issue,71149,medium,issue.body,"ftplus(sig)) The calculation of eps is copied from the actual implementation of MultivariateNormal. I thought that maybe this is related to #71149, but I don't know. Unfortunately, I can't easily test this on cpu since the framework I'm working with (fastreid) isn't exactly bu...",https://github.com/pytorch/pytorch/issues/99042,e57770f23b25413b91b09dacb0814934f6e1aee46ec2adac2f82f700faa75925 references,issue,98541,issue,98537,medium,issue.body,"m"", ""CPU"") def my_sum(*args, **kwargs): return args[0] x = torch.tensor([1, 2]) self.assertEqual(torch.ops.foo.sum(x), x) Likely related to #98537, unsure if same exact issue. Versions master cc @ezyang @bhosmer @smessmer @ljk53 @bdhirsh",https://github.com/pytorch/pytorch/issues/98541,15e35aaf710b98dded2ed1240dd03b8457005687f75d6b44f9d3fb61336e71f5 references,issue,88464,issue,88144,medium,issue.body,"n potentially be negative, which is nonsensical. I think the calcuation of output shape for that case is wrong. This seems related to issue #88144 import torch m = torch.nn.MaxPool1d(kernel_size=4, stride=1, padding=0, dilation=1, ceil_mode=True) input = torch.randn(20, 16, 1)...",https://github.com/pytorch/pytorch/issues/88464,13d63a936fbde840351c771944ebc8f95692ed0dabf2da9af375876c456da383 references,issue,88464,issue,88144,medium,issue.comments[1].body,"The PR #88421 for issue #88144 can also fix this issue. The example above can run pass and output shape will be the same as input -- [20, 16, 1]. I am fixing UT failure i",https://github.com/pytorch/pytorch/issues/88464,752d6fdfe1a88e4d88c4b10be1e772afd751871c0132b0e2d6e7246f20811e18 references,issue,98169,issue,77764,medium,issue.body,"y implemented for the MPS device. If you want this op to be added in priority during the prototype phase of this feature, please comment on #77764. As a temporary fix, you can set the environment variable PYTORCH_ENABLE_MPS_FALLBACK=1 to use the CPU as a fallback for this op....",https://github.com/pytorch/pytorch/issues/98169,6ca50c368a6569b9b9201fda08710c9eb8219d30f4589bdecb285a8223c0db26 references,issue,55375,issue,57534,medium,issue.comments[0].body,Any progress on this issue (and the related #57534)? This is currently a big problem for us and it is increasingly hard to work around it...,https://github.com/pytorch/pytorch/issues/55375,f1f59e2dc7e71e891fef55e095d13a58b057dbf767908294750318a87639caa8 references,issue,97500,issue,96319,medium,issue.comments[0].body,"backward() That being said, there is some discussion around making a deeper refactor to move/having another version counter below autograd. #96319",https://github.com/pytorch/pytorch/issues/97500,88c8a2cc958e04933fae3d18b47043028483f129fd1fb5fcaf9906c6b02fbcaa references,issue,89817,issue,66247,medium,issue.comments[4].body,I have also seen the same issue in #66247 and had to resort to building my own PyTorch container with `USE_MKL=OFF` in order to resolve this issue. My stack trace was also similar t,https://github.com/pytorch/pytorch/issues/89817,caf20057ce589984783e2b675e0309c24cfa1d35318ff9f2dfac87995d67f3e9 references,issue,80017,issue,80104,medium,issue.comments[1].body,"ocesses, ), nprocs=tot_processes, join=True, ) if __name__ == '__main__': main() TODO: we should error out with a more descriptive message, #80104",https://github.com/pytorch/pytorch/issues/80017,fa4ca5a98ee6dec88825f1dd52481af88c1ed0aae1a0b0ce0723d012b9e00b85 competes with,issue,33248,issue,30900,medium,issue.body,"either lack reproducibility or are classified as user error since the programmer attempts to use CUDA on multiple processes in a tree. See #30900, #2517 and #3491. In our environment, using spawn instead of fork to create the child processes does not resolve the issue. cc @ezy...",https://github.com/pytorch/pytorch/issues/33248,6659a52a76787da9cb7bb163081efdb69ad98fe4e3b070bc91abcfb7e670a0e1 references,issue,97196,issue,79967,medium,issue.comments[0].body,Related: #79967,https://github.com/pytorch/pytorch/issues/97196,5db4ab86fd68e402e5dda3909f21609dbb8f4c464ad9d5083dc634d5cc9dc157 references,issue,95776,issue,69786,medium,issue.body,"h the following code. If the commented codes is enabled, the result will output False, which is as expected. In some previous issues (e.g., #69786), there is a similar lack of support for sparse tensor in torch.equal and other APIs. It seems that there are still many APIs that...",https://github.com/pytorch/pytorch/issues/95776,c859896faea9df6fb9a783958b6ee578bbff533e0de3ee509c7ce0bebe4ad38c references,issue,95756,issue,71636,medium,issue.body,"an. I'm not sure if that means some kind of overflow happened in the calculation. I think the root cause of this issue should be related to #71636. To Reproduce import torch torch.manual_seed(0) input = torch.randint(-1,1,[0], dtype=torch.int32) # input = torch.randint(-1,1,[0...",https://github.com/pytorch/pytorch/issues/95756,49814f6e29d5f277c2c123c4cdb539c854833966bfae8ea552530fbf1405f1fe references,issue,94779,issue,29137,medium,issue.body,"() (like in x.mean(axis=())) return a scalar for torch but is a no-op for np, jnp, tf (besides returning float for int array, see bellow). (#29137) torch.ones(shape=()) fail (expect size=), but xnp.ones(shape=()) works (in other frameworks). Same for torch.zeros,... Casting: x...",https://github.com/pytorch/pytorch/issues/94779,93162953f4accb51c4c54aac0b07ba54564c025677d9c55b6f50ce1635e40af2 references,issue,94779,issue,40568,medium,issue.body,"d/numpy.concatenate.html Behavior: torch should accept np.dtype everywhere torch.dtype is valid (e.g. torch.ones((), dtype=np.int32)) (See: #40568). torch.dtype should be comparable with np.dtype: tf.int32 == np.int32 but torch.int32 != np.int32. This allow to have agnostic co...",https://github.com/pytorch/pytorch/issues/94779,c286b214a5d413772adf5c00b29f99258098380677124333d6c35d3e20f9b365 references,issue,94779,issue,46829,medium,issue.body,"f (this is very convenient in tests np.allclose(x, [1, 2, 3])) Other differences (but not critical to fix): Mixing torch and np.array fail (#46829). Both TF and Jax support tf.Tensor + np.array. Mixing float and double Would be nice if torch.asarray was supporting jax, tf tens...",https://github.com/pytorch/pytorch/issues/94779,0ef39c6165a45b832a443ffa915a62ce0aa5da95722946cb2eac1160ac508b3c references,issue,94779,issue,50344,medium,issue.body,"matched more closely numpy API (like tf.experimental.numpy and jax.numpy). This is a highly requested features, like: #2228 (~100 upvotes), #50344, #38349,... Those issues have since been closed even though there's still many issues remaining. This is even more relevant with t...",https://github.com/pytorch/pytorch/issues/94779,4e589be9a4ea94d93067376e669b3c53d2e0b82d93901526e1e110770f5a74c3 references,issue,94779,issue,64359,medium,issue.body,"rch.tensor) torch.ndarray: like isinstance(x, xnp.ndarray) (alias of torch.Tensor) Tensor.astype: like x = x.astype(np.uint8) torch.append (#64359): https://numpy.org/doc/stable/reference/generated/numpy.append.html torch.expand_dims: https://numpy.org/doc/stable/reference/gen...",https://github.com/pytorch/pytorch/issues/94779,eb24cd2e750c1784d4a0350b176e489387b76b32f556e9028d37a135c5db82dc references,issue,94779,issue,40568,medium,issue.comments[0].body,Another NumPy compat issue on supporting string dtypes: #40568 (comment) (useful for porting existing numpy code) And existing issue for append: #64359,https://github.com/pytorch/pytorch/issues/94779,ad4c9cbe4ef5349deb6d8d027d053b98dfa1658dc95dd5bdaabff906e15c6622 references,issue,94779,issue,64359,medium,issue.comments[0].body,er NumPy compat issue on supporting string dtypes: #40568 (comment) (useful for porting existing numpy code) And existing issue for append: #64359,https://github.com/pytorch/pytorch/issues/94779,37cd4931bd6f8380316a57b2b616f560c82061b9ef8136ed451553f5c8317796 references,issue,31461,issue,96036,medium,issue.comments[1].body,"Hi @rohan-varma @mrshenli , it's a long way finding this issue. I found a similar problem and I posted a issue in here #96036. I called the rpc_async and get a future. I found that when I called the future.wait(), it will still hold the GIL so the other thread cann",https://github.com/pytorch/pytorch/issues/31461,7efa1536abab8f9bccfad1a8d8bdf42225e543d13d32bedac3098f96b227ac5e references,issue,96412,issue,91965,medium,issue.comments[0].body,A bit related: #91965,https://github.com/pytorch/pytorch/issues/96412,a9edc53be112639b0d5214d35338ee45f7c170e6b34bdae30638ee0040bc2d20 references,issue,96276,issue,96275,medium,issue.body,"e output is: Segmentation fault (core dumped) The large value and zero in input's shape might be the cause of this bug, which is similar to #96275. Versions Collecting environment information... PyTorch version: 2.1.0.dev20230307+cu117 Is debug build: False CUDA used to build...",https://github.com/pytorch/pytorch/issues/96276,e8a0f2557e4c6b7ca79ffc048cc1d1960c18e2dcb914f457e20ae9c9be3519c9 references,issue,58742,issue,63293,medium,issue.body,"0) to_device astype unique_(all|counts|inverse|values) #70920 linalg.diagonal #70599 linalg.matrix_transpose, matrix_transpose linalg.outer #63293 linalg.tensordot #63478 linalg.trace #62714 linalg.vecdot #70542 cc @mruberry @rgommers @pmeier",https://github.com/pytorch/pytorch/issues/58742,9c809c684c04ede28bb32b2318a3d4b9567e8f26d187137d4a56d80bb05844c3 references,issue,58742,issue,70920,medium,issue.body,"at is compliant and alias concat to it) (#61767) expand_dims (#57116) __ipow__ (#76900) to_device astype unique_(all|counts|inverse|values) #70920 linalg.diagonal #70599 linalg.matrix_transpose, matrix_transpose linalg.outer #63293 linalg.tensordot #63478 linalg.trace #62714 l...",https://github.com/pytorch/pytorch/issues/58742,daaa6c2409feb4387c5e705db4a0a1122e4b2dd829d1bed7104296a16e7112ac references,issue,96136,issue,71683,medium,issue.comments[0].body,batchnorm running stats updates were explicit with calling torch._foreach_lerp_? I pseudo-coded an example of how it could look like here: #71683 (comment),https://github.com/pytorch/pytorch/issues/96136,ea68ac307bd222c66bafc5d80ef1538cf9fd767dc5a7a08759de3e388816faaa references,issue,95944,issue,88,medium,issue.body,"./Python/ceval.c:4327 #87 0x000014ce217d331c in _PyFunction_Vectorcall (func=, stack=0x8e3c530, nargsf=, kwnames=) at ../Objects/call.c:396 #88 0x000014ce21840e10 in _PyObject_VectorcallTstate (kwnames=0x0, nargsf=, args=0x8e3c530, callable=0x14cd97e61e50, tstate=0xe393f0) at...",https://github.com/pytorch/pytorch/issues/95944,8fc64cf1825bb1077bdcd3a462dd7cf65b29fc558df60f7ea3a0024cecb9432f references,issue,95434,issue,76785,medium,issue.body,"f these functions or if the checks for datatype consistency need to be tightened up in the forward process. This issue should be related to #76785. To Reproduce import torch input0 = torch.rand([3, 2], dtype=torch.complex128, requires_grad=True) input1 = torch.rand([3], dtype=...",https://github.com/pytorch/pytorch/issues/95434,dbf3ae4189a98bc794a8ddcafa594b82bf625e43f088df4ca8b0bd639b5caa31 references,issue,95590,issue,77415,medium,issue.body,"🐛 Describe the bug This issue should be related to #77415. In different versions of PyTorch, LazyLinear has different error messages for the inconsistency between the model dtype and the input dtyp",https://github.com/pytorch/pytorch/issues/95590,0c034beacbbcd1a75b95375fd61b6ad52bdbea43ae9093c75c1effa99361dc77 references,issue,95124,issue,86791,medium,issue.body,"timize_for_mobile in the packaged docker environment released by pytorch, a Vulcan support issue occurs, and then crashes. It is similar to #86791. It seems that the currently released docker environments may have some flags not enabled during building, resulting in some funct...",https://github.com/pytorch/pytorch/issues/95124,74f3661cf10347bdfcd999295c9d0d2d0249f4e878c92e311e5b23d88c845b7e references,issue,32293,issue,2129,medium,issue.body,"A truncated normal distribution is useful as initializer of weights or when sampling from ReLU potentials. This has been requested before (#2129, https://discuss.pytorch.org/t/implementing-truncated-normal-initializer/4778/15, etc.). Many packages have this: scipy, matlab, R,...",https://github.com/pytorch/pytorch/issues/32293,7953c65796944315c27b3cafb821f9aea7ccbf413d4241f9b4ff25597199df9c references,issue,32293,issue,31945,medium,issue.body,". Many packages have this: scipy, matlab, R, julia, tensorflow, and others. cc @vincentqb @fritzo @neerajprad @alicanb @vishwakftw Related: #31945",https://github.com/pytorch/pytorch/issues/32293,ac9d6976678d0a6afc090a2a32f5d8e334618a130303dae9fed2baa3079b1efe references,issue,50940,issue,50689,medium,issue.body,"d Simpletransformers. Replacing SequentialSampler with RandomSampler this issue disappear (but RandomSampler has other problems as in issue #50689 ) To Reproduce Steps to reproduce the behavior: Get a huge dataset, e.g. 100M rows Start training Track time over steps Expected b...",https://github.com/pytorch/pytorch/issues/50940,e7d378cad2d64db85a7ac4aa62c42d61c80fbe2960fbd4c95714294ca6d7ea09 references,issue,94808,issue,94668,medium,issue.comments[0].body,This issue might be similar to issue #94668. But the execution result of #94668 is always Segment Fault. This issue ends up with Heap Error sometimes. I am not sure if they have the s,https://github.com/pytorch/pytorch/issues/94808,c989770a1e0d97196ed2f87fa1a934f8e2311e51d86c908af913e0c44c1e69f3 references,issue,68041,issue,58833,medium,issue.body,"nvtx ranges: optimizer duration (CPU) duration (CUDA) FusedLAMB 17.444 ms 17.242 ms FusedMixedPrecisionLamb 8.668 ms 15.587 ms rel: #38655, #58833 cc @VitalyFedyunin @ngimel @vincentqb @jbschlosser @albanD @crcrpar @mcarilli @ptrblck",https://github.com/pytorch/pytorch/issues/68041,be0d8e8b5cd352d25ba415a67024068aadec2f82839998f6b2fba3af4e313c99 references,issue,85157,issue,83775,medium,issue.body,"Summary This is used as a top level doc to track all 1.13 issues for nestedtensors Issues: Complete #83775 (movetorch.nested_tensor) Improve documentation in torch.nested, document op coverage #85186 Improve perf of NT matmul #85064 #85311 Improv",https://github.com/pytorch/pytorch/issues/85157,d3560cf51aceb09b2887afa54ebf1973ddcc45a67859821ba23cd7c8dac5bd4a references,issue,85157,issue,91471,medium,issue.body,n updated to have more reasonable baseline to compare against and utilizes _sdpa under the hood through nn.mha. #91508 Non Blocking issues: #91471 Push Do we want to expose from_padded_tensor or potentially torch._nested_from_padded_and_nested_example? cc @cpuhrsch @jbschlosse...,https://github.com/pytorch/pytorch/issues/85157,46b7fa2ff47fdfc967607821575ccbd3174d3ec1557e8a99499e7f6f5395544a references,issue,85157,issue,91508,medium,issue.body,estedTensor Tutorial has been updated to have more reasonable baseline to compare against and utilizes _sdpa under the hood through nn.mha. #91508 Non Blocking issues: #91471 Push Do we want to expose from_padded_tensor or potentially torch._nested_from_padded_and_nested_examp...,https://github.com/pytorch/pytorch/issues/85157,68beab5e23186029dee12f79af5937b49c8acaa825d9d8917dd5e18867992a72 references,issue,85157,issue,27522,medium,issue.comments[0].body,@cpuhrsch some more nested tensor discussions (also of tensorlist): #65156 #27522,https://github.com/pytorch/pytorch/issues/85157,cc94b09c97243efdc12979481e414482457ae9e490fa891d036afb06203bd09f references,issue,85157,issue,65156,medium,issue.comments[0].body,@cpuhrsch some more nested tensor discussions (also of tensorlist): #65156 #27522,https://github.com/pytorch/pytorch/issues/85157,30add292fc54c92dccd0e2167816abad6f655193fc2ccde66616904f5c1a3355 references,issue,62148,issue,30702,medium,issue.comments[1].body,"At the very least, a function to add unitary dimensions on the right would be nice unsqueeze_multiple: #30702 (comment) (e.g. lambda x, y: y.unsqueeze(-1, repeat = x.dim() - y.dim())",https://github.com/pytorch/pytorch/issues/62148,967d68141c683f978af45c728a8a2fbe4c53b48e2eadcd4acd72f5131dfa4c61 references,issue,90547,issue,81692,medium,issue.body,"lly, now I am fighting with CUDA and CUDNN. While with CUDNN the issue is obvious (libcudnn_static.a got split into several other archives, #81692), I still struggle with CUDA itself. Reading CMake logic in the repository, one can see that PyTorch will prefer static libs over...",https://github.com/pytorch/pytorch/issues/90547,ebd50a4ea1d9e54668b7526474cfa0144a6876acf17fc6ae8dc78692c4a9eaae references,issue,93955,issue,34058,medium,issue.body,"🚀 The feature, motivation and pitch The PyTorch binaries are huge. See #34058. So huge in fact, we've needed to refactor our codebase just so they fit in pip and conda. And that we plan on setting up binary size alert",https://github.com/pytorch/pytorch/issues/93955,30476e04239b2f16f9bd6ec2e625d9486b7a73540bc4d990cd12f38e0730002a references,issue,93955,issue,93081,medium,issue.body,it in pip and conda. And that we plan on setting up binary size alerts: #93991. And that we want to split up our conda for faster installs: #93081 . We should do everything we can to do to keep them small without sacrificing performance. One now commonly supported compiler fea...,https://github.com/pytorch/pytorch/issues/93955,76110b747bca4bd925efcbcc9c0b3d24083c3c1ad04333fda1acb462bb45b262 references,issue,93955,issue,90608,medium,issue.comments[0].body,"LTO reorders object files and can increase the likelihood for static initialization order bugs to appear. I suspect this is why I ran into #90608, #90149 and #90133 while other users didn't. So perhaps #90608 should be fixed first before enabling LTO (and, watch out for simila...",https://github.com/pytorch/pytorch/issues/93955,3f38e53d59e39444290425a56edb30c7d80b52d67eaa8176f9969d3ee81af5aa references,issue,92884,issue,23756,medium,issue.comments[0].body,Related: #26288 #23756,https://github.com/pytorch/pytorch/issues/92884,9a291555020980ce1ac50ab6bb5e7ab3067dfd61b2e76d24edb5440ddff87b6b references,issue,92884,issue,26288,medium,issue.comments[0].body,Related: #26288 #23756,https://github.com/pytorch/pytorch/issues/92884,d23deeb52566d6122ff6ba4c65ef3d9320f8355d131399c0215caba40b40310b references,issue,84524,issue,84290,medium,issue.body,"🐛 Describe the bug Originally discussed in #84290 (comment), more links to related comments / issues in this cited comment The F.softmax / softmin / log_softmax have implicit dim support de",https://github.com/pytorch/pytorch/issues/84524,b109111992865b4c6dce4b4b1fccf4ec72cf922e5dcf50b27b66b18977ebaf7f references,issue,87366,issue,53441,medium,issue.comments[0].body,"As mentioned elsewhere (e.g. #53441), there is a pytorch-based conjugate gradient algorithm available as part of cornellius-gp/linear_operator: https://github.com/cornellius-g",https://github.com/pytorch/pytorch/issues/87366,c06e8966370f48403a70c88c4582a415f25e98559f8c1e86d8e01837659b139d references,issue,86590,issue,66073,medium,issue.comments[0].body,Related: #66073 #68648,https://github.com/pytorch/pytorch/issues/86590,197c06662da8cbdf7b698ddda730019d114db8dffebacb2660e3d95a4332cb4d references,issue,86590,issue,68648,medium,issue.comments[0].body,Related: #66073 #68648,https://github.com/pytorch/pytorch/issues/86590,09e5a83b3e5eed8a1a8e83eba519efb9f224706b2cd111f434a0c76959d53264 references,issue,53937,issue,53678,medium,issue.comments[0].body,"This has the same root cause as #53678, except that the issue in that case was previously hidden, but now it's not. @bhosmer",https://github.com/pytorch/pytorch/issues/53937,2a01bb4851beb9fd88d52d8c0b8dabdb1305c91e4c0891f9cac8a633ecce2e98 references,issue,91293,issue,79715,medium,issue.body,behavior should be that the tracing for both of the following modules should complete without emitting any error messages. Related issues: #79715 #34294 (Seems to contain an (incomplete) fix?) #43248 Code for reproducing the issue: The graph = tracer.trace(nested_cat) line com...,https://github.com/pytorch/pytorch/issues/91293,48ec2414770ddcf7a6ec0fcd6e34f7ed2ac3a35e6d3d9f396f7710b39b78a74c references,issue,71210,issue,18095,medium,issue.comments[0].body,"Related: #68616 , #18095",https://github.com/pytorch/pytorch/issues/71210,a0ceaec943a115f63c46bfbf0d5c22f1a7cef1742b1f5a9c2e00604113f383d1 references,issue,71210,issue,68616,medium,issue.comments[0].body,"Related: #68616 , #18095",https://github.com/pytorch/pytorch/issues/71210,ccb50af8a516c8ae559adf27c79da4e0a5d0dbed7dd327ef5b55f39c6c673603 references,issue,71210,issue,61585,medium,issue.comments[1].body,We should solve this while/by fixing #61585.,https://github.com/pytorch/pytorch/issues/71210,345ea0ae87a080cf477d43489b77c128c10cc6b6ab521d59c5ff00c1c9348f18 references,issue,92801,issue,69431,medium,issue.comments[0].body,Maybe related: #69431 #88540,https://github.com/pytorch/pytorch/issues/92801,6110e998d3a7712c3dfd2d6a4f2899f506abf572ce1edddf3a10ae99c7ba16d4 references,issue,70245,issue,88,medium,issue.body,"names=0x0, nargsf=, args=0x7ffff778f958, callable=0x7ffff78da1f0, tstate=0x555555928d10) at ./Include/cpython/abstract.h:118 #88 PyObject_Vectorcall (kwnames=0x0, nargsf=, args=0x7ffff778f958, callable=) at ./Include/cpython/abstrac...",https://github.com/pytorch/pytorch/issues/70245,c12ee91515dc4413d5bc6530b060cb686ad73e5750fbf16130f22114fafe22fd references,issue,89160,issue,88838,medium,issue.body,"t32 stacktraces (4,000 lines): https://gist.github.com/xwang233/9feb59c637418aa4ce84d917960afbde Versions f20b3f2 V100x8 cuda 11.8 Related: #88838 cc @ngimel @wanchaol",https://github.com/pytorch/pytorch/issues/89160,da4f0fff727c9de16836258d92ed8eb67a8e1c57a44068c979c771d5d8c00bca references,issue,64535,issue,26165,medium,issue.body,th MKL_NUM_THREADS=1 python leak.py or OMP_NUM_THREADS=1 python leak.py there will be no memory leak can be observed. Related Issues #64412 #26165 #24237 might be the related problem. cc @VitalyFedyunin @gujinghui @PenghuiCheng @XiaobingSuper @jianyuh,https://github.com/pytorch/pytorch/issues/64535,7199af4a5f35b96c3e42e4f6f59667f53a93e18d325c5d2987f8f5b0031226ab references,issue,64535,issue,64412,medium,issue.body,ript with MKL_NUM_THREADS=1 python leak.py or OMP_NUM_THREADS=1 python leak.py there will be no memory leak can be observed. Related Issues #64412 #26165 #24237 might be the related problem. cc @VitalyFedyunin @gujinghui @PenghuiCheng @XiaobingSuper @jianyuh,https://github.com/pytorch/pytorch/issues/64535,41abb5645e91281125c0cb76c1838599afcaf3f49dae513b94644f5c4460be87 references,issue,58736,issue,58489,medium,issue.body,"2 torch.complex32 + torch.complex128 = torch.complex32 torch.complex64 + torch.complex128 = torch.complex64 This is not documented well(see #58489), but seems to be intentional. cc @nairbv @mruberry",https://github.com/pytorch/pytorch/issues/58736,addfac03e40938d2a700ac020276d4c4390398acd924f5aee1e39213be296a80 references,issue,58414,issue,58858,medium,issue.body,1. Here is a list of features to implement based on #51390: cuDNN benchmark (#58859) convolution backward / transposed convolution forward (#58858) conv-bias-activation fusion (#58860) better error message with debugging information and python repro when failing (on par with v...,https://github.com/pytorch/pytorch/issues/58414,e62d2080f5d21cbb001f36d16fea0b3d1470dba8e41819a39aec8f9db64e3648 references,issue,58414,issue,58859,medium,issue.body,CI pipeline to test PyTorch with USE_EXPERIMENTAL_CUDNN_V8_API=1. Here is a list of features to implement based on #51390: cuDNN benchmark (#58859) convolution backward / transposed convolution forward (#58858) conv-bias-activation fusion (#58860) better error message with deb...,https://github.com/pytorch/pytorch/issues/58414,637b129d099eaaf1357abceafc7932ba3f506f715e75f3b9484c987e58b521fd references,issue,58414,issue,58860,medium,issue.body,"ement based on #51390: cuDNN benchmark (#58859) convolution backward / transposed convolution forward (#58858) conv-bias-activation fusion (#58860) better error message with debugging information and python repro when failing (on par with v7, reference PR #45023) (#58862) BFlo...",https://github.com/pytorch/pytorch/issues/58414,d460635c3582d8fc2283e26f3933ea48d12fc87c04f978a42d0bb77552f3001c references,issue,58414,issue,58862,medium,issue.body,"vation fusion (#58860) better error message with debugging information and python repro when failing (on par with v7, reference PR #45023) (#58862) BFloat16 support (#58861) Here is a list of issues to resolve based on #51390: The heuristic/benchmark cache of engine config is...",https://github.com/pytorch/pytorch/issues/58414,358b4da9396ba76d963e99461ddc2b3af9ea0d4ed0d0f13bd74eb0c3b1b916ff references,issue,78489,issue,78475,medium,issue.body,"🐛 Describe the bug In opposite, torch.backends.cudnn.version() does collect it correctly. Originally reported by @grazder in a comment: #78475 (comment) This occurred on old-ish version of 1.9.1, not sure if it still happens on 1.11.0 Versions 1.9.1 cc @csarofeen @ptrblck @xwa...",https://github.com/pytorch/pytorch/issues/78489,caec372ff34ad606f1b0fab256506b9ca8830af5b6ca1135d18dfaa8bcaa44db references,issue,91887,issue,24234,medium,issue.comments[0].body,"Also, there are issues with tensorboard not being fast enough to flush events: #24234, hope it's not affecting your case",https://github.com/pytorch/pytorch/issues/91887,37258d4547a28c4830a672ff686f549b9c1f3f0303de105a628414fc6ab7f62e references,issue,91300,issue,19248,medium,issue.comments[1].body,Duplicate of #19248,https://github.com/pytorch/pytorch/issues/91300,2237a4df651262d624fe6509fb0ed5ce5c8264b8cf07c857214571c6ea7b7cb2 references,issue,64492,issue,24422,medium,issue.comments[0].body,"cc @robieta @chaekit as well as the above. Also just added both of you to the oncall at #24422. Pickling is always a challenge. Guys, any thoughts? Thanks.",https://github.com/pytorch/pytorch/issues/64492,4d417706809f339a73c5938999208d1789cdf2ee96f335be526d2901532338ca competes with,issue,73175,issue,58181,medium,issue.body,"tuple of ints, implementing split_with_sizes (I propose torch.split_with_sizes to be then made private / deprecated instead of documenting: #58181) Sometimes it's convenient to pass in sizes as a tensor instead of tuple. Tensor could also be a CUDA tensor then without forcing...",https://github.com/pytorch/pytorch/issues/73175,42dcfeec283770ab0ae082b29cde7b965ed9f7d04e40657382356aaa90a1fc22 references,issue,71329,issue,20117,medium,issue.body,"🚀 The feature, motivation and pitch This thread is a revisiting of #20117, where the possibility of adding various array operations to nn.Sequential were discussed. The rationale behind opening a new issue is as f",https://github.com/pytorch/pytorch/issues/71329,f846fbb7eb55881849574f17fd378b1ef9add3cbabc23ad7b4edf7bd0256f6e6 references,issue,89241,issue,44710,medium,issue.body,"modules exported via torch.jit.trace() with amp.autocast. There is still outstanding issue about lack of true autocast support in C++ here: #44710. For our purposes(inference), we would actually prefer a different solution: I think after jit.freeze() and jit.optimize_for_infer...",https://github.com/pytorch/pytorch/issues/89241,841303a6ddf868c339f8e03773a73bf931ac73007896f4c2302a3d17a11a596d references,issue,89241,issue,72295,medium,issue.comments[0].body,". Maybe there is no issue with pure trace - but LSTM does not really work w/o script. Also, here's previously submitted proposal re freeze: #72295",https://github.com/pytorch/pytorch/issues/89241,76ad414a8829904282275e5bf27fa4caae1035be685bac54d65eb2aa12a83163 references,issue,91374,issue,91156,medium,issue.body,"ions) a.quantile(0.25) # or torch.quantile(a, 0.25) Alternatives No response Additional context I was asked to resubmit this as an RFC. See #91156. cc @ngimel",https://github.com/pytorch/pytorch/issues/91374,09f605796b336721f06486811df390e0fd6f36935b2e9d9788434328fb207a80 references,issue,90663,issue,84625,medium,issue.body,"orch and to use it to increase precision in the sampling of Dirichlet/Beta distributions. This feature proposal is in part related to issue #84625. This issue reveals that for small parameter values, the beta distribution behaves incorrectly. More generally, the Dirichlet dist...",https://github.com/pytorch/pytorch/issues/90663,110ab33294413319bfa4b7d3703953e2273994587957e69c66cb27b217fe207f references,issue,33354,issue,27610,medium,issue.body,"hread) or have a flag to run uncompiled? Improve compile times for particular outliers like densenet161 cc @suo , @fbbradheintz Related to: #27610 Example of slow initial query attached, with output: $ python test_torchscript.py loading model model loaded warmup ['n02123045',...",https://github.com/pytorch/pytorch/issues/33354,f948f6441b37b9c4a7e471fae49dd8305b959de26294fe3e1e8c671796340c87 references,issue,87448,issue,9674,medium,issue.body,"ent"" can be achieved in two steps: 1) dense gradient; 2) subsample Additional context Related issues: #87358 #41128 pytorch/maskedtensor#71 #9674 #86963 cc @nikitaved @pearu @cpuhrsch @amjames @bhosmer @svekars @carljparker @ezyang @albanD @zou3519 @gqchen @soulitzer @lezcano...",https://github.com/pytorch/pytorch/issues/87448,aa07254d2d1f024e097e8a604abf089a02c26815dc119f507afaab68a5d1ca38 references,issue,87448,issue,41128,medium,issue.body,"n PyTorch, ""sparsity-preserving gradient"" semantics is already used in some cases, in particular in torch.sparse.mm, see e.g. discussion in #41128 and #2389 (comment) but also usage in https://github.com/rusty1s/pytorch_sparse and https://github.com/flaport/torch_sparse_solve....",https://github.com/pytorch/pytorch/issues/87448,473f1b0132099f24ac27cea08ba4a80aa601732fe7114446719b832f8db759ee references,issue,87448,issue,87358,medium,issue.body,"se matrix are not made explicit in the documentation and may thus not match user expectations. After discussion in several issues (see e.g. #87358 (comment)), I realised that the default (implicit) semantics of sparse tensors in PyTorch is that the sparse representation only a...",https://github.com/pytorch/pytorch/issues/87448,acd6b52326904f30cdf2483f83e2981079cba48ea06fcb5eb9d3d62a6e5dde84 references,issue,90367,issue,90366,medium,issue.body,"orch.add(torch.zeros((), dtype=torch.bool), True) # jit trace traced = torch.jit.trace(fn, (torch.zeros((), dtype=torch.bool),)) Similar to #90366 but for torch.add. Versions Env [click to expand] """""" Collecting environment information... PyTorch version: 1.14.0.dev20221202+cu...",https://github.com/pytorch/pytorch/issues/90367,fe50d2df22827a4e566fb6c27c222b0c01858e7db7a7f1c35029112db2ee8772 references,issue,71806,issue,62540,medium,issue.comments[0].body,#219 Update our lambdas that in live AWS to support main Once the change is approved. Change the default branch following the guidelines on #62540 Remove support for running pipelines on master. Change viable/strict pipeline to take from main and not master Confirm all pipelin...,https://github.com/pytorch/pytorch/issues/71806,9fa44880688337ad8bc29542f784e929f5e1a6c8794e43fa77c7cb824b95bd11 references,issue,62540,issue,71806,medium,issue.body,lt Branch setting where you can rename the default branch! Rename master to main and merge your PR! Repos: According to get_public_repos.py #71806 #71844 #71843 #71842 #71841 #71840 #71839 #71838 #71837 #71836 #71835 #71834 #71833 #71832 #71831 #71830 #71829 #71828 #71827 #718...,https://github.com/pytorch/pytorch/issues/62540,d045d0940e7c6274e220416420301b5c65c7f645bb625b81581c21501180eb68 references,issue,62540,issue,71810,medium,issue.body,#71830 #71829 #71828 #71827 #71826 #71825 #71824 #71823 #71822 #71821 #71820 #71819 #71818 #71817 #71816 #71815 #71814 #71813 #71812 #71811 #71810 DONE: https://github.com/pytorch/ci-hud https://github.com/pytorch/test-infra https://github.com/pytorch/kineto https://github.com...,https://github.com/pytorch/pytorch/issues/62540,b0072080869991d290ab5c7f728046ba226805f4b7997d5f78cfefec9c2f4b64 references,issue,62540,issue,71812,medium,issue.body,#71832 #71831 #71830 #71829 #71828 #71827 #71826 #71825 #71824 #71823 #71822 #71821 #71820 #71819 #71818 #71817 #71816 #71815 #71814 #71813 #71812 #71811 #71810 DONE: https://github.com/pytorch/ci-hud https://github.com/pytorch/test-infra https://github.com/pytorch/kineto http...,https://github.com/pytorch/pytorch/issues/62540,c315f792c40a9af7439d90272a4e7492d3cf75dd065ff61fa465e79c8f7b108a references,issue,62540,issue,71813,medium,issue.body,#71833 #71832 #71831 #71830 #71829 #71828 #71827 #71826 #71825 #71824 #71823 #71822 #71821 #71820 #71819 #71818 #71817 #71816 #71815 #71814 #71813 #71812 #71811 #71810 DONE: https://github.com/pytorch/ci-hud https://github.com/pytorch/test-infra https://github.com/pytorch/kine...,https://github.com/pytorch/pytorch/issues/62540,1597451dfa3f77faa3fe3605e82c4e7fbf80a096cc7ed31a44a52e0300eb1ea0 references,issue,62540,issue,71814,medium,issue.body,#71834 #71833 #71832 #71831 #71830 #71829 #71828 #71827 #71826 #71825 #71824 #71823 #71822 #71821 #71820 #71819 #71818 #71817 #71816 #71815 #71814 #71813 #71812 #71811 #71810 DONE: https://github.com/pytorch/ci-hud https://github.com/pytorch/test-infra https://github.com/pytor...,https://github.com/pytorch/pytorch/issues/62540,730805fc574f23edfe27ac8897fc194f32820b634643ff9435504cf4bdd53521 references,issue,62540,issue,71815,medium,issue.body,#71835 #71834 #71833 #71832 #71831 #71830 #71829 #71828 #71827 #71826 #71825 #71824 #71823 #71822 #71821 #71820 #71819 #71818 #71817 #71816 #71815 #71814 #71813 #71812 #71811 #71810 DONE: https://github.com/pytorch/ci-hud https://github.com/pytorch/test-infra https://github.co...,https://github.com/pytorch/pytorch/issues/62540,e6d27bafbe80ccb8e34c9d141aba609279b938f8ef6896ed13de724ebe5f301e references,issue,62540,issue,71816,medium,issue.body,#71836 #71835 #71834 #71833 #71832 #71831 #71830 #71829 #71828 #71827 #71826 #71825 #71824 #71823 #71822 #71821 #71820 #71819 #71818 #71817 #71816 #71815 #71814 #71813 #71812 #71811 #71810 DONE: https://github.com/pytorch/ci-hud https://github.com/pytorch/test-infra https://gi...,https://github.com/pytorch/pytorch/issues/62540,b730da374c0e82b550c625dff181ba585c08f58cfb10311090beb598da04821a references,issue,62540,issue,71817,medium,issue.body,#71837 #71836 #71835 #71834 #71833 #71832 #71831 #71830 #71829 #71828 #71827 #71826 #71825 #71824 #71823 #71822 #71821 #71820 #71819 #71818 #71817 #71816 #71815 #71814 #71813 #71812 #71811 #71810 DONE: https://github.com/pytorch/ci-hud https://github.com/pytorch/test-infra htt...,https://github.com/pytorch/pytorch/issues/62540,9d17224c5f8c3d1818914c6c3230d6d2960e8f40a1b82f71cc288e509a7ca7af references,issue,62540,issue,71818,medium,issue.body,#71838 #71837 #71836 #71835 #71834 #71833 #71832 #71831 #71830 #71829 #71828 #71827 #71826 #71825 #71824 #71823 #71822 #71821 #71820 #71819 #71818 #71817 #71816 #71815 #71814 #71813 #71812 #71811 #71810 DONE: https://github.com/pytorch/ci-hud https://github.com/pytorch/test-in...,https://github.com/pytorch/pytorch/issues/62540,e62508e189db0ce9ebbbb6e32e1368c91b31d3d6c801ebbd2f6453cfbf0b8df8 references,issue,62540,issue,71819,medium,issue.body,#71839 #71838 #71837 #71836 #71835 #71834 #71833 #71832 #71831 #71830 #71829 #71828 #71827 #71826 #71825 #71824 #71823 #71822 #71821 #71820 #71819 #71818 #71817 #71816 #71815 #71814 #71813 #71812 #71811 #71810 DONE: https://github.com/pytorch/ci-hud https://github.com/pytorch/...,https://github.com/pytorch/pytorch/issues/62540,caa1549350b439389f65ac3755538587f7ce8b0e7303ae117abdaccabdc06305 references,issue,62540,issue,71820,medium,issue.body,#71840 #71839 #71838 #71837 #71836 #71835 #71834 #71833 #71832 #71831 #71830 #71829 #71828 #71827 #71826 #71825 #71824 #71823 #71822 #71821 #71820 #71819 #71818 #71817 #71816 #71815 #71814 #71813 #71812 #71811 #71810 DONE: https://github.com/pytorch/ci-hud https://github.com/p...,https://github.com/pytorch/pytorch/issues/62540,6d52928b0d8c76350dbdfeb019cbc95181a6e9d17899751d1a5e0646deb194c3 references,issue,62540,issue,71821,medium,issue.body,#71841 #71840 #71839 #71838 #71837 #71836 #71835 #71834 #71833 #71832 #71831 #71830 #71829 #71828 #71827 #71826 #71825 #71824 #71823 #71822 #71821 #71820 #71819 #71818 #71817 #71816 #71815 #71814 #71813 #71812 #71811 #71810 DONE: https://github.com/pytorch/ci-hud https://githu...,https://github.com/pytorch/pytorch/issues/62540,abfc0f762f6d19784f4a495927ac3c65e14e39b35b19991a99eb85b9b3abff7d references,issue,62540,issue,71822,medium,issue.body,#71842 #71841 #71840 #71839 #71838 #71837 #71836 #71835 #71834 #71833 #71832 #71831 #71830 #71829 #71828 #71827 #71826 #71825 #71824 #71823 #71822 #71821 #71820 #71819 #71818 #71817 #71816 #71815 #71814 #71813 #71812 #71811 #71810 DONE: https://github.com/pytorch/ci-hud https:...,https://github.com/pytorch/pytorch/issues/62540,e1eda7b2c8b926abc4a8578f6109a2b6b4c37a7e8b2e26b092eb94cdaf9dc3f2 references,issue,62540,issue,71823,medium,issue.body,#71843 #71842 #71841 #71840 #71839 #71838 #71837 #71836 #71835 #71834 #71833 #71832 #71831 #71830 #71829 #71828 #71827 #71826 #71825 #71824 #71823 #71822 #71821 #71820 #71819 #71818 #71817 #71816 #71815 #71814 #71813 #71812 #71811 #71810 DONE: https://github.com/pytorch/ci-hud...,https://github.com/pytorch/pytorch/issues/62540,7dc2ca3cb5c52e17676a087fdfe9fc02d188d4932e1a49ed55e85a1fde5a3061 references,issue,62540,issue,71824,medium,issue.body,#71844 #71843 #71842 #71841 #71840 #71839 #71838 #71837 #71836 #71835 #71834 #71833 #71832 #71831 #71830 #71829 #71828 #71827 #71826 #71825 #71824 #71823 #71822 #71821 #71820 #71819 #71818 #71817 #71816 #71815 #71814 #71813 #71812 #71811 #71810 DONE: https://github.com/pytorch...,https://github.com/pytorch/pytorch/issues/62540,f0a0883f32c6de44c91f9f34e745f8350c3401931ce58f473ab2f71c56f9319e references,issue,62540,issue,71825,medium,issue.body,#71806 #71844 #71843 #71842 #71841 #71840 #71839 #71838 #71837 #71836 #71835 #71834 #71833 #71832 #71831 #71830 #71829 #71828 #71827 #71826 #71825 #71824 #71823 #71822 #71821 #71820 #71819 #71818 #71817 #71816 #71815 #71814 #71813 #71812 #71811 #71810 DONE: https://github.com/...,https://github.com/pytorch/pytorch/issues/62540,4895898f106398c9e267b2647761b9ab7ff7199470483364e98934f02dd8ab7e references,issue,62540,issue,71826,medium,issue.body,pos.py #71806 #71844 #71843 #71842 #71841 #71840 #71839 #71838 #71837 #71836 #71835 #71834 #71833 #71832 #71831 #71830 #71829 #71828 #71827 #71826 #71825 #71824 #71823 #71822 #71821 #71820 #71819 #71818 #71817 #71816 #71815 #71814 #71813 #71812 #71811 #71810 DONE: https://gith...,https://github.com/pytorch/pytorch/issues/62540,b7880a1b732d75ad39f827a2a3016ff6298751b898e4d10e8523fb0a0d0af834 references,issue,62540,issue,71828,medium,issue.body,get_public_repos.py #71806 #71844 #71843 #71842 #71841 #71840 #71839 #71838 #71837 #71836 #71835 #71834 #71833 #71832 #71831 #71830 #71829 #71828 #71827 #71826 #71825 #71824 #71823 #71822 #71821 #71820 #71819 #71818 #71817 #71816 #71815 #71814 #71813 #71812 #71811 #71810 DONE:...,https://github.com/pytorch/pytorch/issues/62540,29cd8c22fa7a4e0486130e04223a5e70d32596218994a6021bcd6ae8c3657be3 references,issue,62540,issue,71830,medium,issue.body,: According to get_public_repos.py #71806 #71844 #71843 #71842 #71841 #71840 #71839 #71838 #71837 #71836 #71835 #71834 #71833 #71832 #71831 #71830 #71829 #71828 #71827 #71826 #71825 #71824 #71823 #71822 #71821 #71820 #71819 #71818 #71817 #71816 #71815 #71814 #71813 #71812 #718...,https://github.com/pytorch/pytorch/issues/62540,eace72d6fd852e3f5ad2afaf2d30e6b677429aa312c9d67899de153cc02259fa references,issue,62540,issue,71831,medium,issue.body,! Repos: According to get_public_repos.py #71806 #71844 #71843 #71842 #71841 #71840 #71839 #71838 #71837 #71836 #71835 #71834 #71833 #71832 #71831 #71830 #71829 #71828 #71827 #71826 #71825 #71824 #71823 #71822 #71821 #71820 #71819 #71818 #71817 #71816 #71815 #71814 #71813 #718...,https://github.com/pytorch/pytorch/issues/62540,7baacc38cf8dce56546977f430c52fa0f794dcec24bb0417ee9d76e42eb57bbb references,issue,62540,issue,71832,medium,issue.body,your PR! Repos: According to get_public_repos.py #71806 #71844 #71843 #71842 #71841 #71840 #71839 #71838 #71837 #71836 #71835 #71834 #71833 #71832 #71831 #71830 #71829 #71828 #71827 #71826 #71825 #71824 #71823 #71822 #71821 #71820 #71819 #71818 #71817 #71816 #71815 #71814 #718...,https://github.com/pytorch/pytorch/issues/62540,cae1f6a4b27c1f7514349619c9f758185b7d628e4606382985da5414229f4473 references,issue,62540,issue,71833,medium,issue.body,merge your PR! Repos: According to get_public_repos.py #71806 #71844 #71843 #71842 #71841 #71840 #71839 #71838 #71837 #71836 #71835 #71834 #71833 #71832 #71831 #71830 #71829 #71828 #71827 #71826 #71825 #71824 #71823 #71822 #71821 #71820 #71819 #71818 #71817 #71816 #71815 #7181...,https://github.com/pytorch/pytorch/issues/62540,e3d3960bcd458ae0f796296674deee522b70db94f4c92cb216799db3564d77fe references,issue,62540,issue,71834,medium,issue.body,ain and merge your PR! Repos: According to get_public_repos.py #71806 #71844 #71843 #71842 #71841 #71840 #71839 #71838 #71837 #71836 #71835 #71834 #71833 #71832 #71831 #71830 #71829 #71828 #71827 #71826 #71825 #71824 #71823 #71822 #71821 #71820 #71819 #71818 #71817 #71816 #718...,https://github.com/pytorch/pytorch/issues/62540,bafd58754c896646f587b20d865a46deccdb665f2981eba44b03ae646b0c50ff references,issue,62540,issue,71835,medium,issue.body,er to main and merge your PR! Repos: According to get_public_repos.py #71806 #71844 #71843 #71842 #71841 #71840 #71839 #71838 #71837 #71836 #71835 #71834 #71833 #71832 #71831 #71830 #71829 #71828 #71827 #71826 #71825 #71824 #71823 #71822 #71821 #71820 #71819 #71818 #71817 #718...,https://github.com/pytorch/pytorch/issues/62540,50efb03e21daf7baeb273b5731a273aeee40001f7f7b0c18e1833b0835a55909 references,issue,62540,issue,71836,medium,issue.body,me master to main and merge your PR! Repos: According to get_public_repos.py #71806 #71844 #71843 #71842 #71841 #71840 #71839 #71838 #71837 #71836 #71835 #71834 #71833 #71832 #71831 #71830 #71829 #71828 #71827 #71826 #71825 #71824 #71823 #71822 #71821 #71820 #71819 #71818 #718...,https://github.com/pytorch/pytorch/issues/62540,1bc898c80e0f9949499472eac77b4d7a12e243e0bfc32c33dc01b0672e7167b3 references,issue,62540,issue,71837,medium,issue.body,h! Rename master to main and merge your PR! Repos: According to get_public_repos.py #71806 #71844 #71843 #71842 #71841 #71840 #71839 #71838 #71837 #71836 #71835 #71834 #71833 #71832 #71831 #71830 #71829 #71828 #71827 #71826 #71825 #71824 #71823 #71822 #71821 #71820 #71819 #718...,https://github.com/pytorch/pytorch/issues/62540,a35afb17488d69a0182a1b11ec8d861d881d6a0d4a7a00ad381dc1ecfdbccab9 references,issue,62540,issue,71838,medium,issue.body,t branch! Rename master to main and merge your PR! Repos: According to get_public_repos.py #71806 #71844 #71843 #71842 #71841 #71840 #71839 #71838 #71837 #71836 #71835 #71834 #71833 #71832 #71831 #71830 #71829 #71828 #71827 #71826 #71825 #71824 #71823 #71822 #71821 #71820 #718...,https://github.com/pytorch/pytorch/issues/62540,3731ca695603fa2772e6b5bfc0ecc144323242e74c71a2594b2b4c3d63106089 references,issue,62540,issue,71839,medium,issue.body,default branch! Rename master to main and merge your PR! Repos: According to get_public_repos.py #71806 #71844 #71843 #71842 #71841 #71840 #71839 #71838 #71837 #71836 #71835 #71834 #71833 #71832 #71831 #71830 #71829 #71828 #71827 #71826 #71825 #71824 #71823 #71822 #71821 #7182...,https://github.com/pytorch/pytorch/issues/62540,e9f6597d97e1da8484d33326a0de57db2960715d590675d2fdbad36e20fdae3b references,issue,62540,issue,71840,medium,issue.body,ame the default branch! Rename master to main and merge your PR! Repos: According to get_public_repos.py #71806 #71844 #71843 #71842 #71841 #71840 #71839 #71838 #71837 #71836 #71835 #71834 #71833 #71832 #71831 #71830 #71829 #71828 #71827 #71826 #71825 #71824 #71823 #71822 #718...,https://github.com/pytorch/pytorch/issues/62540,982bd25b37a843fcd36974b3a7910e97fc2d8f6221afbc5cdb40ecd1e330da8a references,issue,62540,issue,71841,medium,issue.body,can rename the default branch! Rename master to main and merge your PR! Repos: According to get_public_repos.py #71806 #71844 #71843 #71842 #71841 #71840 #71839 #71838 #71837 #71836 #71835 #71834 #71833 #71832 #71831 #71830 #71829 #71828 #71827 #71826 #71825 #71824 #71823 #718...,https://github.com/pytorch/pytorch/issues/62540,4105eb23ff1fd845ac7b9551839ddaf00077feef63e943d5a1a7e4d1aef3f759 references,issue,62540,issue,71842,medium,issue.body,re you can rename the default branch! Rename master to main and merge your PR! Repos: According to get_public_repos.py #71806 #71844 #71843 #71842 #71841 #71840 #71839 #71838 #71837 #71836 #71835 #71834 #71833 #71832 #71831 #71830 #71829 #71828 #71827 #71826 #71825 #71824 #718...,https://github.com/pytorch/pytorch/issues/62540,223fe4c3910f4b71aca886a74dd7c3599dc4ca3d0feeca334d8487e0704e79a2 references,issue,62540,issue,71843,medium,issue.body,ing where you can rename the default branch! Rename master to main and merge your PR! Repos: According to get_public_repos.py #71806 #71844 #71843 #71842 #71841 #71840 #71839 #71838 #71837 #71836 #71835 #71834 #71833 #71832 #71831 #71830 #71829 #71828 #71827 #71826 #71825 #718...,https://github.com/pytorch/pytorch/issues/62540,5394ba8329234f86b64f62759b59f8d99b483609f1d235986434a47dbef9d330 references,issue,72138,issue,55207,medium,issue.body,arallelism etc.). We have proposed a standardized sharding spec and a few building blocks for users to deal with sharding PyTorch models in #55207. This is a follow-up to the aforementioned proposal after we have looked into the several recent use cases such as Megatron-LM. Pi...,https://github.com/pytorch/pytorch/issues/72138,552680c7e799fcb97a705d8429597d4a368a59a0275670f7aadf34d4e43a25fb references,issue,89054,issue,64162,medium,issue.comments[0].body,Related: #64162 suggesting to accept seed argument to Generator constructor directly,https://github.com/pytorch/pytorch/issues/89054,d5d48ab8ba3d747b05584536235d9fbcc6958b264ef9e9bcc61270dfcd101f80 references,issue,73050,issue,65760,medium,issue.comments[1].body,"ah, the issue was that optional mutable arguments have some issues, so we can't mark the arguments as Tensor(a!)? in native_functions.yaml. #65760 (Our workaround was to add a special case in alias db #66554)",https://github.com/pytorch/pytorch/issues/73050,b903ee7c97a8ad96f8a71a1c6818d5996c4e89ab54426e478683475a06fbdff5 references,issue,80553,issue,1359,medium,issue.body,"be used for code tests. But at least if this feature request is approved, we don't have to start from scratch. This is directly related to #1359. Reference [1] W. W. Hager and H. Zhang, “Algorithm 851: CG_DESCENT, a conjugate gradient method with guaranteed descent,” ACM Trans...",https://github.com/pytorch/pytorch/issues/80553,76231b771c66889d67b3657dc3e98ac3d331ac13c7e55203759806f1d191dd44 references,issue,80553,issue,17902,medium,issue.body,"doi: 10.1145/1132973.1132979. Alternatives No response Additional context There are some issues mentioning implementing CG, like #53441 and #17902. However, I think they are more about linear CG, which is a special case of nonlinear CG. And I feel like linear and nonlinear CG...",https://github.com/pytorch/pytorch/issues/80553,8e2753c5d3ad7ef62864ef48103d7abcf16c8d923883e5ad5ff4df2f53d91238 references,issue,80553,issue,53441,medium,issue.body,"Mar. 2006, doi: 10.1145/1132973.1132979. Alternatives No response Additional context There are some issues mentioning implementing CG, like #53441 and #17902. However, I think they are more about linear CG, which is a special case of nonlinear CG. And I feel like linear and no...",https://github.com/pytorch/pytorch/issues/80553,9db2d703d9d344053d3173e02994e00fcac796c64cda5587a5d5432f1cb6ed62 references,issue,89601,issue,89105,medium,issue.body,"_WRAPPER=0 export USE_CUDA=0 export USE_ROCM=0 python3 setup.py bdist_wheel For 1.13 The first error encountered is already described here: #89105 If the layernorm.glsl file is updated to remove the quotation marks referred to in the linked issue, I eventually hit the followin...",https://github.com/pytorch/pytorch/issues/89601,7630427a3d9ef4430664c97736f61291b477918abc05608010ca9cb826774622 references,issue,89125,issue,89124,medium,issue.body,NestedTensorImpl doesn't support sizes. Please file an issue on https://github.com/pytorch/nestedtensor Alternatives If this were possible #89124 it might provide a good/easy workaround. Additional context No response cc @cpuhrsch @jbschlosser @bhosmer @drisspg @mikaylagawarecki,https://github.com/pytorch/pytorch/issues/89125,55093fb9f9b856350513f9150b9e3b8381049f4335285a15e547c87cb7b4093f references,issue,88980,issue,31822,medium,issue.body,"il/logging_is_google_glog.h or torch/include/c10/util/logging_is_not_google_glog.h) and glog. I can see Pytorch have similar issues before (#31822, #14724) and a PR (#41504) to solve the redefinition issue when libtorch is not built with glog. However, I still have the redefin...",https://github.com/pytorch/pytorch/issues/88980,266ae0f152ff4e2f95a5646a85701ded78b86987869efe1abb8080a9d54ba11c references,issue,88002,issue,57121,medium,issue.body,sum for discontiguous tensors. Testing out cuTensor perf (to see if there are significant improvements since Yaroslav's last benchmarks in #57121.) Additional Context See previous relevant issues #60295 and #57121 for context. cc @VitalyFedyunin @ngimel @jianyuh @nikitaved @pe...,https://github.com/pytorch/pytorch/issues/88002,3bff767ae6a429a444f62e2b2f3d6cf8c26e8155873bfe6a740c27d8d2b5d7df references,issue,88002,issue,60295,medium,issue.body,"opt-einsum to optimize the contraction path for when there are 3+ tensors to contract and opt-einsum is installed. As of now, section 1 of #60295 has been completed. Tasks remaining include to: Achieve numpy compatibility by adding optimize as an torch.einsum API kwarg (our pr...",https://github.com/pytorch/pytorch/issues/88002,58598962f909fedb02e81c4a27e8badc792b3343e062f22d7e08109497eb088e references,issue,88448,issue,76578,medium,issue.body,"expected (weight and bias can have different dtypes) ========================================= Original Context: Similar to conv-bn folding #76578, lin-bn folding #86706 with autocast freeze on gpu casts the inputs to linear to half, while inputs to batchnorm are not casted, r...",https://github.com/pytorch/pytorch/issues/88448,09181aa189b3bc20eb03fee008df819448aeb4455147f33bb77cac82913c4acf references,issue,84234,issue,62451,medium,issue.body,"be provided as an API argument, assuming this will guarantee determinism. This is unfortunately not the case (and is further compounded by #62451). Option 3 is unfortunately unrealistic and can't be sustained over time as hardware becomes obsolete. Additional context Related i...",https://github.com/pytorch/pytorch/issues/84234,cf7c7e77a2835cda94f14817989819cc2b4f10605ecb1022cd4ec7c51099b579 references,issue,84234,issue,62451,medium,issue.comments[1].body,"ce): Target device for the resulting tensor """""" # FIXME: generator RNG device is ignored and needs to be passed to torch.randn (torch issue #62451) rng_device = generator.device if generator is not None else device image = torch.randn(size, generator=generator, device=rng_devi...",https://github.com/pytorch/pytorch/issues/84234,577eca3737dbf58bcf28505a50b1eaea7ff10c03a3bbc8b3e145c19f3a628120 references,issue,60333,issue,51018,medium,issue.comments[1].body,dea on how well torch.nn.ModuleList is supported in TorchScript? I haven't encountered issues with torch.nn.ModuleList + JIT before (except #51018).,https://github.com/pytorch/pytorch/issues/60333,b88347ecd2407277aaeda85f59a7fa326aae03640a4cf8ce548a0a6ec60d348e references,issue,87033,issue,75242,medium,issue.comments[0].body,Supporting out-parameter tensor-loading would also be a nice thing! also some relevant discussion in #75242 and #52181 / #79967 (maybe does hdf5 already support slice loading/saving directly to the disk file via mmap?) cc @stas00 @albanD,https://github.com/pytorch/pytorch/issues/87033,a0ab5b36c67177a35d861fd4aab398ba34efd650f9eaca731e6b7f13aa9a1889 references,issue,87033,issue,79967,medium,issue.comments[0].body,Supporting out-parameter tensor-loading would also be a nice thing! also some relevant discussion in #75242 and #52181 / #79967 (maybe does hdf5 already support slice loading/saving directly to the disk file via mmap?) cc @stas00 @albanD,https://github.com/pytorch/pytorch/issues/87033,3a506f3f561072fb383ed8496d7f941f35c9cb01a406d68ddb8eeebbdd17a60e references,issue,87033,issue,64327,medium,issue.comments[1].body,and #64327,https://github.com/pytorch/pytorch/issues/87033,9ee530bf79fff0b6bbaa8ed7d2d297da2cea87f72ac5ac8d4ec9ba9dac7b37f8 references,issue,84321,issue,1529,medium,issue.comments[0].body,"Related issue on allocator state stats and visualization: #1529. If such vis is developed within the usecase of profiler, it probably can be useful for a general vis of allocator state even outside of pr",https://github.com/pytorch/pytorch/issues/84321,7fc5fe5d1d419f732bd51078a9c8b27c626a2ecf17b915f1608e8f79b3f62c14 references,issue,86890,issue,82926,medium,issue.comments[0].body,Somewhat related request possible asking for more view based slice indexing #82926,https://github.com/pytorch/pytorch/issues/86890,9740df5b1604ded08c5588bdb66b5fd58498e9e7b1651022cf2f58fbae02ef97 references,issue,86804,issue,85877,medium,issue.body,"000e+00, 0.0000e+00, -6.5622e+85, 1.6687e+94]], dtype=torch.float64) This happens on both cpu and cuda. Not sure whether this is related to #85877 Versions Collecting environment information... PyTorch version: 1.12.1 Is debug build: False CUDA used to build PyTorch: 11.3 ROCM...",https://github.com/pytorch/pytorch/issues/86804,6b603e0bcfabece995730d01249f5c8a9b1e468fb5626c18cc3a929dba0c6ecf references,issue,56485,issue,56697,medium,issue.body,R conversion #57381 Inefficient conversion between COO and CSR formats #56959 Avoid no-op suggest_memory_format call in SparseCsrTensorImpl #56697 Additional context CSR tensor support in PyTorch https://pearu.github.io/csr_tensor_support.html cc @ngimel @aocsa @nikitaved @pea...,https://github.com/pytorch/pytorch/issues/56485,ea3b5f6ef7f058319896ae1bad0d5706966ddcb419e39f41af6bf668454e2349 references,issue,56485,issue,58770,medium,issue.body,_cuda #59101 Issue: Data access pattern in the loop in add_out_dense_sparse_csr_cuda could be pretty bad #59104 Issue: MKL csr matmul issue #58770 #59011 CSR layout: CPU addmm Fix #59010 CUDA support in the CSR layout: sparse_to_dense/add_sparse_csr Other CSR issues CSR sparse...,https://github.com/pytorch/pytorch/issues/56485,2e6be1d8a4b515ed89d9f51b7f0ec53e8c9be10d4363679bc215b9c03bacc2e4 references,issue,56485,issue,59058,medium,issue.body,port. Status #59012 CUDA support in the CSR layout: matvec Related issues Issue: support auto generation of device check for sparse tensors #59058 Issue: Add 64-bit indices support to csrmm2 #58899 Issue: Relaxing constraints to s_addmm_out_sparse_dense_cuda_worker #59099 Issu...,https://github.com/pytorch/pytorch/issues/56485,4e4d4938354e5a72a9052362d07b3bda4e8485d8f11f91a26c47004bd078c77a references,issue,56485,issue,59099,medium,issue.body,parse tensors #59058 Issue: Add 64-bit indices support to csrmm2 #58899 Issue: Relaxing constraints to s_addmm_out_sparse_dense_cuda_worker #59099 Issue: Use of storage_offset is not needed in add_out_dense_sparse_csr_cuda #59101 Issue: Data access pattern in the loop in add_o...,https://github.com/pytorch/pytorch/issues/56485,5eb573aea9e8482f3da5e1c63fe689bc3ae5ef1e88dceb275b02edf1146bff2d references,issue,56485,issue,59101,medium,issue.body,xing constraints to s_addmm_out_sparse_dense_cuda_worker #59099 Issue: Use of storage_offset is not needed in add_out_dense_sparse_csr_cuda #59101 Issue: Data access pattern in the loop in add_out_dense_sparse_csr_cuda could be pretty bad #59104 Issue: MKL csr matmul issue #58...,https://github.com/pytorch/pytorch/issues/56485,b4243a1e6871cf7c568e53989aabd31d68fae0ce838e3d4989e8a6a21d1243b3 references,issue,56485,issue,59104,medium,issue.body,needed in add_out_dense_sparse_csr_cuda #59101 Issue: Data access pattern in the loop in add_out_dense_sparse_csr_cuda could be pretty bad #59104 Issue: MKL csr matmul issue #58770 #59011 CSR layout: CPU addmm Fix #59010 CUDA support in the CSR layout: sparse_to_dense/add_spar...,https://github.com/pytorch/pytorch/issues/56485,3bc03404a6b0fb295d6e844bee1220252f6901e5ec1bbecb654c719812298473 references,issue,75747,issue,18631,medium,issue.comments[0].body,"eneral, group convolutions are very slow. Nothing has been done to fix it despite many people asking for a fix over the years. E.g. #73764, #18631, #70954, https://discuss.pytorch.org/t/group-convolution-takes-much-longer-than-normal-convolution/92214, https://twitter.com/wigh...",https://github.com/pytorch/pytorch/issues/75747,e0dce3d7ea5559f4f3197eb6308144252e003c95f0f7be3b8f049fbd56e87016 references,issue,75747,issue,70954,medium,issue.comments[0].body,"group convolutions are very slow. Nothing has been done to fix it despite many people asking for a fix over the years. E.g. #73764, #18631, #70954, https://discuss.pytorch.org/t/group-convolution-takes-much-longer-than-normal-convolution/92214, https://twitter.com/wightmanr/st...",https://github.com/pytorch/pytorch/issues/75747,3960df9c0e1ea0edfdbe755609d05f0449af0626bf9733bdaba24fdbe414ecfd references,issue,75747,issue,73764,medium,issue.comments[0].body,"e more general, group convolutions are very slow. Nothing has been done to fix it despite many people asking for a fix over the years. E.g. #73764, #18631, #70954, https://discuss.pytorch.org/t/group-convolution-takes-much-longer-than-normal-convolution/92214, https://twitter....",https://github.com/pytorch/pytorch/issues/75747,e10ac02f2eb5034b40dff68b4a5fa30c0992fc95bd9a6a8034e089e1bdf654cf references,issue,66491,issue,65868,medium,issue.body,"allow list, since there are no overloads accepting scalars, and no wrapping took place during python arg parsing, it results in error (see #65868) Some operations have overloads in native_functions.yaml. To support scalar as either argument of a binary pointwise operation, one...",https://github.com/pytorch/pytorch/issues/66491,0a9dc8badfc1791778b7378201da1a4e7a82cb4c290d5c201bfb002d9f291e03 references,issue,63725,issue,32651,medium,issue.body,"intuition of one event file per one experiment, I cannot use the tensorboard log intuitively. This discussion may be related to this issue #32651 Pitch Current implementation torch._C._log_api_usage_once(""tensorboard.logging.add_hparams"") if type(hparam_dict) is not dict or ty...",https://github.com/pytorch/pytorch/issues/63725,55ffd4c441a3279dad69278849ee84d034d0fdc0bdf0e583858cc59308050fc0 references,issue,85538,issue,34257,medium,issue.comments[0].body,This seems to be a duplicate of #34257. Is there any plans on fixing this?,https://github.com/pytorch/pytorch/issues/85538,fe3ee7a47011760252b4006c150a622cf8490ddb15a3b25100a1172923e12660 references,issue,85607,issue,80104,medium,issue.comments[1].body,"_worker_threads=64) # ADD THIS ) worker_init() # no-op rpc.shutdown() if __name__ == '__main__': main() I've filed this as an issue before (#80104), but will mark it as high-pri since its been hit twice now so we can provide a better error message",https://github.com/pytorch/pytorch/issues/85607,9e70d863ff7d37e6e9433bda6a9ce6d4cc42de4a4ca46fd50a9186d6feaf7e2e references,issue,86055,issue,28090,medium,issue.comments[0].body,dimension. Some limitations: flatten/unflatten require the input tensor to be contiguous and not just those dimensions under modification: #28090 Some other related issues: #71403 #62352,https://github.com/pytorch/pytorch/issues/86055,cc6c733a49f2a697cb0d16ecef7c67357fb5718eedbc7b80e32d37f98eebd788 references,issue,86055,issue,62352,medium,issue.comments[0].body,latten require the input tensor to be contiguous and not just those dimensions under modification: #28090 Some other related issues: #71403 #62352,https://github.com/pytorch/pytorch/issues/86055,880ff219cbf2c789cf54ae833ba75150783821a6eba69a736a370819ebd226ef references,issue,86055,issue,71403,medium,issue.comments[0].body,ten/unflatten require the input tensor to be contiguous and not just those dimensions under modification: #28090 Some other related issues: #71403 #62352,https://github.com/pytorch/pytorch/issues/86055,e5aef74e884e8dafa9bca363dfa8804be6e99583de16784688b7e2fe2b88e979 references,issue,82565,issue,71446,medium,issue.body,"g Schrödinger equation to begin with) and to perform gradient-based optimal control. Alternatives No response Additional context Related to #71446, but it is not active and specific to MIT's library. Don't hesitate to close this issue if you find it inappropriate, and prefer t...",https://github.com/pytorch/pytorch/issues/82565,9bae331f2d6185b9913d4cb61425e292e4ae6746d80a6ad0840baa4db9c51445 references,issue,85834,issue,78071,medium,issue.body,"asses_mask, valid_classes_mask)) # TODO: This check does not work with FakeTensor inputs # Explicit cast for class_check to bool; See Issue #78071 utils.check( isinstance(target, FakeTensor) or bool(class_check.item()), lambda: ""A target class is out-of-bounds and not the igno...",https://github.com/pytorch/pytorch/issues/85834,75929a07dfcffa1257b8a08b6082e9a75fbe93ab7375a5872d413d116d93a94d references,issue,85652,issue,1529,medium,issue.comments[0].body,"it would also be nice to be able to have some standard alloc/free tracing hook/callback available + python-bound methods to install it: #1529 (comment) - for helping debug allocation/deallocation issues to find where the memory usage is coming from, e.g. as in #85698)",https://github.com/pytorch/pytorch/issues/85652,e6d8e13c796e5b783db75c5f252fdc02dbcf412bbd6f7363bc88a2d34e5bb4cf references,issue,85652,issue,85698,medium,issue.comments[0].body,"to install it: #1529 (comment) - for helping debug allocation/deallocation issues to find where the memory usage is coming from, e.g. as in #85698)",https://github.com/pytorch/pytorch/issues/85652,237fd13b79a2be04158460cae01adf17947c2030778d0d7b598202846a095fd2 references,issue,85792,issue,80561,medium,issue.comments[0].body,Found the same error #80561. The TorchScript used in my code is also a part of a BERT model.,https://github.com/pytorch/pytorch/issues/85792,3dfc966657ff24cb7281eaeab8a7f3c082ee806eb11408e3efc879daf0f5a40b references,issue,64254,issue,29973,medium,issue.body,ensor (for large Python list the motivation would rather be saving memory that would be spent on the copy) Motivation: #46138 (comment) and #29973 (comment) Alternative: make copy_ and assignment do this directly For hot-path assignments of small sequences (e.g. 2 or 3 element...,https://github.com/pytorch/pytorch/issues/64254,41a650714c2581c0606140471f2b6ff50ebf3ac2ca11c5f71f9afa9cea2c935f references,issue,60295,issue,21760,medium,issue.body,ion. cc @jianyuh @nikitaved @pearu @mruberry @heitorschueroff @walterddr @IvanYashchuk @xwang233 @lezcano @VitalyFedyunin @ngimel see also: #21760,https://github.com/pytorch/pytorch/issues/60295,d7010116e1d54b173e9786d1443c2d418243b68cda711b2d4f36d06bacc4127d references,issue,60295,issue,57121,medium,issue.body,"re contiguous, call torch.tensordot, otherwise call a custom function to compute a sum_of_products over discontiguous dimensions. The issue #57121 illustrates a case where PyTorch is much slower than NumPy because the dimensions are discontiguous. A sum_of_products function co...",https://github.com/pytorch/pytorch/issues/60295,c4cf5f2e4af635a115cfcc1036070bf9904d4453046a8b98ea7709cae24af587 references,issue,60295,issue,57121,medium,issue.comments[0].body,"hich would not change the perf numbers (matmul is doing this internally), but would make the code less clear. This won't work for case 4 in #57121 where contraction dimensions are discontiguous.",https://github.com/pytorch/pytorch/issues/60295,d7f0daf654ba317c862ec619b537f56a30cf4aa56abfbe448e6086a0ec49e916 references,issue,80104,issue,80017,medium,issue.body,🐛 Describe the bug See issue raised in #80017. We should provide a better error message when we can no longer send requests due to exhaustion of threads in thread pool. Docs: https://py,https://github.com/pytorch/pytorch/issues/80104,a43f1bc67740ff29c0127a6d6d4d9b1d03409835b2aedd3156980f7c0d64018b references,issue,78050,issue,77731,medium,issue.body,"ementations with PyTorch's eager mode. This has proved to be an issue for several reasons: 1) PyTorch eager's striding is inconsistent. See #77731 and #77553 for some examples. @ngimel has fixed several of these issues on CUDA, too. See #77610 and #77585. These issues suggest...",https://github.com/pytorch/pytorch/issues/78050,ed860e567e2af71bb7717e350c26410289f0a46e0567fbab86d454b0b1b8d0b8 references,issue,78050,issue,72341,medium,issue.comments[0].body,"icularly affected by inconsistent striding when it does occur. I disagree. People do notice when strides are not preserved correctly; e.g., #72341 I accept that there is probably a cottage industry of bugs where we don't handle things correctly and no one really notices, but t...",https://github.com/pytorch/pytorch/issues/78050,f250553b01f33d78ac49624ec01aed1cfb7e361a7c5e8f3fe4fdaef07ff6f55d references,issue,79039,issue,83773,medium,issue.body,ew parts to making this work and the sub components will be tracked in related issues referenced below. Linked Issues: #79040 #79447 #79044 #83773 Nice to have but there is a work around: #79046 #79048 cc @cpuhrsch,https://github.com/pytorch/pytorch/issues/79039,d75553ff60329ef0421056409ebe9251076f22f8f6b3a44c5bc7bc5fa060264f references,issue,84565,issue,69991,medium,issue.body,"Tensor subclass, causing it to go down a non-differentiable path. See #84137 for example. This problem is related to composite compliance: #69991 cc @ezyang",https://github.com/pytorch/pytorch/issues/84565,757817fcdfd641554fb435d7455bc276f855f2791aeb6955435249e81bef5b08 references,issue,83968,issue,77764,medium,issue.comments[1].body,: operations that are not implemented yet on MPS devices. This is in the domain of the pytorch team/community and there is an open issue at #77764 you can comment on.,https://github.com/pytorch/pytorch/issues/83968,38910bdbe35a255399333574a400018f4a1702f7291b3d86daf466534d3222b8 references,issue,67590,issue,44511,medium,issue.body,"that optimizer.step() has been called if they don't know their model produced these types of values. An example of a bug induced by this is #44511 where we want to run scheduler.step() only after optimizer.step(), but there's no easy way to be sure that happened when the GradS...",https://github.com/pytorch/pytorch/issues/67590,cdab91efea49d18797ae7fe671d0de8d79bbd2f31845acb977c6e92994a5cdda closes,issue,67590,issue,55585,high,issue.body,error-using-gradscaler/92930/8. IMO this is an ugly solution that relies on the details of the scaler values. Additional context Also fixes #55585 The warning UserWarning: Detected call of lr_scheduler.step() before optimizer.step() is impacting Lightning (Lightning-AI/pytorch...,https://github.com/pytorch/pytorch/issues/67590,c2e42d788639e3d47af7dbc116d4570f0562f1c52b486f323ee23ccd5b12d66e references,issue,67590,issue,67589,medium,issue.body,"turn retval + optimizer_state[""stage""] = OptState.STEPPED + return This feature request is particularly valuable with the implementation of #67589. With the addition of both, one could do: scaler.scale(loss).backward() scaler.step(optimizer) if scaler.state(optimizer) is OptSt...",https://github.com/pytorch/pytorch/issues/67590,d3aa524ddd6eaf9b56b84587396c36d3b8b7f85d8c1a162654d4da0e910fd062 references,issue,42812,issue,7617,medium,issue.comments[1].body,"For future readers, please see #7617.",https://github.com/pytorch/pytorch/issues/42812,c28ce65d92449f10285bbecd2ff56beb5eaa36695b60dbaa807e018cc9d55d5d references,issue,61653,issue,58828,medium,issue.body,"e with the rest of torch.linalg Implement the backwards for non-symmetric matrices (would fix #38948) Address the performance concerns from #58828 Review the API and perhaps split it into two functions, one with a simpler API and a simpler name and a full one (needs some propo...",https://github.com/pytorch/pytorch/issues/61653,90c83423a07f8749d5a0b33363358be4934960a3be473fe9a77bc43ac7789021 references,issue,61653,issue,58828,medium,issue.comments[0].body,"@lezcano I'm wondering if there's an update on this. Have the performance concerns from #58828 been addressed in the latest linalg/sparse modules? If lobpcg is still slow, might you consider adding the Arnoldi/Lanczos solvers that I s",https://github.com/pytorch/pytorch/issues/61653,87780dd12334da26e79bb2831db881ef90405f3e4e250facb54eddf20575ffc4 references,issue,72146,issue,71683,medium,issue.body,ames lists as I do above. Similar normalization is happening already when a list of params is converted to a dict-like param group Related: #71683 Alternatives No response Additional context No response cc @vincentqb @jbschlosser @albanD,https://github.com/pytorch/pytorch/issues/72146,92de13a8de5d7b122e92ae4bc631ee325a8e9c5f3fe219c9c3dcdc8606627abf references,issue,83271,issue,68385,medium,issue.comments[0].body,Seems like the same issue as #68385. @malfet @janeyx99 @seemethere do we have plans to build distributed as part of the M1 PT binaries?,https://github.com/pytorch/pytorch/issues/83271,efbfff1b1ae48d7107158940beb99100bd38abc2b2551bd9c2aec9650df45af6 references,issue,81680,issue,55267,medium,issue.body,"o CUDA_GCC_VERSIONS that were done in 86deecd #63230 I do think they are a step in the right direction toward improving the situation (xref #55267) I'm wondering if you could add an option, as an environment variable, for us to ignore it. The motivation is that I believe that...",https://github.com/pytorch/pytorch/issues/81680,991cc362844b8b7684ef3e52e68f2631f9f9c5ad6c95517ec0bdea0fddfba5ea references,issue,82761,issue,16897,medium,issue.body,"yTorch could also implement the equivalent of numpy's np.random.choice(), which has an outstanding pull request as described in this issue: #16897 Additional context No response",https://github.com/pytorch/pytorch/issues/82761,0883d975144c7801f1c571e3445ef53e50bf0557195bcb2727bb73b19b5c5218 references,issue,54408,issue,47908,medium,issue.body,"ch 1.8.0 + CUDA 10.2 = 0.6 ms RTX 3090 + PyTorch 1.8.0 + CUDA 11.1 = 0.94 ms RTX 3090 + PyTorch 1.7.0 + CUDA 11.0 = 1.12 ms I also found in #47908 that PyTorch with CUDA 11.0 has problem in efficiency, right? Since RTX 3090 cannot use PyTorch with CUDA 10.2, it is a little pit...",https://github.com/pytorch/pytorch/issues/54408,69595e362d069bcdec4514f68bd22bd828c8f9bd8adbae52053f0e4eec9a232f references,issue,54408,issue,47908,medium,issue.comments[1].body,"I tested the time on each GPU for 1000 rounds. I don't think it is measuring error. I do not mean RTX 3090 itself is low. Since I also read #47908 , I guess it is because PyTorch with CUDA 11 has the problem in efficiency, and RTX 3090 has not choice but to use it. If you are...",https://github.com/pytorch/pytorch/issues/54408,f7301c34ee416b2a16c0ec7ff72386816b533cc1d2fd7f3fa2a36d45ab5c7088 references,issue,81541,issue,82072,medium,issue.comments[1].body,Identified a possible fix in #82072,https://github.com/pytorch/pytorch/issues/81541,aa2d1bb82a28a480a6e0b127a4648e9a4f456be09c109c9ce88dcb58387ee332 references,issue,65357,issue,64905,medium,issue.body,"sion, static graph, etc Show how to do performance analysis/tuning Show how to do inference in distributed setting and aggregate results as #64905 requests. Recover from batches that cause OOM/inf/nan grads. Overall this will make the DDP tutorial more comprehensive and simila...",https://github.com/pytorch/pytorch/issues/65357,bed94d93afd94d458ca862bc045db7b3f632ac18ab3f1ded58dc6aa31fa5f0b1 references,issue,65357,issue,64636,medium,issue.comments[0].body,"Also, OOM recovery and skipping batches: #64636 (per replica or globally) And error / exception / stack trace marshaling Correct implementation of training log logging (both into text jso",https://github.com/pytorch/pytorch/issues/65357,b6881dd104bf9a83b051f1d2b9708b2043047bafc4a2ae0e70dda726acb093e6 references,issue,25310,issue,25032,medium,issue.comments[0].body,"@johncwok we are working on Nested Tensors, a generalized version of Lists of Tensors. Tracked in #25032 Maybe that'll help. As to the question of why more stuff doesn't support packed sequences, they were written as a limited datatype to be co",https://github.com/pytorch/pytorch/issues/25310,3f005a036642cfceedad3bdf6747cac93d7f2dbe240f522a049ca4b6282dae41 references,issue,59629,issue,25310,medium,issue.body,"d passing that as an argument to the recurrent layer. Indeed this is a technique recommended in other issues and the forums. For example in #25310 @andreaskoepf elegantly suggests creating a function such as def squash_packed(x, fn=torch.tanh): return torch.nn.utils.rnn.Packed...",https://github.com/pytorch/pytorch/issues/59629,72ea8d6c2ed05d30ef31bfb6dde0a15db3c0429632889bbf3b345aa4821abae4 references,issue,38767,issue,42843,medium,issue.comments[0].body,"No offense, but it is very badly documented. This issue has been raised several times, e.g. #50464 #42843 #43055 The Adam algorithm implements the L2 regularization, and for long time it implied otherwise. Now it is just short on documentation,",https://github.com/pytorch/pytorch/issues/38767,95c8e5a5c6a3162bda55812d07d70551a2afa5f5bcc6014637e43171be4278de references,issue,36107,issue,24870,medium,issue.body,This issue is expanded from #24870 for reference and for additional discussions on implementation details. See that issue for context. Add an option for the grid/flow input t,https://github.com/pytorch/pytorch/issues/36107,4967024a7e3d09b7b39d8fe06427c924608d6557460401855ce6ccfcd92d144d references,issue,36107,issue,36108,medium,issue.body,"1] coordinates, which is prone to user error, (especially when users don’t know which setting of align_corners to target). This, along with #36108, would help simplify the implementation of convolutional networks to do Optical Flow Estimation and Image Registration tasks. cc @...",https://github.com/pytorch/pytorch/issues/36107,5991c3d92d34f7ee67a0d168011951e2805f2c050b1b3f6bd3dee2bc785644b0 references,issue,36107,issue,81868,medium,issue.comments[1].body,"As detailed in #81868, I am encountering precision issues with grid_sample in bilinear mode at discrete pixel locations. Simple experimentation seems to point to",https://github.com/pytorch/pytorch/issues/36107,257748305e5359b9a5802903b70e6a2e4abb06f30dfdd8ab061e6b9867eb3b9b references,issue,80966,issue,80187,medium,issue.comments[0].body,"ative_layer_norm_backward, like var_mean, should not be decomposed for nvfuser and should instead be directly sent to it, see discussion in #80187 And embedding_dense_backward shouldn't be sent to nvfuser at all, nvfuser doesn't support it.",https://github.com/pytorch/pytorch/issues/80966,fc77663e1b1433781413fd023282ce6c8fc98d9d50a9c168ef6b8c347ef3d593 references,issue,80966,issue,80187,medium,issue.comments[1].body,"ative_layer_norm_backward, like var_mean, should not be decomposed for nvfuser and should instead be directly sent to it, see discussion in #80187 And embedding_dense_backward shouldn't be sent to nvfuser at all, nvfuser doesn't support it. This implies a few things we need to...",https://github.com/pytorch/pytorch/issues/80966,9928cd8c94034c086285ad6b773d025bf0b6177219784219c402d4d648193821 references,issue,77332,issue,77223,medium,issue.body,", type(arg)) ElementwiseMulScalarIntModule()(TrivTensor(torch.randint(10, (3, 4)))) > aten.mul.Tensor > 4 might be related to #77223 cc @Chillee @ezyang @zou3519 @albanD @samdow @silvasean @cathyzhyi @powderluv Versions PyTorch version: 1.12.0.dev20220511+cpu Is...",https://github.com/pytorch/pytorch/issues/77332,4a23857743142e819e4a1fe40e1c20568cf1cb61848cdfa63e74dea95ca3e0a9 references,issue,70485,issue,67970,medium,issue.body,backward_compatible=True) TraceError: symbolically traced variables cannot be used as inputs to control flow This issue might be related to #67970 Versions PyTorch version: 1.10.1 Is debug build: False CUDA used to build PyTorch: 10.2 ROCM used to build PyTorch: N/A OS: CentOS...,https://github.com/pytorch/pytorch/issues/70485,329fe1621d0959ff488594ebcba3178a0131de28d264a4100dd4d22f3f216925 references,issue,54602,issue,44901,medium,issue.body,🐛 Bug if conditional cannot be symbolically traced. Related to #44901 To Reproduce Steps to reproduce the behavior: import torch # Simple module for demonstration class MyModule(torch.nn.Module): def __init__(,https://github.com/pytorch/pytorch/issues/54602,170925a61a636d6dd672c92b403faccf47ff71ede846b0f271fa73c60ff1b31d references,issue,54602,issue,51803,medium,issue.comments[1].body,"Hi @jamesr66a With custom Tracer, the control flow can be used by Trace,Please see #51803 (comment)",https://github.com/pytorch/pytorch/issues/54602,188b6c992adda099297a4c047c2ac85030eb77bf7ba3ddd4df42bb33dafb872d references,issue,62021,issue,53534,medium,issue.body,"not work - as the docs say: ""This function can be called at module-level scope"") Another similar feature request with a similar motivation: #53534 Alternatives I saw some talk somewhere of creating an is_leaf_function method to override in the tracer class. Also seems viable....",https://github.com/pytorch/pytorch/issues/62021,47a5ed663ebd9242b0e522f22d4bb209530f7be146a503d73bf9f90b6eac3022 references,issue,71404,issue,66335,medium,issue.body,"ically all models, but have different names: reset_parameters, _reset_parameters, _init_parameters etc Related issues: pytorch/vision#3410, #66335 Periodically reinitializing parts of model is useful in some training scenarios, as it can prevent overfitting (e.g. used in http:...",https://github.com/pytorch/pytorch/issues/71404,3e9b212f9e56d96dacb23e7e7651726086f7a5cd6c31ec4f244c02ea587dd765 references,issue,71404,issue,66335,medium,issue.comments[0].body,ue highlighted here by @vadimkantorov is related to the fact that nn.Module class info is lost for custom layers during tracing (related to #66335). This means that standard model initialization idiom of traversing the modules() of the model and checking their instance via isi...,https://github.com/pytorch/pytorch/issues/71404,15e6bd2614673b1843b2f735dd854202e63b08c59cf876d4ebcff69bec5cf36e references,issue,58036,issue,54138,medium,issue.body,"e it can be converted with np.asarray. As side-effect, this would also eliminate the need for torchvision's F.to_tensor(pil_image) Related: #54138 cc @mruberry @rgommers @heitorschueroff",https://github.com/pytorch/pytorch/issues/58036,92310c24df2bd2d0ff6c6abc3258bf68c73ce98737819fd0b20bb62b522d362f references,issue,73537,issue,73513,medium,issue.body,"ybind11, pybind11::args) const () from /data/users/ezyang/pytorch-tmp/torch/lib/libtorch_python.so I'm guessing it has something to do with #73513 Versions #73441",https://github.com/pytorch/pytorch/issues/73537,fd38e1c94d1f6552501f23aaf2116722445a5b00ddf6f03f8f44145cd441698f references,issue,81413,issue,52743,medium,issue.comments[0].body,Related: #52743,https://github.com/pytorch/pytorch/issues/81413,793fa5e64ebae14825c223bb710700aa95b256a36ec07c5dfe1c393807eb0d04 references,issue,80301,issue,3867,medium,issue.comments[0].body,"nks for the request! I think this is something we do want to support, but there are some implementation subtleties to get right. Details in #3867 and this post.",https://github.com/pytorch/pytorch/issues/80301,3cbd690fa816233f908ebe3abe0a03f4dbad94552171547fa20693cb7ba0e7cb references,issue,80337,issue,51872,medium,issue.body,"s the indices of non-zero elements, or raises a RuntimeError that can be caught and recovered from. Additional Context: Possibly related to #51872? Versions PyTorch Version: 1.11.0+cu113 Is debug build: False CUDA used to build PyTorch: 11.3 ROCM used to build PyTorch: N/A OS:...",https://github.com/pytorch/pytorch/issues/80337,99fdf5fd98364b0b54c6146bf46840c9616cf11c1604c0225ed799784f08433e references,issue,80208,issue,78878,medium,issue.comments[0].body,"y do I have to manually convert targets to float when F.cross_entropy expects a LongTensor and not FloatTensor. This was also brought up in #78878. As mentioned there, we'd accept a PR implementing integral target support within binary_cross_entropy_with_logits to avoid the ne...",https://github.com/pytorch/pytorch/issues/80208,7162acba238061472975696936914099c00dbeb7d62f67497fff7578aa4953e3 references,issue,80157,issue,76324,medium,issue.body,c functions and integrals as PyTorch operators. Enjoy! One of a five-part series of special functions issues: Bessel and Related Functions (#76324) Elliptic Functions and Integrals (#80157) Gamma and Related Functions (#78065) Orthogonal Polynomials (#80152) Parameterization M...,https://github.com/pytorch/pytorch/issues/80157,16d5c755a5ab93f20bc34ba34251449ce5f583ec62acb0c2365cebfdcfe8d9c4 references,issue,80157,issue,78065,medium,issue.body,"s of special functions issues: Bessel and Related Functions (#76324) Elliptic Functions and Integrals (#80157) Gamma and Related Functions (#78065) Orthogonal Polynomials (#80152) Parameterization More than any other special functions, elliptic functions and integrals are expr...",https://github.com/pytorch/pytorch/issues/80157,1e4e2f44c0f6a01b9c23d66e5098de74e50554540f418f340a34f20bb64e5591 references,issue,80157,issue,80152,medium,issue.body,"essel and Related Functions (#76324) Elliptic Functions and Integrals (#80157) Gamma and Related Functions (#78065) Orthogonal Polynomials (#80152) Parameterization More than any other special functions, elliptic functions and integrals are expressed in various ways. In partic...",https://github.com/pytorch/pytorch/issues/80157,bf177d3fcaa5f53273956451678f45d24745e2e1a3be2790873edd615f8c8195 references,issue,80152,issue,76324,medium,issue.body,operators. Enjoy! One of a five-part series of special functions issues: Gamma and Related Functions (#78065) Bessel and Related Functions (#76324) Orthogonal Polynomials (#80152) Elliptic Functions and Integrals (#80157) API Chebyshev Polynomials chebyshev_polynomial_t(input:...,https://github.com/pytorch/pytorch/issues/80152,40de737e9baed2a433d790a3c6cbde70f28d515a615cedfd23054781128a747a references,issue,80152,issue,78065,medium,issue.body,of orthogonal polynomials as PyTorch operators. Enjoy! One of a five-part series of special functions issues: Gamma and Related Functions (#78065) Bessel and Related Functions (#76324) Orthogonal Polynomials (#80152) Elliptic Functions and Integrals (#80157) API Chebyshev Poly...,https://github.com/pytorch/pytorch/issues/80152,e2aeacd14fbea2ed41d31b4fe8d4ba4c130ba4869140c05232c5252b6d118d7d references,issue,80152,issue,80157,medium,issue.body,"amma and Related Functions (#78065) Bessel and Related Functions (#76324) Orthogonal Polynomials (#80152) Elliptic Functions and Integrals (#80157) API Chebyshev Polynomials chebyshev_polynomial_t(input: Tensor, n, *, out=None) → Tensor Chebyshev polynomial of the first kind $...",https://github.com/pytorch/pytorch/issues/80152,f0e4052f0ed8a7ef037d9fd937128e04f5e9bb8cab8927d50487665eac17ca16 references,issue,79888,issue,53712,medium,issue.body,issues There seems to be a similar issue in the ReduceLROnPlateau scheduler described in issue #62475 Similar issues are also described in #53712 Versions Collecting environment information... PyTorch version: 1.10.2+cu113 Is debug build: False CUDA used to build PyTorch: 11.3...,https://github.com/pytorch/pytorch/issues/79888,52c0230acea389915d38f089cfd1eef5fff5182a4ed99a283635157b50280871 references,issue,79888,issue,62475,medium,issue.body,does not match it's current length. Related issues There seems to be a similar issue in the ReduceLROnPlateau scheduler described in issue #62475 Similar issues are also described in #53712 Versions Collecting environment information... PyTorch version: 1.10.2+cu113 Is debug b...,https://github.com/pytorch/pytorch/issues/79888,bcd250b38bae402262a073e8c35bce29940aa1c9711c36aa180a3fe7118dbf83 references,issue,62094,issue,51471,medium,issue.body,"state_dict. This limits control over what is serialized. Use Cases Serialize arbitrary objects within a module (flags, integers, etc.) See #51471 Define ""buffers"" that are serialized but don't take part in device moves / floating pointing dtype conversions (e.g. to(), cuda(),...",https://github.com/pytorch/pytorch/issues/62094,9f312be18dede33d29a86df1bea4b4d6168f7e23a61910eb9510337ed373b9ec references,issue,31945,issue,2129,medium,issue.body,"urn log(erfcx(x)) Motivation These special functions are very useful whenever we have to work with truncated normal distributions. Related: #2129, #32293 cc @mruberry @rgommers @heitorschueroff",https://github.com/pytorch/pytorch/issues/31945,572f153480f1782c1b67f682a6f5968846524b789c793e63af7d834bd741aa67 references,issue,31945,issue,32293,medium,issue.body,"(erfcx(x)) Motivation These special functions are very useful whenever we have to work with truncated normal distributions. Related: #2129, #32293 cc @mruberry @rgommers @heitorschueroff",https://github.com/pytorch/pytorch/issues/31945,a96b948b732a3b665b84dda30eb30d73dc7f810bc80a0b7b628ffccfafb18e18 references,issue,68871,issue,61693,medium,issue.body,ine #set(USE_LAPACK 0) then the build run successfully with LAPACK included (no error when running on iOS). Some related issues are: #66543 #61693 https://discuss.pytorch.org/t/lapack-library-not-found-in-compilation-pytorch-mobile-android-java/126849 https://discuss.pytorch.o...,https://github.com/pytorch/pytorch/issues/68871,31a5ada75942884ed8d9abd28942bb88366092df7bcce44c6ef48acea2934251 references,issue,78917,issue,68332,medium,issue.comments[0].body,"e current lr, not an absolute value of lr. I find this super-unintiutive. Some of my grievances about the scheduler design are listed here: #68332",https://github.com/pytorch/pytorch/issues/78917,f1ed278e566650a4437cff32936bdb871799f087dfefe47c66508a83eb3d6193 references,issue,77200,issue,76865,medium,issue.body,ow object collectives / rpc to send tensors around without copying them to the CPU first. Additional context This is a known issue for PTD: #76865 cc @mruberry,https://github.com/pytorch/pytorch/issues/77200,7110e1b601d6aa017399e8efe0d3a23859aae6233da8c034db11c2253e45cd30 references,issue,78738,issue,37160,medium,issue.comments[0].body,related #37160,https://github.com/pytorch/pytorch/issues/78738,de4b3208e135e0a6963c648726082d83753d48f70aba37fa49913cf3de805de1 references,issue,77869,issue,60953,medium,issue.body,"bug Currently pytorch will not thrown any obvious warnings / errors when PYTHONOPTIMZE(-O and -OO) flags are used. [#76619, #76034, #76659, #60953] This might imply that the behaviour is consistent whether the flags are enabled or disabled. However this is not true. Currently...",https://github.com/pytorch/pytorch/issues/77869,0256fc43dc4325da09c45db6bcebf7cc639a243451856d7f22908c19a4add245 references,issue,77869,issue,76659,medium,issue.body,"ibe the bug Currently pytorch will not thrown any obvious warnings / errors when PYTHONOPTIMZE(-O and -OO) flags are used. [#76619, #76034, #76659, #60953] This might imply that the behaviour is consistent whether the flags are enabled or disabled. However this is not true. Cu...",https://github.com/pytorch/pytorch/issues/77869,bde2dc53f54871252302173376adc11763ef7ecad1e4a6794d4d4b703ade16a5 references,issue,47702,issue,47688,medium,issue.comments[0].body,"See also #47688, which is another issue on the same topic.",https://github.com/pytorch/pytorch/issues/47702,adff29753af54b017ab44daeae245ffdfd5fe92f5cf1ecf512070ed6277d7245 references,issue,28341,issue,22402,medium,issue.comments[0].body,mplementation. Related issue -- #24015 A way to pass LinearOperator into functions that traditionally expect torch.Tensor. Related issue -- #22402 Other examples of structured linear operations: doc,https://github.com/pytorch/pytorch/issues/28341,3aabf5c034567ad8a66c7fb0660f0b640495ae6ad4d71b086e27c9e242ad5c8f references,issue,74519,issue,63034,medium,issue.body,"rror, for TestNllLossBackward) @wconstab ASAN: test_ts_opinfo.py failure for test_correctness_all_cpu_float32 (error log paste) (Related to #63034) Cuda: TestAmpForeachNonFiniteCheckAndUnscale fails, needs to be debugged",https://github.com/pytorch/pytorch/issues/74519,9268c8050c21360cb22bf0d22fd5c4c10220a483491d8bb9f92adbadd4a72405 references,issue,76558,issue,52332,medium,issue.comments[0].body,The perf difference seems to be due to manual fast implementation (but less accurate implementation see also #52332) of operator/ for c10::complex vs the llvm implementation. c10 implementation pytorch/c10/util/complex.h Lines 247 to 257 in a0cc38e conste,https://github.com/pytorch/pytorch/issues/76558,0418d460c8babb2b5fdf2cfff3e49e79a211af0dfb0f8c48e608f29d85c2f75e references,issue,76354,issue,75363,medium,issue.body,🐛 Describe the bug failure logs nn_functional_binary_cross_entropy_with_logits (#76768) nn_functional_conv_transpose3d (#75363) nn_functional_nll_loss (#76768) nn_functional_prelu (#76768) unique_consecutive (#76571) unique (#76571) Versions https://github.com/pytor,https://github.com/pytorch/pytorch/issues/76354,6e55fd14db0bbbadea76eeb58ea247915d16f7f57afd2ac92cad4644c11b1bfe references,issue,77184,issue,60341,medium,issue.body,"undefined symbol: _ZNK2at6Tensor5dtypeEv Things we have tried include explicitly specifying extra_ldflags=[""-ltorch_cpu""] (as suggested by #60341) and clearing the torch_extension cache (as suggested by #68905), but neither of these has worked. We suspect this issue is due to...",https://github.com/pytorch/pytorch/issues/77184,3263f2276c84b7de30c990b993a46caf3d5d386a3064c02afdfbd7b27f751473 references,issue,77184,issue,68905,medium,issue.body,"nclude explicitly specifying extra_ldflags=[""-ltorch_cpu""] (as suggested by #60341) and clearing the torch_extension cache (as suggested by #68905), but neither of these has worked. We suspect this issue is due to some compiler version mismatch (note that we use GCC 5.4 in the...",https://github.com/pytorch/pytorch/issues/77184,6bd40c2eef4223175e9de8daf1d229965ce85b4c7a57f0a0d5bac9a79cdf7786 references,issue,77166,issue,61658,medium,issue.comments[0].body,"ve cholesky_solve, and expect that the user uses the construction above if they want to materialise the inverse. This is already tracked in #61658",https://github.com/pytorch/pytorch/issues/77166,9050bbd7693c5316970b3be8ae116533df577a5467ebda45f576b516d456a00a references,issue,77140,issue,42502,medium,issue.comments[0].body,"Somewhat related: #42502 in the sense that if the end-goal is to call randperm multiple times, maybe a batched mode could express it better - but the threads concer",https://github.com/pytorch/pytorch/issues/77140,45b09799e219a3f8b6ddda3a1be57922f3953cc443e9c1adce6275c0842e477d references,issue,76853,issue,76483,medium,issue.body,"large, extremal) were skipped do to errors of the form: Greatest absolute difference: nan at index. Followed from working on these issues: #76483, #74279 Versions Not version specific cc @ezyang @anjali411 @dylanbespalko @mruberry @lezcano @nikitaved",https://github.com/pytorch/pytorch/issues/76853,acc41b90375a6ec34361ce7e200a90fea980afdf37aec1c6d301ffe66b2aab33 references,issue,54133,issue,24243,medium,issue.comments[0].body,"This seems to be related to #24243. If I replace x.requires_grad_(True).split(1, -1) by x_splitted = list(x.requires_grad_(True).split(1, -1)) It throws another error (only o",https://github.com/pytorch/pytorch/issues/54133,730e4c202084ad29226e49b86a0fe2da92b77b75dceddbbcfb3b5e5491d8d068 references,issue,47953,issue,53879,medium,issue.body,ch 1.9 linear algebra development plan See #47953 (comment) For detailed MAGMA mechanism See #47953 (comment) See also #42666 #26996 #42403 #53879 cc @ezyang @gchanan @zou3519 @bdhirsh @ngimel @vishwakftw @jianyuh @nikitaved @pearu @mruberry @heitorschueroff @walterddr @Vitaly...,https://github.com/pytorch/pytorch/issues/47953,0a7834249f539519967492b0966d90b3420a8c6c6dd2abf068a617b3511652c8 references,issue,47953,issue,23940,medium,issue.comments[0].body,Cross-posting an issue I created a year ago: #23940 and an issue I created two years ago: #13546.,https://github.com/pytorch/pytorch/issues/47953,67eed97d00bb6e0f4bdec8d8c87f16d54003bfc798367834fee46d8a8189d491 references,issue,76806,issue,74616,medium,issue.body,"dtype torch.double: b / a : tensor([1.0000, 0.5000, 0.3333], dtype=torch.float64) If we pursue this we should consider it in the context of #74616. I think type promotion works as expected in scenarios where the dunder is expected to be used: 1 / torch.tensor((1.), dtype=torch...",https://github.com/pytorch/pytorch/issues/76806,e3aa4abd92a2dc92d3aed8bfd55fb8501ea8abc392eadac52842ae4306901043 references,issue,46164,issue,45897,medium,issue.body,to torch.mm. Remove spmm and dsmm functions from pytorch API starting from PyTorch version 1.9 (?). Additional context Related discussions: #45897 #45400 (comment) cc @vincentqb @aocsa @nikitaved @pearu @mruberry @vishwakftw @jianyuh @heitorschueroff,https://github.com/pytorch/pytorch/issues/46164,9879b7b82eb95f692fd67514c64270c6c3a26680e7d516631df7ae9eb6e19a1e competes with,issue,76644,issue,46164,medium,issue.comments[1].body,using torch.sparse.mm instead of torch.smm. torch.smm (and related torch.sspaddmm) are planned to be deprecated and removed eventually (see #46164).,https://github.com/pytorch/pytorch/issues/76644,cb7064899241e765eaeb48429c5e5a46338cab901d0da9ea8f38a28fcfdd9a96 references,issue,76655,issue,76654,medium,issue.body,try to sort 1 dimention tensors. Alternatives Maybe you could implement merge sort or other parallel sorting algorithms Additional context #76654 I have a report and code here import torch import torch.nn as nn import torch.multiprocessing as mp import torch.distributed as dis...,https://github.com/pytorch/pytorch/issues/76655,1e75f199ac2e174b0c4b48eb619f06a90a9b4412e2419dda1a40f2200d1460b2 references,issue,76654,issue,76655,medium,issue.comments[1].body,"and the remaining dimensions will have more than >33k elements, you will see a nice speed-up on a multi-core machine. Sure! I opened here. #76655",https://github.com/pytorch/pytorch/issues/76654,28e0419a4b474b296d38d355d4ad98a88bc61533217b693a2b4b3b3b388ec651 references,issue,76555,issue,66491,medium,issue.comments[1].body,semi-related: #66491,https://github.com/pytorch/pytorch/issues/76555,d8e321c41a05ebc9c9882ab9fe7e1cb2d69d2c4094b6e515385d5dd47619ebd8 references,issue,71249,issue,20117,medium,issue.comments[1].body,@jaketae Looks like there was some previous discussion about append in #20117; unclear why the attempt to add it failed. I think we'd accept a modern PR for this since I don't see any reasons listed not to. @albanD wd,https://github.com/pytorch/pytorch/issues/71249,ba111f1fab75028e432baad614995155fb1a79ad213cfd14998e762408446cc6 references,issue,76514,issue,76433,medium,issue.body,"n these cases, we can provide decompositions so that backends that don't want to deal with them no longer need to. Removable: _log_softmax (#76433) _softmax (#76433) _log_softmax_backward_data (#76433) _softmax_backward_data (#76433) _unique2 (previous attempt: #18655) Decompo...",https://github.com/pytorch/pytorch/issues/76514,416c4ddac5dbbf399257736c402ee0dec611a1bb7ff1842611c463bc47ddeb9f references,issue,76012,issue,71465,medium,issue.comments[0].body,"ls-last format to begin with, you wouldn't need to pay permutation cost. We have an issue open for layerNorm to work on arbitrary dimension #71465",https://github.com/pytorch/pytorch/issues/76012,bc71461e2edd342849602ef0d6b83f40452905c5b9116e27c28ac5f517ec3626 references,issue,75912,issue,21018,medium,issue.comments[0].body,"Possibly related to #21018, however in that issue it sounds like the error occurs right away. In my case, the error only occurs at the end of the script. (In the real",https://github.com/pytorch/pytorch/issues/75912,fdf48a6c0a2abc26d3e141bd8a43f1d90e35a3c760af2b9bd09495e55711c227 references,issue,72345,issue,69687,medium,issue.body,see #69687 example PR: #71129 Benchmark the forward mode AD perf for your change!,https://github.com/pytorch/pytorch/issues/72345,2abd806dad065581eb9e33327c50a92e2414271bea50e6f7f0694711c950262d references,issue,72362,issue,69687,medium,issue.body,see #69687 example PR: #71129 Benchmark the forward mode AD perf for your change!,https://github.com/pytorch/pytorch/issues/72362,1ae555f2f572d85ea2c0bf6b5cda5959c4dcd20a171e2eb69bb9a778b42bd030 references,issue,72359,issue,69687,medium,issue.body,see #69687 example PR: #71129 Benchmark the forward mode AD perf for your change!,https://github.com/pytorch/pytorch/issues/72359,9a87c2ccf9a7dae73db1c28540a016be42e64b6aeb207610a1722bb903555824 references,issue,75419,issue,49642,medium,issue.comments[0].body,See also #49642,https://github.com/pytorch/pytorch/issues/75419,5ab312354319d1caab1cebed11588813385cce945e35b8f6bd58aae4421c5aed references,issue,72112,issue,69033,medium,issue.body,"ith 1 in the graph before parsing. Is there a way around this issue, or another way to parse the generated IR? Note: @wconstab mentioned in #69033 (comment) that ""manually extracting the IR from lazy tracing and feeding it to a compiler isn't really the envisioned use case for...",https://github.com/pytorch/pytorch/issues/72112,e0e682ead538b0784c66ab66a8cbc5d1b5dc32d22105fb6bf2d94044e2e299fa references,issue,74148,issue,49716,medium,issue.comments[1].body,"GD is implemented in a future chain. I tried using two concurrent future chains when I developed this hook, but it didn't work it out. See: #49716",https://github.com/pytorch/pytorch/issues/74148,c7538f0fbe01f550f733f74233f9b0a189961ac1546d9b18eb2bad34ec92d991 references,issue,74771,issue,50034,medium,issue.comments[0].body,"related #50034, fixing this behavior without perf penalty is pretty hard.",https://github.com/pytorch/pytorch/issues/74771,b2c91b9563d507314733a8f647efb63fde47b8fada98996374eda6d619ec4027 references,issue,50345,issue,31945,medium,issue.body,(NOTE: We are not adding ops until we can reduce their impact on build time and CUDA context size.) Add iv Add ive Add kv Add logerfc (see #31945) Add logerfcx (see #31945) Add hyp2f1 Add betainc (being added in #58700) #74630 gammaincinv gammainccinv Bugs Probable mathematica...,https://github.com/pytorch/pytorch/issues/50345,3280ff2ff020ea95f02cda9803910b1af17651d240adfd10fa226abe741ef510 references,issue,50345,issue,55299,medium,issue.body,gammainccinv Bugs Probable mathematical error in torch.functional.kl_div() #57459 [feature] torch.polygamma : Support Tensor for argument n #55299 torch.polygamma inconsistent with scipy.special.polygamma for n >= 1 #55357 TODO Review Docs (especially logit) torch.special.expi...,https://github.com/pytorch/pytorch/issues/50345,86011e216a9550781eb0cb226b25b06e870fc563142d79ebc7e87f0177ca8d53 references,issue,41292,issue,12672,medium,issue.body,individual files becomes impractical and inefficient. This can be addressed using sequential storage formats and sharding. Related issues: #12672 #24985 #26957 #42405 Improved Sampling. Sampling is important component of data loading pipeline. Next requirements should be consi...,https://github.com/pytorch/pytorch/issues/41292,c26739d8d6a921f7421d512ac8b9864df70acea8e711210658ffab9132a2f28d references,issue,41292,issue,22924,medium,issue.body,ake DataLoader return readable and actionable exceptions. Make DataLoader return usable traces in the case of Ctrl+C and similar OS signals #22924. Issues with CPU Utilization. Usage of DataLoader frequently ends with oversubscribing to CPU threads or CPU underutilization. We...,https://github.com/pytorch/pytorch/issues/41292,cbc115a8d29966473b1d7bc6404bc412a1b67404c840186fac070c229e43dcd0 references,issue,41292,issue,24985,medium,issue.body,dual files becomes impractical and inefficient. This can be addressed using sequential storage formats and sharding. Related issues: #12672 #24985 #26957 #42405 Improved Sampling. Sampling is important component of data loading pipeline. Next requirements should be considered...,https://github.com/pytorch/pytorch/issues/41292,354532c736e3475e7d515499601f1075e5827c56184771edc13be3eabd3a2822 references,issue,41292,issue,25162,medium,issue.body,Sampling. Sampling is important component of data loading pipeline. Next requirements should be considered in case of architecture changes. #25162 Cascading Sampling Seq2Seq requirements #25743 #28743 Introduce ability to choose between multi-threading and multi-processing. Ri...,https://github.com/pytorch/pytorch/issues/41292,79e2ba0b1c2ffffbd19c42df42ca212fe265789453a35af8e6a2db60c44ea328 references,issue,41292,issue,25522,medium,issue.body,cases is overcomplicated due to cryptic Exception traces and/or inability to reproduce quickly. Examples #39570 #36375 #33296 #31758 #30147 #25522. We plan to: Implement proper timeouts handling and timeouts control. Make DataLoader return readable and actionable exceptions. M...,https://github.com/pytorch/pytorch/issues/41292,af3167404cf06049d9b137ff8f1ba18db0167e30a66a0f37ec7f05dd719615ce references,issue,41292,issue,25691,medium,issue.body,etween training epochs. Memory usage issues. In case of unbalanced loading/training speed DataLoader behavior leads to OOM. Examples #31101 #25691. We plan to: Add ability to control amount of pre-loaded data. Documentation Clear Initialization process doc. Some users getting...,https://github.com/pytorch/pytorch/issues/41292,65b6d298126d9268f12923cd0a71d10c5b99d520c32890c35d5823fe9af14744 references,issue,41292,issue,25743,medium,issue.body,ta loading pipeline. Next requirements should be considered in case of architecture changes. #25162 Cascading Sampling Seq2Seq requirements #25743 #28743 Introduce ability to choose between multi-threading and multi-processing. Right now we provide only multi-processing soluti...,https://github.com/pytorch/pytorch/issues/41292,07e428e9b287462617c7d707ba3437570f0b76d33cfac36be872d4c334c202cc references,issue,41292,issue,26957,medium,issue.body,les becomes impractical and inefficient. This can be addressed using sequential storage formats and sharding. Related issues: #12672 #24985 #26957 #42405 Improved Sampling. Sampling is important component of data loading pipeline. Next requirements should be considered in case...,https://github.com/pytorch/pytorch/issues/41292,fe722b1456793a8ba467a19424ddebd278554a97f97d328fdb30420e1613c38e references,issue,41292,issue,27617,medium,issue.body,ize of single epoch is too large. We can introduce method to save/restore data pipeline state. #36650. Improve collate_fn experience #33181 #27617 Unify Transforms Interface. Unlock ability to make JIT-able transforms Make it painless to make GPU transforms Make it possible to...,https://github.com/pytorch/pytorch/issues/41292,cc4ed6d4b7e91e5aecb63ccccd87bb0dae953887f67984a8c96d3282175be176 references,issue,41292,issue,28743,medium,issue.body,"ing pipeline. Next requirements should be considered in case of architecture changes. #25162 Cascading Sampling Seq2Seq requirements #25743 #28743 Introduce ability to choose between multi-threading and multi-processing. Right now we provide only multi-processing solution, mak...",https://github.com/pytorch/pytorch/issues/41292,79482cd7724d90f0dab49630d2aef2d21aa973cd6a0a4469f1f08172ca7acf0f references,issue,41292,issue,31758,medium,issue.body,lysis of such cases is overcomplicated due to cryptic Exception traces and/or inability to reproduce quickly. Examples #39570 #36375 #33296 #31758 #30147 #25522. We plan to: Implement proper timeouts handling and timeouts control. Make DataLoader return readable and actionable...,https://github.com/pytorch/pytorch/issues/41292,24dfc5114379fd1684e5e46fb597ef2ee57eca55f8ea0e3a9a208ae688f89bf1 references,issue,41292,issue,33181,medium,issue.body,when size of single epoch is too large. We can introduce method to save/restore data pipeline state. #36650. Improve collate_fn experience #33181 #27617 Unify Transforms Interface. Unlock ability to make JIT-able transforms Make it painless to make GPU transforms Make it possi...,https://github.com/pytorch/pytorch/issues/41292,9bccd5b7bbe6331734157cc7d4938ec619f1d2a7b6e3ed279a6193a1e18180ad references,issue,41292,issue,33296,medium,issue.body,use analysis of such cases is overcomplicated due to cryptic Exception traces and/or inability to reproduce quickly. Examples #39570 #36375 #33296 #31758 #30147 #25522. We plan to: Implement proper timeouts handling and timeouts control. Make DataLoader return readable and act...,https://github.com/pytorch/pytorch/issues/41292,f815b5527e379798daec4f48180096971570ca05f1681a998870bc1e7c4904ac references,issue,41292,issue,35642,medium,issue.body,e cases caching is the easiest way to store entire DataSet in the memory. We can unify caching API to avoid multiple non optimal solutions. #35642 #39274 State saving / restoration for DataSet / DataLoader / Sampler. Restarting training from specific checkpoint is problematic...,https://github.com/pytorch/pytorch/issues/41292,83956b7c523f27899a3c95d464f4ad0ad32e0144ecb64659b273676b151772a6 references,issue,41292,issue,35759,medium,issue.body,f used threads/processes. GPU underutilization. Multiple users reported problems with the CUDA context and multiprocessing. Examples #40403 #35759. We plan to: Review the process of spawning processes to make sure it is easy/understandable to use. We plan to: Unlock the abilit...,https://github.com/pytorch/pytorch/issues/41292,4afb71a6333f2d9468af3fb3b3ec6d6e4f6934d363317cf1e5b9a6f65ca7759d references,issue,41292,issue,36650,medium,issue.body,rom specific checkpoint is problematic when size of single epoch is too large. We can introduce method to save/restore data pipeline state. #36650. Improve collate_fn experience #33181 #27617 Unify Transforms Interface. Unlock ability to make JIT-able transforms Make it painle...,https://github.com/pytorch/pytorch/issues/41292,7a63792969741adf5c2f7dc491a0ff32068ea78b0ae48801c4cabb5a7ccc0673 references,issue,41292,issue,38419,medium,issue.body,"rious sizes of batch (1-N) on every step of the process. We need to make it possible to specify it on every level of loading graph. WebData #38419. As datasets become larger and larger, storing training samples as individual files becomes impractical and inefficient. This can...",https://github.com/pytorch/pytorch/issues/41292,c0225f462990490be3ef1e0f393a9ac2903223f4cff255dbc757c59b5a3fdc43 references,issue,41292,issue,39570,medium,issue.body,eason. Root cause analysis of such cases is overcomplicated due to cryptic Exception traces and/or inability to reproduce quickly. Examples #39570 #36375 #33296 #31758 #30147 #25522. We plan to: Implement proper timeouts handling and timeouts control. Make DataLoader return re...,https://github.com/pytorch/pytorch/issues/41292,6cda902cf4c088f13e12e972566a76911bdd3f31621a4981b43cd3359a29aa50 references,issue,41292,issue,42405,medium,issue.body,omes impractical and inefficient. This can be addressed using sequential storage formats and sharding. Related issues: #12672 #24985 #26957 #42405 Improved Sampling. Sampling is important component of data loading pipeline. Next requirements should be considered in case of arc...,https://github.com/pytorch/pytorch/issues/41292,f2082c7900234343b618dc003d15c084755adc5ea6aa80f17fdef73e979b4bd8 references,issue,36108,issue,24870,medium,issue.body,This issue is expanded from #24870 for reference and for additional discussions on implementation details. See that issue for context. Add a function flow_sample to torch.nn.,https://github.com/pytorch/pytorch/issues/36108,3d93a7a1b6f8f439aac0af5c0ff73c7af2dc3853826f15105aa5f46707ae85b9 references,issue,36108,issue,36107,medium,issue.body,"d_sample will be nearly identical, and both will dispatch to the same underlying kernel (the existing grid_sample kernel). This, along with #36107, would help simplify the implementation of convolutional networks to do Optical Flow Estimation and Image Registration tasks.",https://github.com/pytorch/pytorch/issues/36108,4adf1cd8c14f6a3b87bb3c19bfe73836b06900fa7b724d4e0fe3c493969bb788 references,issue,38228,issue,22924,medium,issue.body,+ exceptions from destructors. This triggers the call to std::terminate and the abort. Here is some code that reproduces the issue. I think #22924 is essentially the same issue. I've also seen this happen in research code that mixes C++ threads with Python. import torch import...,https://github.com/pytorch/pytorch/issues/38228,56c36cdbf1c1e11c8e89963d9c67c781c55a3064ef8c5306682f95790a45511f references,issue,73710,issue,72960,medium,issue.body,"acking issue for refactor/cleanup wishlists. Please contribute ideas/desires as a new issue or add to existing issues below! Codegen Infra: #72960 LazyTensor, LTCTensorImpl, Data: #73711 Shape Inference stuff: TODO open tracking issue move kNullValue out of LazyIr.h, as it's b...",https://github.com/pytorch/pytorch/issues/73710,93d3e9d0095baf169504964633c91b339fc8c7230807e0f373324b295c433c6d references,issue,73710,issue,73711,medium,issue.body,"sts. Please contribute ideas/desires as a new issue or add to existing issues below! Codegen Infra: #72960 LazyTensor, LTCTensorImpl, Data: #73711 Shape Inference stuff: TODO open tracking issue move kNullValue out of LazyIr.h, as it's bad form to use a static in a header Comp...",https://github.com/pytorch/pytorch/issues/73710,ab147f4a263aa67f12d2fdb918f3071f58ef014eb2e11ef2ec1ca3f941ee86f3 references,issue,73870,issue,73190,medium,issue.body,"stdin>"", line 1, in RuntimeError: max_pool1d() Invalid computed output size: Note This bug was inspired by the high-priority issue #73190 which used much larger inputs. Versions PyTorch version: 1.10.1+cu102 Is debug build: False CUDA used to build PyTorch: 10.2 ROCM...",https://github.com/pytorch/pytorch/issues/73870,462f84d829bd8987700f45f6bbded7ddeb8ea356338334ff8a3432ea80c80563 references,issue,45242,issue,39279,medium,issue.body,"ed in the context of distributed work in #44715 #44791 (cc @jbschlosser @vincentqb @albanD @wanchaol, internal doc) and of meta-learning in #39279 (cc @egrefen @seba-1511), we want to refactor the optimizer in order to provide a functional form to them (with no other algorithm...",https://github.com/pytorch/pytorch/issues/45242,917b1a36f4377562e042176a90dd2279b1e797b46a781bd0e2ff90a814c87131 references,issue,58111,issue,62998,medium,issue.comments[1].body,"ve a boolean to ensure this lambda is run only once. This could be a temporary solution until the below is possible. (better option) - Once #62998 is resolved, we will have a module-level pre-backward hook which we can then use in DDP, and also deprecate DDPSink. cc @zhaojuanmao",https://github.com/pytorch/pytorch/issues/58111,097025a017224c3b23016afd0af71ce8f2b47fa6641e513806a33ec864f39093 references,issue,58111,issue,70865,medium,issue.comments[1].body,"This is also contributing to #70865, as when we run with static_graph, num_iterations accounting is buggy with multiple forward pass. There are 2 options: (hacky) - Using a si",https://github.com/pytorch/pytorch/issues/58111,a392b502206652a79c66cefee3052feb49ad16ead3cf70e898710432d69055cc references,issue,73640,issue,65049,medium,issue.comments[0].body,Thanks for the request! Is this a duplicate of #65049? There was a PR #65139 to address this but it stalled out. I believe the problem was that TorchScript doesn't support the Generator type. c,https://github.com/pytorch/pytorch/issues/73640,b7c7758b9809337214367add06b5527078d46e0abd4c595d9f75f611e9dfe090 references,issue,60858,issue,51720,medium,issue.body,"untime. Currently, PyTorch uses a 32-bit MKL version, linking with the 64-bit (ILP64) might break the internal code that relies on MKL (see #51720, #56959), but it was not tested. Include the recommendation in the documentation that on CPU 32-bit indices are preferred to avoid...",https://github.com/pytorch/pytorch/issues/60858,1faab651886ca595b08968e5f26d4526d0c8d8e273aae5e915db2a94b0f3f6c8 references,issue,60858,issue,58770,medium,issue.body,"d by priority): mkl_sparse_?_mm (sparse matrix - dense matrix multiplication, dense result) Implement descriptor wrappers for CSR matrices. #58770 Use as a backend for: torch.mm torch.addmm torch.baddmm (if batched CSR is enabled in PyTorch) torch._sparse_mm PR in progress: #6...",https://github.com/pytorch/pytorch/issues/60858,2e4e573cda07748bb6159fc4701e708ada8b87e185fffd955ac073a7bb181eca references,issue,70533,issue,27522,medium,issue.comments[0].body,regular sum also seems not very nice in terms of computational complexity (unless JIT can optimize it). A related feature request: #27522,https://github.com/pytorch/pytorch/issues/70533,0f8725a7a0346cb475c915c0fec1aa17f82d29ad4dba415e50aeac6a6d31483f references,issue,29042,issue,28417,medium,issue.body,ature request is to officially support building CUDA code with clang and include this build in CI. This build is not officially supported. (#28417 (comment)). There has also been related bug reports (e.g. #28417). Motivation Due to the lack of C API and the instability of C++...,https://github.com/pytorch/pytorch/issues/29042,c4c19a4b08a9dc2129276fb0a873b48e853965bb6b5f426994b04edd52edf6c7 references,issue,67712,issue,67448,medium,issue.body,", like 5min or something. Motivation This would help save resources and help us pinpoint culprit tests when they time out, like the ones in #67448 cc @seemethere @malfet @pytorch/pytorch-dev-infra",https://github.com/pytorch/pytorch/issues/67712,70d52f79f615fc9062b7b0b50c8eada582311478b0ad18de2ab6556117005701 references,issue,73412,issue,73113,medium,issue.comments[0].body,Similar to #73113,https://github.com/pytorch/pytorch/issues/73412,f4b70250fa802676031888c9d369e518c17e367031d166b9da4d9f26a940d434 references,issue,70350,issue,64812,medium,issue.comments[0].body,ing the pytorch api at runtime using https://github.com/xloem/mempickle/blob/master/patch_pytorch.py . This approach may work to workaround #64812 and #68349 which are likely the same issue.,https://github.com/pytorch/pytorch/issues/70350,d524d10efccae9248a5378d240d94e9bfa2748e508598cef96b74fcaa0f4996b references,issue,70350,issue,68349,medium,issue.comments[0].body,orch api at runtime using https://github.com/xloem/mempickle/blob/master/patch_pytorch.py . This approach may work to workaround #64812 and #68349 which are likely the same issue.,https://github.com/pytorch/pytorch/issues/70350,0b1dacb90c46f4024c723ddf490e7fbfc1c0297642b345b01cba292607ed9b54 references,issue,70350,issue,64812,medium,issue.comments[1].body,"python3 -c ""import torch;print(torch.mm(torch.zeros((1,1024)), torch.zeros((1024,1024))))"" ... run ... bt ... disasm Likely a duplicate of #64812 a should be fixable by building OpenBLAS with a dynamic backend",https://github.com/pytorch/pytorch/issues/70350,0b4caf8fe4f6222122a20b06d957ab7f6305cc3f6e3158399347543942d1ec63 references,issue,71843,issue,62540,medium,issue.body,Please look at the parent issue for instructions.(#62540),https://github.com/pytorch/pytorch/issues/71843,9c148419a9d90fcca1c347427c27a6424059be0adeb3621cdeaad748a6adc450 references,issue,72911,issue,69991,medium,issue.comments[0].body,"Thanks for the bug report, @cw-tan. Notes to maintainers: this is probably a composite compliance problem: #69991",https://github.com/pytorch/pytorch/issues/72911,4d542ed9f3de70ace7b3d442154984b528c822197c614b7c47bf980b4dc946de references,issue,72824,issue,56356,medium,issue.comments[1].body,"y @TestSomething22, thanks for the report! Type promotion support in losses does look to be a bit arbitrary. According to the discussion in #56356, we ideally want to support type promotion only for unary pwise, binary pwise, and reduction ops. Since it's common for losses to...",https://github.com/pytorch/pytorch/issues/72824,57a6318baffa61de320372997729500cb2d779d35b9ac2ad6d9b8a12662b71c6 references,issue,60234,issue,58128,medium,issue.comments[0].body,"I didn't look closely into this, but it's probably something like #58128 -- it's not a bug.",https://github.com/pytorch/pytorch/issues/60234,e351de5ff02ffc0610cd7752245452f7270e44e944a7be1873ff049bac035cc3 references,issue,60234,issue,58128,medium,issue.comments[1].body,"I didn't look closely into this, but it's probably something like #58128 -- it's not a bug. Thanks for your reply, I have changed the title to numerical-reproducibility issue but I don't think it is the same as #",https://github.com/pytorch/pytorch/issues/60234,5a6c90cb09da3cd1744ed42965ac12026ac792a12fe4eccc6e7955ffd891014b references,issue,69960,issue,69801,medium,issue.body,🐛 Describe the bug Related #69801 FeatureAlphaDropout also has similar problem. import torch import torch.nn as nn from torch.testing._internal.common_utils import freeze_rn,https://github.com/pytorch/pytorch/issues/69960,f40e520765ad9b33081b8eb53ed8849bf220743ce27b1f9108a88afdd95c3f57 references,issue,69972,issue,41081,medium,issue.body,avior Option 1: torchscript ignores the wrapped module Option 2: The user is expected to unwrap the batchnorm layers. Feature request here: #41081 Additional Context This issue was reported in PyTorch Lightning: Lightning-AI/pytorch-lightning#11076 A related feature request: #...,https://github.com/pytorch/pytorch/issues/69972,a7ced12b074cbee804c6d88077fc7de4809388d633b819b6404caa6eed57f748 references,issue,69972,issue,41353,medium,issue.body,re: #41081 Additional Context This issue was reported in PyTorch Lightning: Lightning-AI/pytorch-lightning#11076 A related feature request: #41353 Versions Collecting environment information... PyTorch version: 1.11.0a0+gita31aea8 Is debug build: False CUDA used to build PyTor...,https://github.com/pytorch/pytorch/issues/69972,4ead2fc93befa6dcec4cd388501715672a119ef776617620ce6acba15f076eb5 references,issue,41499,issue,6662,medium,issue.body,"📚 Documentation Hello, I noticed this older issue #6662 is still open and looked through this PR #24435 about adding doctest to jit. If it would be helpful I can work on other parts of the docs t",https://github.com/pytorch/pytorch/issues/41499,8a63402e7f0b1cd45447f292bd93b4109bedfee101aef1e19ab5229b7addb69e references,issue,70678,issue,70242,medium,issue.body,"1) is not desirable for obvious reasons. (2) seems easier than (3), but I'm not sure about what performance issues we would have. See also: #70242 #70433 #65993 @ezyang @bdhirsh @lezcano",https://github.com/pytorch/pytorch/issues/70678,660b1a5a94dd9e1963fd43613f3b61a216d336ae48322bdfed4f21113c89fbed references,issue,72258,issue,29981,medium,issue.comments[1].body,"feature request rather than a bug. There are some interface questions about what cdf should mean in multiple dimensions; see discussion in #29981 For examples of the challenging math to implement this, see @martinjankowiak's GaussianScaleMixture distribution in Pyro",https://github.com/pytorch/pytorch/issues/72258,dd1eb8386367b8dbfbd0234a48676810e611b3bd59c0e1f094f7dfcf7719b2c5 references,issue,72558,issue,72556,medium,issue.body,Followup to #72432: Expand current linter that checks all installs are pinned to also check other Would be easier to do AFTER #72556! Pin all pip installs in our ci scripts Expand the current existing lint in Lint.yml to cover more than the .github directory Modify the li,https://github.com/pytorch/pytorch/issues/72558,3dda598ea795799d34afc759cfc5ecdb5a13c008cae9cabfc7e9a0d1010c999f references,issue,4358,issue,3619,medium,issue.comments[1].body,"something similar was reported before in #3619 (comment) It's MKL, it is not being fork-safe for some reason, not sure why. When import torch is done, I guess some MKL state is being ini",https://github.com/pytorch/pytorch/issues/4358,afe6f50f270fee3a3f3417a14c96d9cf25a3629044426b8dfb676ae9e64d2c33 references,issue,69744,issue,69741,medium,issue.body,"uts[0] torch.autograd.functional.jacobian( func=wrapper inputs=(all_inputs[1],) ) But this has no hope of being TorchScript compatible (see #69741) and is also pretty ugly. Additional context No response cc @ezyang @albanD @zou3519 @gqchen @pearu @nikitaved @soulitzer @lezcano...",https://github.com/pytorch/pytorch/issues/69744,c9f7a24fa3cece1e0b98695582c89d5ac9f0cfa6243b8c4cd813fa38db62e69f references,issue,72341,issue,71465,medium,issue.comments[0].body,"This is a duplicate of #67976? Currently norms (except batchNorm) don't have memory-format-preserving implementations, see also #71465",https://github.com/pytorch/pytorch/issues/72341,f0e0b04b9c783935757b5df0fd3cea3fe3f2dba9203b37cd391abd77dadc74fa references,issue,72361,issue,69687,medium,issue.body,see #69687 example PR: #71129 Benchmark the forward mode AD perf for your change!,https://github.com/pytorch/pytorch/issues/72361,44a45c17fde080055ffcef1acc49e5363e4e93fae6de8342d437476c82193368 references,issue,72358,issue,69687,medium,issue.body,see #69687 example PR: #71129 Benchmark the forward mode AD perf for your change!,https://github.com/pytorch/pytorch/issues/72358,87e55e53b060ef8529754591dbdcb2408b3b5740dacc24b4059569cfd2a92cfd references,issue,72347,issue,69687,medium,issue.body,see #69687 example PR: #71129 Benchmark the forward mode AD perf for your change!,https://github.com/pytorch/pytorch/issues/72347,510c97759808f78169201c146fff888f9960e41910f7badd09124c18a0f3d8e8 references,issue,62812,issue,58742,medium,issue.body,"g.einsum among other linear algebra functions. Currently, PyTorch supports the same functionality with torch.einsum. Array-API Ops tracker (#58742) lists a few aliases and it would be nice to add an alias torch.linalg.einsum to torch.einsum making it compliant with Array API....",https://github.com/pytorch/pytorch/issues/62812,29aa302fd7eb0e803e16aa5e6129149e1ac09f631377cfba1d5536f1b4107411 references,issue,71359,issue,48932,medium,issue.body,"🚀 The feature, motivation and pitch I ran into #48932 with a slightly different model getting exported via torch.jit, but ultimately the fix is straightforward -- install libtorchvision, find i",https://github.com/pytorch/pytorch/issues/71359,104e370ea26853debb3123d4bbee4e9a5fd7a72527708883b07f2b317bb2cdf9 references,issue,71911,issue,43567,medium,issue.body,"nctions, is it an error or inf/nan? Tests that are currently disabled: test_norm_extreme_values (when ord='nuc' or 2 or -2) Related issues: #43567 #52633 Versions Master branch with either MKL 2022+ or OpenBLAS 0.3.15+ cc @jianyuh @nikitaved @pearu @mruberry @walterddr @IvanYa...",https://github.com/pytorch/pytorch/issues/71911,199c340be707853cca1c6323597c0bd80cbe5ace6cbb56283f1fe3d4969ca795 references,issue,71911,issue,52633,medium,issue.body,", is it an error or inf/nan? Tests that are currently disabled: test_norm_extreme_values (when ord='nuc' or 2 or -2) Related issues: #43567 #52633 Versions Master branch with either MKL 2022+ or OpenBLAS 0.3.15+ cc @jianyuh @nikitaved @pearu @mruberry @walterddr @IvanYashchuk...",https://github.com/pytorch/pytorch/issues/71911,eacd6f6a31f4b52ad484da9d5030ee5226aebaeb8e9f9214292a40f484f6edb3 references,issue,71842,issue,62540,medium,issue.body,Please look at the parent issue for instructions.(#62540),https://github.com/pytorch/pytorch/issues/71842,3704cc45a4712d9d7a5a0f62efde340c7c8a583178a212cdbc2754d7a6d68a58 references,issue,71841,issue,62540,medium,issue.body,Please look at the parent issue for instructions.(#62540),https://github.com/pytorch/pytorch/issues/71841,0bb16db92d6b261f3e3bdfcc0e2bf376f07512f2b92113f67df0da1c564b0805 references,issue,71815,issue,62540,medium,issue.body,Please look at the parent issue for instructions.(#62540),https://github.com/pytorch/pytorch/issues/71815,96d238c6f343790781322b1b8701d261443151c33c518bfa15a5bf5f43ef454e references,issue,71814,issue,62540,medium,issue.body,Please look at the parent issue for instructions.(#62540),https://github.com/pytorch/pytorch/issues/71814,e43b44939838637ebea3c89338ea32dddd6759a72ae757c021a272a40aa19136 references,issue,71813,issue,62540,medium,issue.body,Please look at the parent issue for instructions.(#62540),https://github.com/pytorch/pytorch/issues/71813,05b3ca894181b3892fa116fe8f8cdf7bc6e55669c419951cd38a3dbb9bbe2c52 references,issue,71812,issue,62540,medium,issue.body,Please look at the parent issue for instructions.(#62540),https://github.com/pytorch/pytorch/issues/71812,a296c932caa27c13f308dae976128267e761ee13fc4ecd7a16b37e228686a96b references,issue,71810,issue,62540,medium,issue.body,Please look at the parent issue for instructions.(#62540),https://github.com/pytorch/pytorch/issues/71810,d1ca21598bdc720fb05464a02aa881df08a8188c3f14cda2e1e49b67d5baa7bb references,issue,71840,issue,62540,medium,issue.body,Please look at the parent issue for instructions.(#62540),https://github.com/pytorch/pytorch/issues/71840,8435a691f1b7570cbc8fc710e50897bc89528f1213bdadae1ce755dd62ead220 references,issue,71839,issue,62540,medium,issue.body,Please look at the parent issue for instructions.(#62540),https://github.com/pytorch/pytorch/issues/71839,9508db3d0b63439731c28fee0e3339f7ae5cb91a3bf72d26db2cfec64bcc91ab references,issue,71838,issue,62540,medium,issue.body,Please look at the parent issue for instructions.(#62540),https://github.com/pytorch/pytorch/issues/71838,0b7345dc4522dc8f79099ea5461c7f61aa1fd178464bfa1be19326221fc8f737 references,issue,71837,issue,62540,medium,issue.body,Please look at the parent issue for instructions.(#62540),https://github.com/pytorch/pytorch/issues/71837,8e28fb7969181e8e008a7604a9cd3b8bb1a7a02085d95719d24cac2e2a5bcfd8 references,issue,71836,issue,62540,medium,issue.body,Please look at the parent issue for instructions.(#62540),https://github.com/pytorch/pytorch/issues/71836,ad5e04feec1a6a459935adf0cec000590998808aa6a5a186138c2c48b07e0e16 references,issue,71835,issue,62540,medium,issue.body,Please look at the parent issue for instructions.(#62540),https://github.com/pytorch/pytorch/issues/71835,91010246bcc9ef5558c14d5c9cc6b65379d8469f885e1fbaeca367d8fe5052ef references,issue,71834,issue,62540,medium,issue.body,Please look at the parent issue for instructions.(#62540),https://github.com/pytorch/pytorch/issues/71834,d0b5692d030cd9f90c74e4272253d7404ec170ba67565d5db7503d430a0a35c7 references,issue,71833,issue,62540,medium,issue.body,Please look at the parent issue for instructions.(#62540),https://github.com/pytorch/pytorch/issues/71833,1fe4cfe331a12bef6110f50d33c1e54b15bf1ef1664192488561190cd1d24252 references,issue,71832,issue,62540,medium,issue.body,Please look at the parent issue for instructions.(#62540),https://github.com/pytorch/pytorch/issues/71832,aa677b29b268df764b2a02820a68d40f2874f1a61994b199fac15a6f2b26c462 references,issue,71831,issue,62540,medium,issue.body,Please look at the parent issue for instructions.(#62540),https://github.com/pytorch/pytorch/issues/71831,44554a739c866cb1e85feef6f498335a6ee3ee13a51eecc982a0ed8ad0154887 references,issue,71830,issue,62540,medium,issue.body,Please look at the parent issue for instructions.(#62540),https://github.com/pytorch/pytorch/issues/71830,17966667700366e6528f51a03f678ee6aecc34ffcc9ad36274c3698475dc92e2 references,issue,71828,issue,62540,medium,issue.body,Please look at the parent issue for instructions.(#62540),https://github.com/pytorch/pytorch/issues/71828,1128a82a6d5ef77ff50c28658cb13837b087101611c58097fd7244b20654d382 references,issue,71826,issue,62540,medium,issue.body,Please look at the parent issue for instructions.(#62540),https://github.com/pytorch/pytorch/issues/71826,cd222f540e29a94087fd11b95bccac14f268afb504b775d0c6f3e179a07c1bb1 references,issue,71825,issue,62540,medium,issue.body,Please look at the parent issue for instructions.(#62540),https://github.com/pytorch/pytorch/issues/71825,b8d9211643b611143330aaf2523b008ad0b7601f9d66cf36d368c1ffeaf97fb8 references,issue,71824,issue,62540,medium,issue.body,Please look at the parent issue for instructions.(#62540),https://github.com/pytorch/pytorch/issues/71824,5ae533229eba019e55d7f1f4d60bc0754338cefbc80bd32d09c12fa96456acdc references,issue,71823,issue,62540,medium,issue.body,Please look at the parent issue for instructions.(#62540),https://github.com/pytorch/pytorch/issues/71823,431c3f99a9be779556bdd32e10be236f157395a054c588131d5f58ceeca8a2fb references,issue,71822,issue,62540,medium,issue.body,Please look at the parent issue for instructions.(#62540),https://github.com/pytorch/pytorch/issues/71822,683a8e06aa411dc5474a1897ecf1e89df709947c5e4ccda380525bc8ef3c522a references,issue,71821,issue,62540,medium,issue.body,Please look at the parent issue for instructions.(#62540),https://github.com/pytorch/pytorch/issues/71821,54e009950d5a9a942bd9c4d08537ad24fe28f5fdb4247aa1c56094266ab6054c references,issue,71820,issue,62540,medium,issue.body,Please look at the parent issue for instructions.(#62540),https://github.com/pytorch/pytorch/issues/71820,14df0c4c314622fdd34bdef830be3040dc2471ea9f9745bf926abaa4d99edbd9 references,issue,71819,issue,62540,medium,issue.body,Please look at the parent issue for instructions.(#62540),https://github.com/pytorch/pytorch/issues/71819,6bb914935c317a5eb642525b5735fa6a610f5de3c1578716e644a8067d6e1c38 references,issue,71818,issue,62540,medium,issue.body,Please look at the parent issue for instructions.(#62540),https://github.com/pytorch/pytorch/issues/71818,763d609a65c1788fca5c975c4523f09ba99550eccbe775b7ff16ebc14f175e48 references,issue,71817,issue,62540,medium,issue.body,Please look at the parent issue for instructions.(#62540),https://github.com/pytorch/pytorch/issues/71817,8999cf8c8bdf0a9bdf4db51af899ca02a152c4382242f75f9857a8cf54d78028 references,issue,71816,issue,62540,medium,issue.body,Please look at the parent issue for instructions.(#62540),https://github.com/pytorch/pytorch/issues/71816,fcdd295379f736359ad0c283ba64a502b9f43ad7e6df18c64c9863dc45b877b5 references,issue,56126,issue,24834,medium,issue.comments[1].body,I met a similar issue on PyTorch 1.8.1. A comment in #24834 was helpful for me. This issue seems duplicate.,https://github.com/pytorch/pytorch/issues/56126,8921b6b29ee5e47c452bc5cc1ed9024a8cf8af76a41972a1b6062e2126f6da4b references,issue,70559,issue,62396,medium,issue.comments[1].body,2.5 3.3 4. ] Comparision with scikit-image is not exact as they rescale as if recompute_scale_factor=True. I think this issue is related to #62396 where pytorch computes the output size in a bit different way vs opencv or scikit-image and thus rescaling tensor with size 3 is i...,https://github.com/pytorch/pytorch/issues/70559,e5b086e0b8b590162aca031f3af6da915b52aadac36882c07849dea526473ce8 references,issue,71595,issue,67570,medium,issue.body,"🚀 The feature, motivation and pitch A process to overlap optimizer + allreduce in DDP is being implemented as described in #67570, although as a prototype feature there is limited support for some use cases. This issue exists to maintain a list of currently missing fea",https://github.com/pytorch/pytorch/issues/71595,6130b0a8467230c3a4eac8185bcce8358444555bc36bb6f67cbf7db6f6763625 references,issue,71209,issue,70914,medium,issue.body,"name dim, bool keepdim, *, Tensor out) The same applies to the following other ops: any max mean prod std sum var This issue is related to #70914, which likely uses the same underlying functionality to parse inputs. cc @mruberry @rgommers @pmeier @asmeurer @leofang @AnirudhDag...",https://github.com/pytorch/pytorch/issues/71209,7e1726864470aefc77a5923ae7961415593e29b5f81eb3ac161376bac75a5ee7 references,issue,71548,issue,31772,medium,issue.body,"chance to implement this feature soon? :) The interface trick and the NormalizingFlow usecase are already mentioned in #47496 #53410 #43737 #31772 (comment) Alternatives AFAIK there is no way how to iterate over ModuleList in reverse without the interface trick, but the freezi...",https://github.com/pytorch/pytorch/issues/71548,db21eaa6d0d5796f6bb3fa17996a0ad660193e9c5eafdbe2ac378fec579b88e5 references,issue,53712,issue,62475,medium,issue.comments[1].body,"same here. any workaround? related issues: #62475, #20997, Lightning-AI/pytorch-lightning#8720",https://github.com/pytorch/pytorch/issues/53712,aa23724c3e065ff364c350c6ffc67e625fdbd3b224f3f92b2aa6053358bc7619 references,issue,70191,issue,60341,medium,issue.body,"c1010TensorImpl23shallow_copy_and_detachERKNS_15VariableVersionEb Looking at the already existing issues I have found something similar in: #60341 #13541 Versions For every version I have tried: 1.7.0, 1.7.1, 1.8.1, 1.9.0 and 1.9.1 cc @ezyang @seemethere @malfet",https://github.com/pytorch/pytorch/issues/70191,2dd1434b30d0b6b18e495c4a90a9b09ed895adfa160a41dd62edca88155600ae references,issue,70931,issue,67153,medium,issue.comments[0].body,"Pytorch 1.10 doesn't support cuda 11.5, that support was added later. Duplicate of #67153",https://github.com/pytorch/pytorch/issues/70931,726b113ff3552cfaabd1c6d341ddc1f83acb45e149ac2c9ffb929ffed7338bf7 references,issue,59566,issue,49992,medium,issue.body,"ify this relationship in the docs? Perhaps there should be a section pointing to docs for the wrapped functional form, if there is one. See #49992 Anything missing above that should be in the docs? cc @brianjo @mruberry @albanD @jbschlosser",https://github.com/pytorch/pytorch/issues/59566,131f8e5b9262a25617652df1d775c07c0b597bf43e46a75e6f19b0794421788c references,issue,69688,issue,58466,medium,issue.body,te the docs to reflect that torch.cuda.set_per_process_memory_fraction() does not consider overheads by e.g. CUDA context (refer #48172 and #58466) Versions Collecting environment information... PyTorch version: 1.8.2+cu111 Is debug build: False CUDA used to build PyTorch: 11....,https://github.com/pytorch/pytorch/issues/69688,532174681f02161fef350d2c959481815aca84c94ce042b27b2b81405b5a4d9e references,issue,64023,issue,61523,medium,issue.comments[0].body,"confusing which overloads are used for particular operations. Another example issue, where unexpected intermediate truncation leads to nans #61523",https://github.com/pytorch/pytorch/issues/64023,d7cfa614492466c70be0aec17371ba2bb9829a64a7aa09760159da51775ecb47 references,issue,70241,issue,56794,medium,issue.comments[1].body,Probably #56794,https://github.com/pytorch/pytorch/issues/70241,cf9588eafdb5de40560470fbcf7c4db35362f2ede498219ae5bb6048fdbb0890 references,issue,66073,issue,41243,medium,issue.body,ass _BatchNorm and a variable is_frozen or frozen on the module instance Related request to keep only BatchNorm class and drop BatchNorm*d: #41243 cc @albanD @mruberry @jbschlosser @walterddr,https://github.com/pytorch/pytorch/issues/66073,6cd4fc7f70bd8eabf646539561d6a63393714a471b90ecc00d5ed5fac3fbb4c0 references,issue,69892,issue,69822,medium,issue.body,"ime ago, it wasn't supported back then, but now NumPy also supports it :) So maybe it's a stronger argument now. Most recently proposed in: #69822 (comment) Alternatives No response Additional context No response cc @mruberry @rgommers @gchanan",https://github.com/pytorch/pytorch/issues/69892,b9d194259e8934588bcd73e904b62e37c9979d20f2b0e537cd7aeee147a81666 references,issue,69386,issue,20323,medium,issue.comments[1].body,"normal_ variant. The first two were added to support the size parameter (to behave similar to numpy) as mentioned in: xzhu1900@bfc4eb6 and #20323. The inplace variant uses the Meta keyword, which shouldn't really be used. Not sure why the former of the first two variants uses...",https://github.com/pytorch/pytorch/issues/69386,f0e975f51f73e373da8056cb416629bc3cc1d6da1b099abee077cb7b06bfbc2c references,issue,68979,issue,67760,medium,issue.body,"ers. Alternatives Keep as is and only allow scheduler that inherit from _LRScheduler. However then I propose to make this class public, see #67760 cc @vincentqb @jbschlosser @albanD",https://github.com/pytorch/pytorch/issues/68979,524435923a196a4772793efe2bb250bc4eea24c38b90f827c6d7b999c98d7a1b references,issue,68979,issue,68978,medium,issue.body,🚀 Feature SequentialLR should allow for optional arguments in step(). This is a follow-up to #68978 which only requests support for ReduceLROnPlateau while this issue requests for a broader support of arbitrary (custom) schedulers. Motivat,https://github.com/pytorch/pytorch/issues/68979,0846b8259cbbac5179832d150bd05509d4f01327583cbe94a9978b6a344b953a references,issue,69467,issue,52973,medium,issue.body,on pytorch.distributions.Distribution is fantastic but it lacks a lot of practical distribution functionality. Additional context See also: #52973 cc @fritzo @neerajprad @alicanb @nikitaved,https://github.com/pytorch/pytorch/issues/69467,adb1541ece5b1464d5ca09ec6fdd0ebb00cba6f5a8c99d60cf3e812ec45a213f references,issue,69469,issue,52973,medium,issue.body,on pytorch.distributions.Distribution is fantastic but it lacks a lot of practical distribution functionality. Additional context See also: #52973,https://github.com/pytorch/pytorch/issues/69469,0a868e5f1af0b19c3e6cb1d2f8f7e12754ec9b07709db7ec867140f7056f1376 references,issue,68648,issue,66073,medium,issue.body,"currently do: self.model_with_dropouts.eval() self.model_with_dropouts.train = lambda _: None (same usecase for BatchNorm - ""frozen mode"", #66073) Instead maybe one could use self.model_with_dropouts.eval(force = True), and then it would not change mode unless the upper-level...",https://github.com/pytorch/pytorch/issues/68648,3c36f8bca9ea0fcf1d26091ebacd3ef527f930c68941032339236b7ffaa78049 references,issue,68288,issue,68287,medium,issue.body,"Add autograd wrappers for ops that use sizes in their derivatives: #68287 Add autograd wrappers for ops that take sizes as their arguments in forward. #68298 Support zeros and zeros_like in Lazy TS backend, as aut",https://github.com/pytorch/pytorch/issues/68288,fc689b2f691c94909b8ad8f4b52319a5b479e1135def5cafda5a92ce517db2e5 references,issue,68288,issue,68298,medium,issue.body,"d wrappers for ops that use sizes in their derivatives: #68287 Add autograd wrappers for ops that take sizes as their arguments in forward. #68298 Support zeros and zeros_like in Lazy TS backend, as autograd calls zeros_like and zeros : pytorch/torch/csrc/autograd/custom_funct...",https://github.com/pytorch/pytorch/issues/68288,57d821a2a3968c96546296ea51e08895e5eda12f6e2d23d86f8ede9a4d081d88 references,issue,68105,issue,56356,medium,issue.comments[1].body,"Type promotion for matmul operations is currently not supported #56356. Note that even if torch.mm(A, B, out=result) with the different types was supported, it would not be differentiable (functions with out kw",https://github.com/pytorch/pytorch/issues/68105,242c84e5e8d47cddf31e827a922e73dda0fe03e3a9aba88ab492dc49dfca36b2 references,issue,53879,issue,47953,medium,issue.body,"* N]); } printf(""\n""); } } See also Linear Algebra tracking issue #42666 Linear algebra GPU backend tracking issue [magma/cusolver/cublas] #47953 cc @ngimel @jianyuh @nikitaved @pearu @mruberry @heitorschueroff @walterddr @IvanYashchuk @ptrblck",https://github.com/pytorch/pytorch/issues/53879,5ebc778d10b70cbb25245384503f7f394fec5eb17c19730cbb8201e20b28ffb2 references,issue,32365,issue,32300,medium,issue.body,built-in class. Environment Google Colab with a fresh GPU runtime on which !pip install -q torch==1.4.0 has been run. This is a spin off of #32300 at @driazati's request. cc @suo,https://github.com/pytorch/pytorch/issues/32365,46aeff7917fa3ab2cf01e68e4b468926a8996a98909cf3ab7002700c2dbbb3ef references,issue,32300,issue,32365,medium,issue.comments[1].body,"Done. #32365 Would there be any value in providing deferred script decorators? The decorator could schedule the compilation but not do it immediately, t",https://github.com/pytorch/pytorch/issues/32300,d8b0316f694589c0b66998734c08dde441ce0b4f1e413dd743f4bd3fd5223456 references,issue,56397,issue,11982,medium,issue.body,"ches the correct debug information from pytorch.org, then installs it so gdb will transparently symbolicate crashes. This will also address #11982 Alternatives Do nothing, keep repro-ing issues manually and being sad.",https://github.com/pytorch/pytorch/issues/56397,258845323454e330a885582263b1b1c759580d8bf773813355001de8da5ce6de references,issue,61079,issue,59682,medium,issue.body,"Similar to #59682, we should have an automated way to check things like in #60976 (comment) We can implement this as a custom clang tidy check in pytorch/tes",https://github.com/pytorch/pytorch/issues/61079,64cc09a0a0e3fa028c26c7873a9e5cb36f890031633523d770de45f69bf21916 references,issue,68204,issue,54982,medium,issue.comments[1].body,Related RFC: #54982 (comment),https://github.com/pytorch/pytorch/issues/68204,cfe2b94db9e8b8d4f0d8abd9d484af80eb33fca22a1de33fbaebf6fc71d153f9 references,issue,68104,issue,10471,medium,issue.comments[1].body,Filed on the lightning repo with issue #10471.,https://github.com/pytorch/pytorch/issues/68104,c5b85dc4a40b72d0c6b14b501db0e5be9e8564ef916614b40fb6e1d47dd04b25 references,issue,64897,issue,55366,medium,issue.body,thout an explicit cast This is also quite wasteful to copy-allocate torch.bool to torch.int64 to perform these basic operations :( Related: #55366 cc @heitorschueroff @brianjo @mruberry,https://github.com/pytorch/pytorch/issues/64897,48cf8ee26da7eaa2c003ecb7cf2e555cb2e982ac1d2671988c3a7c8f5771329b references,issue,65913,issue,39279,medium,issue.comments[1].body,"Hi, Thanks for sharing these details. We are indeed looking into providing these functionalities (via different APIs). You can see #39279 and #49171 for example. These two would provide similar functionality as what you're looking for right?",https://github.com/pytorch/pytorch/issues/65913,6995b81486c9407089aaed5a44b79d1d4899dd832c1f1f72a15c0e0e72935504 references,issue,67589,issue,67590,medium,issue.body,"Optimizer) -> str: return self._per_optimizer_states[id(optimizer)][""stage""].name This addition is also useful for the reason described in #67590 Alternatives Make self._per_optimizer_states public. Additional context We are accessing the protected attribute in PyTorch Lightni...",https://github.com/pytorch/pytorch/issues/67589,53ad0aa1e6ef6aedfe9c7fe537f25f9985f505aff08bd4f8109aa88fcf7153ce references,issue,66627,issue,55571,medium,issue.comments[0].body,Related #55571,https://github.com/pytorch/pytorch/issues/66627,c91149bd67753077a37b63dcfa4a53f16fda5d2ad1c0085273034c05cb870d08 references,issue,66824,issue,64067,medium,issue.body,"#65514 introduced a Python framework for running C++ tests(the benefits and reasoning are described in #64067). While that PR added a framework for running tests, we still have many manual calls to C++ test binaries in the scripts used in CI. This i",https://github.com/pytorch/pytorch/issues/66824,7677aa354633ba351baff028ae6a564d5ae8a2b9a1d8eb4f5f4b477f66f8965e references,issue,66750,issue,66751,medium,issue.comments[0].body,Possibly similar issue to #66751 ?,https://github.com/pytorch/pytorch/issues/66750,6d41ef384262bdd7340de18724f995ba7819f9408fbf019abd421cd98dd37980 references,issue,43127,issue,66640,medium,issue.comments[1].body,Same Issue: #66640,https://github.com/pytorch/pytorch/issues/43127,d01f69df8aa0a143f7a869a1f9db58f66c9dff4eee9c05369369be18a94145e9 references,issue,55340,issue,53623,medium,issue.body,"a particular feature is tested, and slight semantics differences in using fork/spawn/multiprocessing for the 2 files, such as the issue in #53623. Desired Outcome The desired outcomes of this effort is to 1) Make it easier and clearer for developers and OSS contributors to wri...",https://github.com/pytorch/pytorch/issues/55340,33a29d538af4d1e2ed76828dbbca7a52ee92246c1024592c15bdc77f699c9c14 references,issue,64730,issue,66073,medium,issue.comments[1].body,"Hi, Thanks for the detailed issue. Isn't this similar to #66073 ? In particular if you set the Module in eval mode, this is the behavior you get right?",https://github.com/pytorch/pytorch/issues/64730,6240b460e0dc2e8cecf178022fea5bb1db3edc7aa9365e98d7df947fca07826d references,issue,45769,issue,36333,medium,issue.comments[0].body,"#36333 might be related, quadro 8000 + cuda 10.2 + cudnn 7.6.5",https://github.com/pytorch/pytorch/issues/45769,27b14bdc0b526646935266595372407af11275ce3e2338c41e0600ac1fbf1975 references,issue,13222,issue,56891,medium,issue.comments[1].body,This is also related to #56891 where the slowness has the same origin as here.,https://github.com/pytorch/pytorch/issues/13222,0fed3a13de998792eb85294ab9051a76f0d0b1333b936f9d7145846732db8396 references,issue,65760,issue,65788,medium,issue.comments[0].body,Fixing this might be relevant re: #65788,https://github.com/pytorch/pytorch/issues/65760,15cd6f55e2252a211fdf7c22440f14fd62a717695ba31fc7dcd60a12dcbaeab1 references,issue,50802,issue,49949,medium,issue.comments[0].body,"Great idea, @shaibagon. Linking to the more general issue here: #49949.",https://github.com/pytorch/pytorch/issues/50802,c8e98845b5c881f21c352e197aa9d5171d24be73b38f3f7411c7740f093cb5cc references,issue,65132,issue,65131,medium,issue.comments[0].body,"ass(cls, torch.as_tensor(data, dtype=torch.float32, **kwargs)) print(XLATensor([0, 1], device='cpu')) Maybe this also fixes your problem at #65131",https://github.com/pytorch/pytorch/issues/65132,d0fdcfadf983d4b8e0adc718490b9429cdb0008fed88bc9beda7fedbbdb10c59 references,issue,50962,issue,34306,medium,issue.comments[0].body,Related about 1d conv_tbc: #34306,https://github.com/pytorch/pytorch/issues/50962,05671cdbe4ed22c2e4c34f378b8c87abe6aac5dbb1533c1cef0971ef5fbfc1c0 references,issue,44077,issue,22155,medium,issue.comments[0].body,Dup of #22155,https://github.com/pytorch/pytorch/issues/44077,13ffa2ac37005d1065b6eab70313752ca4320b46ded9fd82a009d93fd1d52d3c references,issue,64881,issue,49440,medium,issue.comments[0].body,e system to support more flexible data-loading process. Then you should be able to load data directly into GPU. Please checkout out our RFC #49440,https://github.com/pytorch/pytorch/issues/64881,b5f078940697b1291a06d6c356a1d74998c26e1e5f96e1e1d3063db891a6952f references,issue,17966,issue,13207,medium,issue.comments[0].body,#13207 related?,https://github.com/pytorch/pytorch/issues/17966,ca03efbf0ea9bea1eb0b2177423b69693de2dd02adbfabb23f790aa15e9a49e6 references,issue,30873,issue,2478,medium,issue.comments[1].body,"reas of noisy loss, and ReduceLROnPlateau kicks in and leaves me stranded at the top of a spike, then the LR collapses and training stalls. #2478 is another attempt to solve a similar problem by adding backtracking, and that got as far as a PR in #2544 but progress there seems...",https://github.com/pytorch/pytorch/issues/30873,dfa15eff6de2d51ef340d8e9b92a4e8f9b3de0adef2bf6f897d4641b857f2d46 references,issue,64407,issue,62032,medium,issue.body,"As documented in #62032, an autograd ""not implemented"" kernel would provide a function/macro to register an autograd kernel that forwards but raises an error on ba",https://github.com/pytorch/pytorch/issues/64407,57ca2718a1bce405773ad8c0fcd88a83de408b2b15ab30fd2583f506590f1238 references,issue,46168,issue,26288,medium,issue.body,uffer (this is often the case for normalization or elementwise ops). This would lead to better memory utilisation. Originally discussed in: #26288 (comment) @albanD's argument was that this may break hooks. I propose to allow some ways for user to explicitly guarantee that no...,https://github.com/pytorch/pytorch/issues/46168,89971846afaf86d3a96e32102a3e233f46d03dd1af0fcc806ea297d78cb769e0 references,issue,62032,issue,35041,medium,issue.body,"mpl to also automatically apply this registration when someone registers a CPU/CUDA implementation directly, but we would need to implement #35041 so that subsequent registrations of autograd would override this behavior.) A simpler, albeit less efficient, way to solve this pr...",https://github.com/pytorch/pytorch/issues/62032,6e90999c3e697f543277e8631ed9a9dbbd46d81d6926f0fe089f93c12b6845bc references,issue,38185,issue,634,medium,issue.comments[0].body,Related feature request: #634,https://github.com/pytorch/pytorch/issues/38185,03dc4b4ba7e544f3569a68977d46a0066360174d2c1cc27180b7b4446e76886c references,issue,50344,issue,48628,medium,issue.body,ch.special for existing PyTorch functions analogous to scipy.special functions (#50345) Doc Improvements Create a NumPy Compatibility note (#48628) RFC: Update PyTorch function documentation to identify the corresponding NumPy function (#50343) cc @ezyang @gchanan @zou3519 @bd...,https://github.com/pytorch/pytorch/issues/50344,353f6a2a78547a2efa170b35df6922e68c932bd54b6382e06087e9885f9a0067 references,issue,50344,issue,50343,medium,issue.body,vements Create a NumPy Compatibility note (#48628) RFC: Update PyTorch function documentation to identify the corresponding NumPy function (#50343) cc @ezyang @gchanan @zou3519 @bdhirsh @jbschlosser @anjali411 @mruberry @rgommers @heitorschueroff,https://github.com/pytorch/pytorch/issues/50344,8960e5cc2727fc984bee733e99931a1705244cddf79c9e54bd4301f4f85f94f2 references,issue,50344,issue,50345,medium,issue.body,"ing"" functions in the NumPy namespace (#38349) Implement torch.special for existing PyTorch functions analogous to scipy.special functions (#50345) Doc Improvements Create a NumPy Compatibility note (#48628) RFC: Update PyTorch function documentation to identify the correspond...",https://github.com/pytorch/pytorch/issues/50344,7c5b9e3d2cd137aa4a8cb3b49383df834799651e0a0535d399b2f460a189ba52 references,issue,46544,issue,46225,medium,issue.body,"ue is for discussing how PyTorch should handle NaN values and how we should design our operator API to do that. We have many issues such as #46225, reporting inconsistent handling of NaN values. One solution is to follow NumPy's design and have nan* ops which ignore NaN values...",https://github.com/pytorch/pytorch/issues/46544,6cc9dda3370b9252cd796c420d707ebddb21f049620f4ed8702a8dfccc32e832 references,issue,52984,issue,11578,medium,issue.body,xref #11578 (comment) for the current status on using pytest. improve the pytest output on test failures (probably through settings in a pytest.ini in,https://github.com/pytorch/pytorch/issues/52984,919f847cab7cb72f91f5725c8b6967b9e08f7e20131e1e4591858df49472fac8 references,issue,54006,issue,52984,medium,issue.body,"GPU tests. CPU sec 1 (missing, sorry) 2 253 4 225 8 251 16 250 32 233 64 234 GPU sec 1 2572 2 1385 4 745 Originally posted by @jeffdaily in #52984 (comment) cc @ezyang @seemethere @malfet @walterddr @lg20987 @pytorch/pytorch-dev-infra @mruberry @VitalyFedyunin",https://github.com/pytorch/pytorch/issues/54006,a3aed5067b20bdb858460dce8fe550fec959ce8a952632542e44844f124f317c references,issue,55159,issue,55152,medium,issue.body,This proposal is loosely related to #55152 but I preferred to split them because I think this one is more controversial. The op_db database is getting bigger and bigger and it's beco,https://github.com/pytorch/pytorch/issues/55159,6103c25c297cdc5132e1acb710e3b3c5fe596406f3be412e52a6cdcd89e5101f references,issue,42847,issue,63085,medium,issue.body,"Update: the test_nn.py part of this is now tracked by #63085 Currently, many files, such as test_jit.py and test_nn.py, are larger than what GitHub can handle. For these files, you will not be able to",https://github.com/pytorch/pytorch/issues/42847,72644185c1c30b7a24a4a45d0242993929e7239ada7a6983c3f0603687cfea85 references,issue,49909,issue,49758,medium,issue.body,See #49758 for an example (this particular example turns out to not have extra cross device synchronizations). But it would be good to give ourselves,https://github.com/pytorch/pytorch/issues/49909,a9713298a9aadafc91b7350a2ea2ed9c02cf8a940e1959cc144002e459af2a94 references,issue,17234,issue,7795,medium,issue.comments[0].body,#7795 is probably related to this.,https://github.com/pytorch/pytorch/issues/17234,05541b33011222b5bf4ae78df2c8b52c583fdc49f1e70e16c16692ab7fc41e28 references,issue,63618,issue,43947,medium,issue.body,🐛 Bug There are cases where calling torch::cuda::synchronize() appears to speed up training. There have been some other issues on this (#43947 and #44103 ) I've noticed the counter-intuitive slow downs from the c++ side when replacing tensor.item calls with non-synchronizing a...,https://github.com/pytorch/pytorch/issues/63618,52c1594b68051958593666072cd84d50a28a1a296ea4ab8935244a3ec859dfc6 references,issue,63618,issue,44103,medium,issue.body,here are cases where calling torch::cuda::synchronize() appears to speed up training. There have been some other issues on this (#43947 and #44103 ) I've noticed the counter-intuitive slow downs from the c++ side when replacing tensor.item calls with non-synchronizing accumula...,https://github.com/pytorch/pytorch/issues/63618,5c7deed39f5c73d4493e45ec6d971536c0bf10e6012b800b95cc042e27b7208d references,issue,63485,issue,47117,medium,issue.body,"🚀 Feature Preserve tensor subclasses when unpacking a SavedTensor Motivation This is motivated by #47117 Pitch Currently, a new tensor object is created when we unpack a saved variable. This loses any subclass information during the backward. U",https://github.com/pytorch/pytorch/issues/63485,c800e6013d1b60463f8eef9c0c0d7017c8734f3206f40d63715cb514200fc56f references,issue,63295,issue,59855,medium,issue.body,"() File ""test/run_test.py"", line 905, in main raise RuntimeError(err_message) RuntimeError: test_autograd failed! Some of it is the same as #59855, but there seem to be more failures this time. cc @malfet @seemethere @walterddr @ezyang @albanD @zou3519 @gqchen @pearu @nikitave...",https://github.com/pytorch/pytorch/issues/63295,062b15224cf306dab9b3855f162850a8739789b26ad5af12f5fcb4b4290e4dc0 references,issue,63023,issue,48353,medium,issue.comments[0].body,related: #48353,https://github.com/pytorch/pytorch/issues/63023,b56c5a124a2452b519be986464ecc7e442a731e7eea5fe41cb33d3c1018fe8c2 references,issue,50012,issue,5212,medium,issue.body,"Note this issue was previously discussed in #5212, too. Despite having the same name, torch.split and np.split are different functions. The distinction is a little subtle when just reading",https://github.com/pytorch/pytorch/issues/50012,3c237a54b6a820126b9af11c0bbc9a278c00b8053147d5c63cb1b922ebe6a4c9 references,issue,61490,issue,58745,medium,issue.body,th the values when reducing over one dimension. This is incompatible with the Python Array API Standard as described in the following issue #58745 and furthermore makes it difficult to support reducing over multiple dimensions. Alternatives a dedicated operator for computing t...,https://github.com/pytorch/pytorch/issues/61490,503957391b25b5e0e2c657a9ed61a622930243d623d4f49b06049bdc5301d6a9 references,issue,63145,issue,11967,medium,issue.body,"nd configuration: N/A Any other relevant information: N/A Additional context I found another problem with autograd in torch.jit.trace here: #11967 However, it seems like the error message is different (with much older pytorch version), so I'm not sure if it comes from the same...",https://github.com/pytorch/pytorch/issues/63145,24c19725f88b1ce3b9b89c3667526151fa0326ff815bc1b2f37576695aa3a65d references,issue,57815,issue,28594,medium,issue.comments[0].body,A related issue: #28594,https://github.com/pytorch/pytorch/issues/57815,8c4e11c0f99b6788c5395f1241b98e3b06a91d7d3561695b8816e767ec79aa2f references,issue,42427,issue,42355,medium,issue.body,hope build_android can support nn::module nn::Functional nn::Linear want to train in android app also refer to #42355,https://github.com/pytorch/pytorch/issues/42427,e8971a2f4ba1ac861028408fa7cc406ccf45fd9b886ae7360b3ebdba8248c1d6 references,issue,62171,issue,47149,medium,issue.comments[0].body,Related request: #47149,https://github.com/pytorch/pytorch/issues/62171,6c79fc3bddbe35f3afc272135fb5b5a686b143f3c9ea13a16fe68710104bba8a references,issue,62931,issue,52675,medium,issue.comments[0].body,Related in the sense of theme of avoiding CUDA initialization for functioning: #52675,https://github.com/pytorch/pytorch/issues/62931,5c2ee1d7b2c462e444ed9908fe6c8ffa58c5ab53d411ce0ebd9e6b9fa8d988b6 references,issue,16708,issue,13188,medium,issue.comments[1].body,"ne nasty issue that prevents detection of correct headers (even if the correct directory is in CPATH at the first place) is GCC's multilib: #13188. mkldnn detection problem may be from another cause, but a smoke test during cmake that checks that detected headers indeed come f...",https://github.com/pytorch/pytorch/issues/16708,b6362243b754a643dbcd00c369046f98bed4b2725e7252fb3b5a80696f0cf9cd references,issue,62784,issue,61486,medium,issue.comments[0].body,"Thanks for the report @842974287. This is a known issue, using dim=None is not possible for any operator. See #61486 (comment)",https://github.com/pytorch/pytorch/issues/62784,39065d6eeb91abdffdce129451c099de8f02c24f0b39a56aeeaff054e4fc10af references,issue,16668,issue,1529,medium,issue.body,it is only safe to reuse the block if we are returning it on the same stream (which will cause us to setup the correct ordering). Related: #1529 CC @zheng-xq cc @ngimel,https://github.com/pytorch/pytorch/issues/16668,7505be92548bf7ed82c6a8c429778838723a9a54cdc8766dcfbd6e75ec2e476c references,issue,62448,issue,23756,medium,issue.comments[1].body,"related issues on invertible/inplace nets and saving memory (that also recover the activations without storing them), but not DTR-related: #23756, #46168, #26288 :)",https://github.com/pytorch/pytorch/issues/62448,ac1468441fc0c02e3a98a994000c820854872f55511dcb4f5a82c67b534bfbe3 references,issue,62448,issue,26288,medium,issue.comments[1].body,"on invertible/inplace nets and saving memory (that also recover the activations without storing them), but not DTR-related: #23756, #46168, #26288 :)",https://github.com/pytorch/pytorch/issues/62448,045d211a0d8697b22378772f4a1b40e4d09a47dae6476f5aff16e2f553a5f90b references,issue,62448,issue,46168,medium,issue.comments[1].body,"issues on invertible/inplace nets and saving memory (that also recover the activations without storing them), but not DTR-related: #23756, #46168, #26288 :)",https://github.com/pytorch/pytorch/issues/62448,e05229fdde143fd5cd1b9676d704a0380450fe24b57764b82b68f47c2ce08c1c references,issue,62147,issue,52265,medium,issue.comments[0].body,"Related: #52265 As described in the discussion for that issue, the function is_subclass() is provided to get around this. See the example here.",https://github.com/pytorch/pytorch/issues/62147,5233cea3e04544a30d838b8a05fd707a8a3479d4202ddf883ed0b4eea2dd90e0 references,issue,54982,issue,26889,medium,issue.body,"am interested in helping and talking to anyone interested in working on any of these things Related issues: #52753, #46948, #43501, #40373, #26889 cc @gmagogsfm",https://github.com/pytorch/pytorch/issues/54982,f3980333dde1b549bf3cc36ac9d175fc11bbc7621c6006f55af985a100e59201 references,issue,54982,issue,40373,medium,issue.body,"n but I am interested in helping and talking to anyone interested in working on any of these things Related issues: #52753, #46948, #43501, #40373, #26889 cc @gmagogsfm",https://github.com/pytorch/pytorch/issues/54982,7b4148146996c4db6d23334436d5e29ecbef62b1995a4085d369dbc558f74cd1 references,issue,54982,issue,43501,medium,issue.body,"orking on but I am interested in helping and talking to anyone interested in working on any of these things Related issues: #52753, #46948, #43501, #40373, #26889 cc @gmagogsfm",https://github.com/pytorch/pytorch/issues/54982,4f1cc95239dd925b92fc69a47a4f18243f8dc9a737a7cec762e754a4c044b5ee references,issue,54982,issue,46948,medium,issue.body,"ing on working on but I am interested in helping and talking to anyone interested in working on any of these things Related issues: #52753, #46948, #43501, #40373, #26889 cc @gmagogsfm",https://github.com/pytorch/pytorch/issues/54982,87498ff43caaa6cd5fb9406bee4f02c660f0fa3f5bb9b05f4d97dd72be2124f4 references,issue,38995,issue,25416,medium,issue.comments[0].body,cc @agolynski related to #25416,https://github.com/pytorch/pytorch/issues/38995,cbe2d33c19920bb3fa4a224aeda0f86cf184e53268e78fcec1ed7b3c5ba904f5 references,issue,62009,issue,2001,medium,issue.comments[0].body,Somewhat related: #2001 @duwangthefirst would the torchinfo package satisfy your need for structure printing functionality?,https://github.com/pytorch/pytorch/issues/62009,2ed0be08f03babc6f2ded3073540feb6cc4a4302ff8bd24e2546663b8f0e9fba references,issue,57534,issue,55375,medium,issue.body,"s not clear why someone would like to do that, this problem appears naturally when using DataParallel with complex tensors. As mentioned in #55375, DataParallel casts tensors back to real, so one needs to call view_as_complex. However, while on GPU0 all tensors have a storage...",https://github.com/pytorch/pytorch/issues/57534,aea185a596d55fe114d98b339eb3d039b295ccbbcd195c1ea8289b7617b8b03a competes with,issue,51138,issue,50444,medium,issue.body,"be in float format (i.e., 1. instead of 1). Please also check if multi-dimensional tensors are printed out in the right format. Related to #50444 cc @gmagogsfm",https://github.com/pytorch/pytorch/issues/51138,ff0781ec01e86c589b509746e981b18bbeab4ddb642c530c7520f48eb53f7549 references,issue,61650,issue,56891,medium,issue.body,rix and its rank would be very helpful to make them more usable. Format the docs for them to be in line with the rest of torch.linalg Solve #56891 cc @jianyuh @nikitaved @pearu @mruberry @heitorschueroff @walterddr @IvanYashchuk @xwang233 @lezcano,https://github.com/pytorch/pytorch/issues/61650,a0f284cbb1d01d10148bf260606ca28b15458a2440a2398359573a4223e93072 references,issue,61528,issue,50341,medium,issue.body,A request to implement SciPy's scipy.ndimage.map_coordinates. First requested here #50341 (comment) cc @mruberry @rgommers @heitorschueroff,https://github.com/pytorch/pytorch/issues/61528,c0ec72f7ca7e0e335b3cb637d5b54e5b4e6ad557a0d91ef6ada7ec0d3429d164 references,issue,15617,issue,499,medium,issue.body,"yer but without weight sharing. Motivation Similar to the necessity of the transposed convolution layer, the locally connected layer (issue #499, PR #1583) should also have a transposed version. cc @albanD @mruberry @jbschlosser",https://github.com/pytorch/pytorch/issues/15617,23dc9b8e990994a0a2e8672802e39e076afb35c32687262f8aeb78578dcdac67 references,issue,61411,issue,30532,medium,issue.comments[0].body,"The CUDA Compute arch (3.7) of the K80 is not supported by the pre-built binaries. See: #30532 (comment) In the linked issue, a user provides custom built versions to support old Tesla cards.",https://github.com/pytorch/pytorch/issues/61411,ef1932db92661ca17e095145092cf2341b51c6ca6cc2703e5755852676ecbfff references,issue,61453,issue,61417,medium,issue.comments[0].body,"@pyscorcher Thank you for this issue, we do have planned to add support for this very soon. See #61417",https://github.com/pytorch/pytorch/issues/61453,db520a38644c3753545209906b67d988bdfc59f08459c1d57cff5788d1e35beb references,issue,59218,issue,32407,medium,issue.body,"nclear what the benefit of choosing Eigen + *BLAS is over just choosing a BLAS lib It would be great, if this was clarified. Maybe related: #32407 (comment) which hints that maybe using Eigen could be dropped as a BLAS lib is always required. cc @malfet @seemethere @walterddr...",https://github.com/pytorch/pytorch/issues/59218,64daa55eea727c285720b45a591267e18bf5dcecd5fe8356808192f75ffdac85 references,issue,61260,issue,60341,medium,issue.comments[0].body,"Seems related to #60341 Never compiler should solve the issue, or explicitly link your project against libtorch_cpu (and libtorch_cuda)",https://github.com/pytorch/pytorch/issues/61260,5539e2b3003f94ec9ec75daefa1b68ee0f2727971f310a6b0d9cc100c50710d7 competes with,issue,61213,issue,61211,medium,issue.body,"Fourier space by transform + low pass filtering + inverse transform if the high frequency modes are truncated rather than set to zero, cf.: #61211",https://github.com/pytorch/pytorch/issues/61213,ea7c55a57f39404c1637441d9c3571e7c508f5ea2cb7285b653f091b5674939e references,issue,45565,issue,44316,medium,issue.body,. These parallel implementations may need to distinguish between C++-level overloads so they can create their own parallel overloads. (xref #44316) Pitch The proposal is to add a get_current_overload_default_kwargs function which takes no arguments into torch.overrides. It wil...,https://github.com/pytorch/pytorch/issues/45565,9918713745057f586cf8c0e83043523fac1e2167a1f54318fe53586022f675cb references,issue,60541,issue,28224,medium,issue.comments[0].body,"#28224 My goal is the same as him, but the method is different. In fact, this is the synchronization problem between CPU and GPU. I want to block",https://github.com/pytorch/pytorch/issues/60541,9f2f6f7e97120edca960581460c403876298ba2d77744ee10af05d5d2804156d references,issue,44634,issue,14945,medium,issue.body,"ls) - PR gh-49158 Add sparse tensor support for all element-wise functions that map 0 to 0 (this is low-hanging fruit) #45113 #45897 #45996 #14945 Create public API to replace use of _indices, _values, _nnz. See gh-45695. 1.9 release fill-value support in COO and GCS sparse te...",https://github.com/pytorch/pytorch/issues/44634,cdb87830f1f0ba86fbfc6d78f90f485891241ca9045ed4a65b4eefd390492140 references,issue,44634,issue,45897,medium,issue.body,"here for details) - PR gh-49158 Add sparse tensor support for all element-wise functions that map 0 to 0 (this is low-hanging fruit) #45113 #45897 #45996 #14945 Create public API to replace use of _indices, _values, _nnz. See gh-45695. 1.9 release fill-value support in COO and...",https://github.com/pytorch/pytorch/issues/44634,fa6ae8d5756a7c7503186fbfa721535d58a0c1821995821b9a9c89926e48f822 references,issue,44634,issue,45996,medium,issue.body,"r details) - PR gh-49158 Add sparse tensor support for all element-wise functions that map 0 to 0 (this is low-hanging fruit) #45113 #45897 #45996 #14945 Create public API to replace use of _indices, _values, _nnz. See gh-45695. 1.9 release fill-value support in COO and GCS sp...",https://github.com/pytorch/pytorch/issues/44634,1ec34502cd91a5ed8aa71c49ff6363a5c1eea8f5cc84fa02433e702303baf4ac references,issue,44634,issue,56485,medium,issue.body,"ultiplication, gh-3158 CSR sparse tensors The state of CSR sparse tensor support as of March 2021 #50937 Sparse tensor CSR layout, CPU only #56485 Sparse tensor CSR layout for CUDA CSR sparse tensor: ops not implemented yet #59060 Addition of sparse CSR tensor csr_sparse @ csr...",https://github.com/pytorch/pytorch/issues/44634,b1904c0527511fa232ca51fe8418de132496eb3ad3e3bbf7e3a75556613ed587 references,issue,44634,issue,60858,medium,issue.body,support #56371 - modernize CSR test-suite #56369 - eliminate global usage of torch.set_default_dtype cuSPARSE backend: #60854 MKL backend: #60858 cc @ezyang @gchanan @zou3519 @vincentqb @aocsa @nikitaved @pearu @v0dro @mruberry,https://github.com/pytorch/pytorch/issues/44634,35ee2c5992381c5cb52fc2cd5efe8d98378cf3374567fb8cff5040975dfb4713 references,issue,44634,issue,14945,medium,issue.comments[0].body,"I suggest adding #14945 to 1.8, performance of sparse_coo_tensor call is very bad, we should either fix it or expose unsafe version to python, or, preferably, both",https://github.com/pytorch/pytorch/issues/44634,2c65b2830250cb9edd1eb3bb3c29804b9385fa1c445fe674cd00819237df8673 references,issue,44634,issue,14945,medium,issue.comments[1].body,"I suggest adding #14945 to 1.8, performance of sparse_coo_tensor call is very bad, we should either fix it or expose unsafe version to python, or, preferably, both",https://github.com/pytorch/pytorch/issues/44634,987c8fc22e88855d45bd2e0872c43eaddb9c532a9cdecf0c34d84ed6b5f55909 references,issue,24823,issue,19826,medium,issue.body,This issue is an expansion of the issue reported in #19826. The discussion there diagnoses the segfault that occurs in the vectorized 2D CPU kernel. This issue covers the wider problematic handling,https://github.com/pytorch/pytorch/issues/24823,57b4983e4f817db0b65877f60bd0d79da535bf9de551d129da39e5d204d31e30 references,issue,12659,issue,7313,medium,issue.comments[0].body,I don't think this fits well with how optimizers operate on parameters currently. I wonder if something like #7313 - changing what Parameters are - could help here. You could collect the gradient updates in a calculated parameter (which would still be a,https://github.com/pytorch/pytorch/issues/12659,a78bb5692351e89f43b61147e4017ddd2bd3a6cbb31b9fa6d5b85ca891cb370f references,issue,60153,issue,60149,medium,issue.body,and doesn't have a (Source) link because it is not a python function. e.g. I'm trying to research when an API has changed as explained here #60149 but I have no quick way to start the research by following a (Source) link. Thank you! cc @brianjo @mruberry,https://github.com/pytorch/pytorch/issues/60153,aa8b1be1d2bfaf924e8f51d56db8208c225092a3825ea631a607cba637f0ceb5 references,issue,59868,issue,56356,medium,issue.comments[1].body,"but these global flags are not silver bullets because libraries built on PyTorch may or may not support them. Related ideas are in this RFC #56356, which proposed a moratorium on implementing type promotion in PyTorch beyond the three op classes discussed above.",https://github.com/pytorch/pytorch/issues/59868,b040da905774fb2c54ebdb97d62b1307da7a2d5f22b9795d1b73f1c5dc268df1 references,issue,56440,issue,59526,medium,issue.comments[1].body,I think this is blocked by #59526,https://github.com/pytorch/pytorch/issues/56440,ca572b5c7e473a9ec108ee01bd1cdae114bc98d56251679153b360cd09e7bf9d references,issue,58997,issue,37410,medium,issue.body,"sed more conveniently in bi-level optimization models. Method step() would then call forward(), and also be callable from __call__ Related: #37410 cc @albanD @mruberry @jbschlosser @vincentqb @iramazanli",https://github.com/pytorch/pytorch/issues/58997,6cd0cd0a07ff444b54e737fe676fd46725826eda48e73c34999f98666df6c96e references,issue,55056,issue,40932,medium,issue.body,"t in NaNs, leading to training divergence, it is often preferred to effectively ignore these cases so training can continue. See #41508 and #40932 for more info. Pitch Add a new, optional eps parameter with default None to the various softmax forms: from torch import nn from t...",https://github.com/pytorch/pytorch/issues/55056,28d0c0af829b090dff1205aec0826fc519f1dd7863cdce4d0cdcd9ef26000601 references,issue,55056,issue,41508,medium,issue.body,"ently result in NaNs, leading to training divergence, it is often preferred to effectively ignore these cases so training can continue. See #41508 and #40932 for more info. Pitch Add a new, optional eps parameter with default None to the various softmax forms: from torch impor...",https://github.com/pytorch/pytorch/issues/55056,c1504a0d9b71bc5bb6c8ed21f3e05dee96340a73a62a6dee7e6dff929cec3e81 references,issue,40988,issue,17901,medium,issue.body,"🐛 Bug I am attempting to compile pytorch on Debian 32-bit VM. Compilation errors about AVX have been mentioned before #17901, but I still came across similar errors mentioning _mm256_extract_epi64, and they seemed to come from ATen not Caffe2. To Reproduce Steps t",https://github.com/pytorch/pytorch/issues/40988,be64d964a2f95455f08e9a226a2311de7f6e65b5930b3974004748f6e60dea9d references,issue,14436,issue,13993,medium,issue.body,"In #13993, a user noticed that when compiling with fbgemm, AVX2 instructions started being used in unrelated invocations of std::unordered_map. The r",https://github.com/pytorch/pytorch/issues/14436,0659315ce0f85aeeadecd804416726138fa67c86c708eb6703ca12b9dad55131 references,issue,59251,issue,58779,medium,issue.body,"build commands, ninja without -v doesn't output all build commands. Motivation Verbose output is hard to read at times. In combination with #58779, I once spent 20 minutes trying to find a build failure. Allowing extension writers to turn off the verbosity will lead to greater...",https://github.com/pytorch/pytorch/issues/59251,6956c1dc0227bf3327407e0172361c10fb1524feb7a7fafc81d51ce4492bc6bb references,issue,48491,issue,88,medium,issue.body,"90b8, callable=0x7ffff7661f70, tstate=0x555555921bd0) at /tmp/build/80754af9/python-split_1605449976777/work/Include/cpython/abstract.h:118 #88 PyObject_Vectorcall () at /tmp/build/80754af9/python-split_1605449976777/work/Include/cpython/abstract.h:127 #89 call_function (kwnam...",https://github.com/pytorch/pytorch/issues/48491,d4608b14d1c87c4bc8226c2ebe9fa2aa3565848a90ecd7cee874eb72379bac01 references,issue,58962,issue,17199,medium,issue.comments[0].body,"ain with the desired number of threads after creating the subprocess. BTW, had you tested it on v1.4 today? A similar issue was reported in #17199 but on v1.0.1.",https://github.com/pytorch/pytorch/issues/58962,089acc3f88d9604ea36035eadeedf6d71141ec42495c526dda941193ffea96a2 references,issue,58212,issue,47117,medium,issue.comments[0].body,"Hi @nelhage ;) This is pretty similar to #47117 but there's no save_for_backward involved here. This might actually just be straight up fixed by #56017, if you are able to build PyTorch f",https://github.com/pytorch/pytorch/issues/58212,d0185d1db7ef8487177e7dc0119152b16cbd3d7eb57e5b620ae58713f36e9062 references,issue,58212,issue,47117,medium,issue.comments[1].body,") I built from source on #56017 and it doesn't seem to have resolved the problem. Let me know if there are other tests I can run. I noticed #47117 when searching for this and was hoping the lack of save_for_backward made this one more minimal or straightforward, but then forgo...",https://github.com/pytorch/pytorch/issues/58212,5dc3293cab85576b32bc36c7cc86a27bcc93eb6da9d95b12c4abf02027f0f76c references,issue,13993,issue,88,medium,issue.body,"ythonrun.c:949 #87 0x00005578178cb16f in PyRun_SimpleStringFlags () at /tmp/build/80754af9/python_1540319457073/work/Python/pythonrun.c:445 #88 0x00005578178cef7a in run_command (cf=0x7ffcd85f149c, command=0x557819537f40 L""import torch\n"") at /tmp/build/80754af9/python_1540319...",https://github.com/pytorch/pytorch/issues/13993,8b57b42bcc32cfd81921ac9808132e9410b1916f94b87e95ac920298d019d816 references,issue,53903,issue,41530,medium,issue.body,"3.7,5 CUDA/cuDNN version: 10.2 GPU models and configuration: RTX 2070 Any other relevant information: Additional context This is related to #41530 cc @ezyang @albanD @zou3519 @gqchen @pearu @nikitaved @soulitzer",https://github.com/pytorch/pytorch/issues/53903,6e617d73bd660d0c7215e3c3e3112e04bebacab1aa1217cce820f2685b8b6107 references,issue,58136,issue,42109,medium,issue.comments[0].body,Duplicate of #42109. Here's a profiling result that I had for #42109 GPU activities: 50.05% 9.05349s 12666 714.79us 713.56us 716.09us void at::native::indexAdd,https://github.com/pytorch/pytorch/issues/58136,c9229133c851773e5df10eb495a371a632e6ed2b0ade5cf26e595e4203bb279d references,issue,57617,issue,41186,medium,issue.body,🐛 Bug The gradients for complex128 on POWER (ppc9le) are wrong/inexact making the tests fail. This may be related to #55754 or #41186 To Reproduce Steps to reproduce the behavior: build pytorch and run tests ERROR: test_fn_grad_angle_cpu_complex128 (__main__.TestGradientsC,https://github.com/pytorch/pytorch/issues/57617,fabf2adb5ff998e8626ef88ec5c64b542c251deeba64647c8b4be29bce078a47 references,issue,46653,issue,46559,medium,issue.body,"#46559 mentioned that it's hard to debug async RPC, as the error is only thrown when users call Future.wait() but users might miss that. @lw sugge",https://github.com/pytorch/pytorch/issues/46653,740b6dd94ea8edb3a7ee473e67790738a4e097e6b440141ce3bb9c1d026405ef references,issue,55661,issue,36444,medium,issue.comments[0].body,Some context #36444,https://github.com/pytorch/pytorch/issues/55661,d0b583356e4bbca7be533ac07d88e5e646561234e69dcf9698c9a0fa8d96c713 references,issue,51962,issue,28341,medium,issue.comments[0].body,"I think @Balandat has been doing some work around LinearOperator, #28341",https://github.com/pytorch/pytorch/issues/51962,f67233aba6e8d6f450cac960d850ed8584ca4bf63e99e879bb3304543db321fa references,issue,56975,issue,5405,medium,issue.comments[0].body,Related: #5405,https://github.com/pytorch/pytorch/issues/56975,d335e64f6928e8aec7e5b60588b496e1cae8c2bba42c31ae7ef6d594f8556f79 references,issue,13207,issue,4716,medium,issue.body,"I installed master of PyTorch from sources as: TORCH_CUDA_ARCH_LIST=5.2 python3 setup.py install (because of #4716) I'm getting: import torch print(torch.cuda.get_device_name(0)) # 'TITAN X (Pascal)' print(torch.cuda.get_device_capability(0)) # (6, 1) to",https://github.com/pytorch/pytorch/issues/13207,c0227072ec6e280fe4fb6695ea31c29b781476c63b4beced826175b2fd97fffe references,issue,13207,issue,4716,medium,issue.comments[1].body,"rward-compatible as well? Can't 5.2-compiled code run on 6.1? (it seems it can't, but it's strange) The reason I was putting it manually is #4716. I'm using a machine which also has a small desktop GPU with capability of 2.0, and compilation fails, since it tries to target it...",https://github.com/pytorch/pytorch/issues/13207,b59091ff302234fc94121516ff925bed0da5486649b20b9bc56e038c77a0d315 references,issue,56921,issue,56480,medium,issue.body,per #56480 (comment),https://github.com/pytorch/pytorch/issues/56921,a81e85ed22b68d5fc6096089db6d22176c05e430c40ef7b7815e62e3b5872921 references,issue,56244,issue,55757,medium,issue.body,daFuture. extractDataPtrs uses getSubValues which throws on non-torchscript python objs. Additional context Hit this bug while looking into #55757 cc @osalpekar @jiayisuse @lw @beauby @pritamdamania87 @mrshenli @jjlilley @gqchen @rohan-varma @pietern @zhaojuanmao @satgera @aaz...,https://github.com/pytorch/pytorch/issues/56244,e62960885f26e702954a603b821a669bc7693dd789c8bce6ce5e0f6de6fbe815 references,issue,18536,issue,12484,medium,issue.body,"While debugging #12484 I wanted PyTorch to tell me which cuDNN convolution algorithm it selected. But I have no way of getting it to do so, except running nvprof",https://github.com/pytorch/pytorch/issues/18536,196f703965578bb8075e99c23db76b0b037a12d03b4c389a959f66f414a96841 references,issue,51896,issue,51782,medium,issue.body,can't assign a module's attribute to a lambda once it was a module before. This + absence of easy shortcuts for anonymous module creation (#51782) prevents some useful stubbing approaches with lambdas (or functions such). import torch class Foo(torch.nn.Module): def __init__(s...,https://github.com/pytorch/pytorch/issues/51896,52abdd483c833e51e3936faaa4d43f08526d1233a1599f8a2ef2fca99307961b references,issue,51137,issue,50444,medium,issue.body,~~~ pass ~~~~ <--- HERE Desired behavior The error message should simply say that interface types cannot define __init__ methods Related to #50444 cc @gmagogsfm,https://github.com/pytorch/pytorch/issues/51137,a39e27f98b8bedd964d150fc70d630a4bf3829cfe616e29fc0b773c63ebbb848 references,issue,52920,issue,50444,medium,issue.body,ython class object A preferred error message should be something like this: torch.jit.trace() expects function or module objects Related to #50444 cc @gmagogsfm,https://github.com/pytorch/pytorch/issues/52920,81ccaacb1740786e0ace2172910839380a2602bfda943d8a528d86eea776c128 references,issue,56297,issue,40137,medium,issue.comments[0].body,"Also relatedly #40137; it would also be nice if the traces reported what locks they held (esp the GIL) so I don't have to sleuth it manually, but that might be d",https://github.com/pytorch/pytorch/issues/56297,e76f645f9422140c36fd76609a09cc4e2feaeea3986658227bdf915ebd4a4d06 references,issue,12672,issue,4959,medium,issue.body,"er. In particular, for TensorDataset objects, this will massively speed up batch creation. (Currently DataLoader is prohibitively slow, see #4959.) It also allows Datasets to be used more flexibly outside of a DataLoader. Pitch I think it's quite natural to embed this collate...",https://github.com/pytorch/pytorch/issues/12672,d2c6420a836d7ef622b12da9acfb554a117563cca2e682699f7693a7bb201866 references,issue,55757,issue,56244,medium,issue.comments[0].body,I think #56244 blocks this because the futures implicitly have a then() callback attached to it when profiler is enabled which causes issues.,https://github.com/pytorch/pytorch/issues/55757,3bdc5d0ffd1b2ec2f70b7a33caf3e39a5c6b96e3716ccc83e2077834e25c9eb4 references,issue,48108,issue,16293,medium,issue.body,"hat makes all us use PyTorch for research. It also does not help that JIT UX improvement issues are almost never worked on: #29092, #29177, #16293, #40914 Related to 2: #25066 Finally, I have two solution proposals. They may not be viable, but are nice from UX perspectives. To...",https://github.com/pytorch/pytorch/issues/48108,ccc46ece2bd6bd0980ad2979b5366f7020081e6fbcc880ad728d62884c1135e9 references,issue,48108,issue,25066,medium,issue.body,"for research. It also does not help that JIT UX improvement issues are almost never worked on: #29092, #29177, #16293, #40914 Related to 2: #25066 Finally, I have two solution proposals. They may not be viable, but are nice from UX perspectives. To solve 1: Have a way to annot...",https://github.com/pytorch/pytorch/issues/48108,1bbd5e0385c1ed9ef8d68de33f0874cdb515a225d3d4d5696f0df04a72261075 references,issue,48108,issue,29092,medium,issue.body,"ally the thing that makes all us use PyTorch for research. It also does not help that JIT UX improvement issues are almost never worked on: #29092, #29177, #16293, #40914 Related to 2: #25066 Finally, I have two solution proposals. They may not be viable, but are nice from UX...",https://github.com/pytorch/pytorch/issues/48108,0adf7e55feb11be44ef889d3759279a87cabdc29ed32ca87a7ee56a0bfea1858 references,issue,48108,issue,29177,medium,issue.body,"thing that makes all us use PyTorch for research. It also does not help that JIT UX improvement issues are almost never worked on: #29092, #29177, #16293, #40914 Related to 2: #25066 Finally, I have two solution proposals. They may not be viable, but are nice from UX perspecti...",https://github.com/pytorch/pytorch/issues/48108,0718d555ee265003ededc4f971f345d61e2f4234c8de68a9bef78774b73075e5 references,issue,48108,issue,40914,medium,issue.body,"s all us use PyTorch for research. It also does not help that JIT UX improvement issues are almost never worked on: #29092, #29177, #16293, #40914 Related to 2: #25066 Finally, I have two solution proposals. They may not be viable, but are nice from UX perspectives. To solve 1...",https://github.com/pytorch/pytorch/issues/48108,58398e620879756c87f0320e3bc67a3b71568094563a2d974d66fad631a141d3 references,issue,56090,issue,55905,medium,issue.comments[0].body,#55905 is a related issue exposed by vectorizing with NNC but it's basically the same root cause of llvm <=9 not really knowing what to do with fp,https://github.com/pytorch/pytorch/issues/56090,5311ffc732d339905b3715d765405f2c8a00509229f51721572b0b6f3949971c references,issue,39443,issue,15849,medium,issue.comments[0].body,#15849 (comment) speed up use_list condition for me. use_spawn: 28.03s; use_spawn+use_list: 52.92s; use_spawn+use_list+FastDataLoader: 27.38s. My,https://github.com/pytorch/pytorch/issues/39443,5fa7106a1608c39738194e5935e4bc18a00d22f7e7ddcc124000fe8e592e4889 references,issue,52310,issue,49666,medium,issue.comments[0].body,Related #49666,https://github.com/pytorch/pytorch/issues/52310,6d17936e0d9037fbcfc7835d0677a8f2ddfbb2c839845e7c902df4b0b5b2dd2f references,issue,55095,issue,45565,medium,issue.body,"re inaccessible to torch function clients (e.g., FX); this includes aliasing and default arguments, which have previously been discussed in #45565. This proposal is to make it possible to call into __torch_function__ from the C++ dispatcher, by way of the Python dispatch key....",https://github.com/pytorch/pytorch/issues/55095,d7d96c58f91281bb9bc3e92b86979b28334780d89dfc8bebc39d3c032a43a817 references,issue,55095,issue,53937,medium,issue.comments[0].body,"Unfortunately, if you delegate in this way, the way to fixing #53937 is closed (because by the time you've gotten to C++, we really are expecting an honest to goodness int64_t, not some symbolic tracer). Non-",https://github.com/pytorch/pytorch/issues/55095,67b0b26f55bd69c6b51f4bf35b09691008433f1d2c008cf56ed7adb7ed698d70 references,issue,54622,issue,52310,medium,issue.comments[0].body,possibly related to #52310,https://github.com/pytorch/pytorch/issues/54622,56b5a8de75c032679aec991aa5ad697b97294531c633637f2d7c45fea3c055a5 references,issue,55297,issue,55296,medium,issue.body,st be no errors on this command. Environment PyTorch Version 1.8.1 OS: Android 10 Installing from source Arch: arm64-v8a Build command: see #55296,https://github.com/pytorch/pytorch/issues/55297,53da4cf8d6559718b5d0485737f921dc233c329ee3eddb9890958bc861cbe18c references,issue,55122,issue,47953,medium,issue.body,"uration of ~0.2 ms on my machine, which means it's mostly negligible for large matrix operations. Additional context See also #42403 #42666 #47953 cc @ngimel @jianyuh @nikitaved @pearu @mruberry @heitorschueroff @walterddr @IvanYashchuk @VitalyFedyunin @ptrblck",https://github.com/pytorch/pytorch/issues/55122,01f6aa81d1de3256705530bd6133a510e8d467880f8ac1e7d76abcc96acc76cd references,issue,55174,issue,44159,medium,issue.comments[0].body,Related #44159,https://github.com/pytorch/pytorch/issues/55174,87a64357c7da1c9528c0e3533fbce2bd19409e93654dbaee92809ed64fc6d8bf references,issue,55104,issue,55095,medium,issue.body,"e you also want to override a composite (you would have to do an actual operator specific registration). In some potential use cases, e.g., #55095, this makes it impossible to do a higher level override of the composite (because the backend fallback is the only mechanism by wh...",https://github.com/pytorch/pytorch/issues/55104,5c4e53dbe68aa238b561da8b7d32ac00daa342708ecde990aff5dd8e02a996ab references,issue,53707,issue,48246,medium,issue.body,"new device to work around the CUDA specialization. As the Pytorch supports more and more new devices other than CUDA, like XPU proposed in #48246, the abstraction of runtime interface is required to decouple the Pytorch modules from CUDA. Pitch Add general frontend device runt...",https://github.com/pytorch/pytorch/issues/53707,a2832d7045beaea927136fe1c5dfdbd7887a20b6ba56a8472d485474a3cdda30 references,issue,3877,issue,50345,medium,issue.body,"Torch (site, code). Is if feasible to wrap Cephes for similar use in PyTorch? EDIT: Request for the Cephes functions should be directed to #50345 (Issue for tracking torch.special) as most of these functions are exposed in scipy via special module. cc @mruberry @rgommers",https://github.com/pytorch/pytorch/issues/3877,fdf225d08df00c35b138f568cc1748f6a06d0db5305c89f1ca041d60be1c563f references,issue,53623,issue,19177,medium,issue.body,"nda] numpy 1.20.0 py39hdbf815f_0 conda-forge [conda] torch 1.9.0a0+git34d9278 dev_0 Additional context This was first reported in #19177 and again in #41337. This was attributed to pytest, but this only came up there, because pytest uses another execution order than...",https://github.com/pytorch/pytorch/issues/53623,5989c27e1dbd9d47f2e685464c2433754898f6ad8416e5be42590d798229f84c references,issue,53623,issue,41337,medium,issue.body,"39hdbf815f_0 conda-forge [conda] torch 1.9.0a0+git34d9278 dev_0 Additional context This was first reported in #19177 and again in #41337. This was attributed to pytest, but this only came up there, because pytest uses another execution order than unittest. The other...",https://github.com/pytorch/pytorch/issues/53623,f158864cb061712e311c78becb6893f71b9456e084945c7d7e82d7a1c236c8f3 references,issue,54138,issue,51156,medium,issue.body,ffers. __array_struct__ can be considered obsolete via PEP 3118- Buffer Protocol. Additional context This feature request is a follow-up to #51156 and https://pearu.github.io/array_interface_pytorch.html discussions. cc @mruberry @rgommers @heitorschueroff,https://github.com/pytorch/pytorch/issues/54138,cfe4124d50be8dff741c022f453de5f935cd0fb77fc5c7bf8be325ee06bffedc references,issue,54138,issue,51156,medium,issue.comments[0].body,"As noted in #51156 (comment), exposing pytorch Tensor data is problematic if the storage of a Tensor object is changed in-place. For example: >>> a = torch.te",https://github.com/pytorch/pytorch/issues/54138,bc2d6d6cfa87a74767b099196223a4ecdf238dc4002557d9419630b067022412 references,issue,54138,issue,51156,medium,issue.comments[1].body,".asarray(a_tensor) will prefer the new __array_interface__ attribute over Tensor.__array__. It's not clear to me that what I pointed out in #51156 (comment), // TODO: This attempts to keep the underlying memory alive by setting the base // object of the ndarray to the tensor a...",https://github.com/pytorch/pytorch/issues/54138,e9ab5c59a67e13986b43b70a0f9d2c0c36ff2276e4c354fbde53f2a1ff1c78a0 references,issue,54139,issue,51156,medium,issue.body,mpy 1.19.5 py39hdbf815f_1 conda-forge [conda] torch 1.9.0a0+git63e0e88 dev_0 Additional context This bug report is a follow-up to #51156 and https://pearu.github.io/array_interface_pytorch.html discussions. cc @ngimel,https://github.com/pytorch/pytorch/issues/54139,04c36eca985cc6e23e5ff293899b8511ba9193e27610a51e40f36cd5d3c3b0fa references,issue,46944,issue,42258,medium,issue.comments[0].body,"This is the root cause for #42258. As @alanhdu mentioned, it works if the class is available when running the model. But this only works for Python and is a real blocker whe",https://github.com/pytorch/pytorch/issues/46944,7a1dd51d689816f1c53800f153a2fb55d5b9eae6385bde1b6b4c12a846d35654 references,issue,22343,issue,3790,medium,issue.body,"licative factor to use. #3740, #21250, #22163 introduce variations on Adam and other optimizers with a corresponding built-in weight decay. #3790 is requesting some of these to be supported. We could instead have a new ""weight_decay_type"" option to those optimizers to switch b...",https://github.com/pytorch/pytorch/issues/22343,4abbf07ef5b317a61ad954c705852d705a1cd7419958e7160c38313b6ed5508d references,issue,53982,issue,10386,medium,issue.body,_dim = batch_dim) else: self.collate_fn = collate_fn ### OTHER STUFF HERE ### A similar (but more limited in scope) request was advanced in #10386 cc @ssnl @VitalyFedyunin @ejguan,https://github.com/pytorch/pytorch/issues/53982,767b93c9d5f71de1afa0f5cd7242f0d96810f64993ac527b09c4d99fe1323cf6 references,issue,53982,issue,49440,medium,issue.comments[0].body,Would be possible after migrating to DataPipes #49440,https://github.com/pytorch/pytorch/issues/53982,691e9c360dab289ad802c70147ecbc4a3c6082a2395d006298432672d34efcdf references,issue,11514,issue,50192,medium,issue.comments[0].body,Related: #50192,https://github.com/pytorch/pytorch/issues/11514,2944ca61c59724fc2f52be9eaf845da56294f69f49df458a12c930ed5489b1b2 references,issue,42849,issue,37002,medium,issue.body,"do-code (although the current implementation is in C++). The four concepts (Grad Reader, Grad Bucketer, Comm Scheduler, and Grad Writer) in #37002 map to four functions below. The grad_ready function triggers DDP backward logics. class ReducerBase: def __init__(self, model, pg...",https://github.com/pytorch/pytorch/issues/42849,a84445448dd063f19b62593148bfb30c70c2675fad2becb141835025251f38ca references,issue,50503,issue,38310,medium,issue.comments[1].body,@asobhy-qnx I also tried to cross compile libtorch for QNX and was having trouble on doing so #38310 . Could you please share how you were able to cross compile libtorch for QNX? This would really help me a lot. Thx,https://github.com/pytorch/pytorch/issues/50503,44d71ff73b53fce5d793b9f1dc2954e3dac839f469f2f10ee319943465b39e99 references,issue,53391,issue,39224,medium,issue.body,"rror checking will get you one way or another, but it would be better to just directly report that strides are illegally negative. See also #39224 #16424",https://github.com/pytorch/pytorch/issues/53391,114c21dfad083be26e77e96689d85520efdf3b06603b5c61b3a17bd78be5920b references,issue,41337,issue,19177,medium,issue.body,🐛 Bug This might be similar or the same as #19177. I'm getting test failures in the distributed tests. The failing tests are: DistributedDataParallelTest.test_accumulate_gradients_module Di,https://github.com/pytorch/pytorch/issues/41337,ecef2439cbedc6241e162865806d7b06b58667e2aa348eb4c3eb40d1636abe5a references,issue,41337,issue,19177,medium,issue.comments[0].body,"@mrshenli, did you by any chance have some context on why #19177 was failing? From some sources such as https://discuss.pytorch.org/t/what-does-runtimeerror-cuda-driver-error-initialization-error-mean/875",https://github.com/pytorch/pytorch/issues/41337,363ff76fd2aa244227c02798a6d02940bc79b84b8c47354b5376fff8b2a9ada0 references,issue,41337,issue,19177,medium,issue.comments[1].body,"@rohan-varma The tests reported in #19177 do not fail when run individually (say with pytest -k) but fails when run them together sequentially. My suspicion was that some test, alth",https://github.com/pytorch/pytorch/issues/41337,15f224ab7a72b2409cb5a3c85ec9828cb0af8b5049a9f29b8221c5ddbc7eb140 references,issue,53017,issue,50444,medium,issue.body,"pe), but not a != None. This asymmetry is unintuitive to users. Desired behavior a != None should be handled like a is not None. Related to #50444 cc @gmagogsfm",https://github.com/pytorch/pytorch/issues/53017,0a09c96d54284ad68641c524cc9c3ebd14bc92bb2218a7a6fef666a82dea892d references,issue,53022,issue,50444,medium,issue.body,"ected None but got int: File ""test-instance-attr-annotation.py"", line 11 self.x = None if x > 0: self.x = x ~~~~~~~~~~ <--- HERE Related to #50444 cc @gmagogsfm",https://github.com/pytorch/pytorch/issues/53022,c62de1d0185556a1b12d4c676e5daab8d09ad76959acfdbd4d15707ef008b1ef references,issue,50444,issue,51135,medium,issue.body,torch.tensor are silently ignored. Instead it should report error when users try to subclass builtin nominal types (opened a separate issue #51135 ) Test case: test_subclass.py. This test case overwrites __len__ of torch.Tensor but the overwritten method is silently ignored >...,https://github.com/pytorch/pytorch/issues/50444,0c2c3c8d83104ff56bcaca8d4df05ffe9cf768769ffd964067ee97640677e27f references,issue,50444,issue,51136,medium,issue.body,"er() ~~~~~~~~~~ <--- HERE When module type constructor is invoked inside TorchScript, the error message is mystic. (opened a separate issue #51136) Test case: test-module.py python test-module.py Eager: 5 Traceback (most recent call last): File ""test-module.bad.py"", line 23, i...",https://github.com/pytorch/pytorch/issues/50444,22d9621bd84cee29c157020e2e640567ebadb435f084ddf3d41695a0ef15ba28 references,issue,50444,issue,51137,medium,issue.body,"pes do not allow to define constructors, but the error message complains about return type of __init__ not annotated (opened separate issue #51137) Test case: test-interface.py > python test-interface.py Traceback (most recent call last): File ""test-interface.py"", line 4, in <...",https://github.com/pytorch/pytorch/issues/50444,7d0a7506655cc85ab007fe0a317e2282261be7c736e1b8ca0fb1aa9a0877eaa6 references,issue,50444,issue,51138,medium,issue.body,"line 7 def __init__(self): ~~~~~~~~~~~~~~~~~~~ pass ~~~~ <--- HERE tensor printout format differs from eager mode (opened separate issue in #51138 ) Test case: test-tensor-printout-differ.py > python test-tensor-printout-differ.py tensor([1., 1., 1., 1., 1., 1.]) Eager: True 1...",https://github.com/pytorch/pytorch/issues/50444,c15120260b7974289c4ea5d0058c938cb79a047c00325756fd9ec724eb27a3d6 references,issue,50444,issue,51140,medium,issue.body,"as wrong: python test-local-annotation3.py Eager: 2 TorchScript: 1.0 Non-intuitive error messages for user errors (opened separate issue in #51140) Test case: test-enum-mystic-error-message.py The test failed because em_fn was annotated as @torch.script.jit, but also supplied...",https://github.com/pytorch/pytorch/issues/50444,bb97aec1c8c5b9d3c6f8c0755baab9c2bf94b799679abca390a7eadad1717a5b references,issue,52800,issue,52788,medium,issue.body,See #52788 and also #50448 This is not the first instance of test misusing unittest.main(). We need a way to ensure all newly added tests are covered,https://github.com/pytorch/pytorch/issues/52800,332ad9056800242faae6d41f170b4390854796084b6ae4db59595e74bf8a1835 references,issue,52788,issue,52800,medium,issue.comments[1].body,This is not the first time such issue happens. Created #52800 to track a solution,https://github.com/pytorch/pytorch/issues/52788,c1ac5557ce757cad3c404a884c8ca952ea0af2f707f49aac48416b8c8eed0422 references,issue,52211,issue,4632,medium,issue.body,🐛 Bug A cudnn error is raised using the code below. I suppose that the issue is related to #4632 #13636. To Reproduce import torch import torch.nn.functional as F torch.backends.cuda.matmul.allow_tf32 = True torch.backends.cudnn.benchma,https://github.com/pytorch/pytorch/issues/52211,8b6da39cfcf6e85730460be3b56320bbfe4f412b5fe7ea7d3429e0a5c7359464 references,issue,52717,issue,52648,medium,issue.body,"This is much like issue #52648, but for torch.norm, not torch.linalg.vector_norm. However, torch.norm does not have the workaround that #51099 will add for torch.linalg.v",https://github.com/pytorch/pytorch/issues/52717,8e3278714c313ef33628f21f14dcde32e6d138edc948789de58d07ae922af316 references,issue,46886,issue,7313,medium,issue.comments[1].body,"Yup, #7313 :) But maybe even regardless, having supported convenience methods for removing buffers / parameters / etc can be useful.",https://github.com/pytorch/pytorch/issues/46886,4fe5bfbe3fc7e502273c78c0cd7e7fbb4da28121f349860b75b6c147c21425cd references,issue,46946,issue,46944,medium,issue.body,"onda with cpuonly Python version: 3.6 Additional context For context, this is biting us due to the backwards compatibility constraints from #46944. cc @gmagogsfm",https://github.com/pytorch/pytorch/issues/46946,95087a60061e4b87fee06a5f353e5f9a0480b8bdd865f0d6755a2d529fd7a792 references,issue,51835,issue,32976,medium,issue.comments[0].body,I think even pytorch's GRU is affected by that! I'm getting the same RunTime error when calling a jit'ed GRU. Sounds like #32976,https://github.com/pytorch/pytorch/issues/51835,8ae406212d1b4df49d2613fc210dd87fb022d9ae893527bd7c171be12b2f8097 references,issue,45465,issue,46749,medium,issue.comments[1].body,@nikithamalgifb - see the description in #46749 (comment) specifically https://colab.research.google.com/drive/1qWUq6PQ4fj67TBbD3xg_BvYJve9E91AZ?usp=sharing,https://github.com/pytorch/pytorch/issues/45465,5e828c98fc389b4a234802a0b0875eee0b7920a817b48c01ef8e68052f36bad4 references,issue,48549,issue,30786,medium,issue.comments[0].body,There is a highly-related issue with a solution: #30786 We can follow same pattern for all ops defined as built_in ops for tensor.py,https://github.com/pytorch/pytorch/issues/48549,b703f3158ef2d6741a26a30503b284a13a209f49fd26f4da697bd4a63ca271b3 references,issue,40107,issue,27647,medium,issue.comments[0].body,"he agent means we can remove the sync and join methods from the RpcAgent interface (I remember seeing an issue for that, it might have been #27647). A way to keep track of all the futures returned by the agent is to have the send method be implemented by the abstract base RpcA...",https://github.com/pytorch/pytorch/issues/40107,46efa19199b1005150fcd6f35207fa9cfddc4b6a46b266da29d87d35c8f4f7d7 references,issue,51136,issue,50444,medium,issue.body,"s object"" is really confusing. Can we say something like ""cannot invoke init() method of subclasses of torch.nn.Module classes""? Related to #50444 cc @gmagogsfm",https://github.com/pytorch/pytorch/issues/51136,ac17dba685e26f9891b9863169f677b602c96b66142f1513ceb3eb636033e2f5 references,issue,51136,issue,51140,medium,issue.comments[0].body,"This is actually same as problem 3 in #51140. Basically instances of classes cannot be scripted. After fixing that, the error becomes: Tried to access nonexistent attribute or method '",https://github.com/pytorch/pytorch/issues/51136,1336b615665a19986ad09a1f2fe2bec1f6061bc25f0f7a706023373ab15fa62e references,issue,51563,issue,50444,medium,issue.body,auto-inferencer should deduce that b is of type Optional[int] (by joining possible types of b along all possible execution path. Related to #50444 cc @gmagogsfm,https://github.com/pytorch/pytorch/issues/51563,05967fb9c7844255eefc15180960205f9e6ce5e3f24f23bd6d6ad8382d615175 references,issue,32306,issue,43867,medium,issue.comments[1].body,Seems raising same issue: #43867,https://github.com/pytorch/pytorch/issues/32306,c889e460bd33b994cf1ebcb761bf2462a1afd7edbdbcfdac0690e7b775dcd9d4 references,issue,51140,issue,50444,medium,issue.body,"ied_name raise RuntimeError(""Could not get name of python class object"") RuntimeError: Could not get name of python class object Related to #50444 cc @gmagogsfm",https://github.com/pytorch/pytorch/issues/51140,5977bb42d198d54464d08f59d647ad25ff58d00a67b1bc1a9941034f0a83297b references,issue,35666,issue,31557,medium,issue.body,"nder the hood that doesn't allow to swap epsilons easily (or some epsilon factory, but that's more complicated). Two related issues: #31829 #31557 cc @ezyang @ssnl @albanD @zou3519 @gqchen",https://github.com/pytorch/pytorch/issues/35666,852b3f44f5dc816325e4f3eef6d2631e598bfde548345c24ec8825f08edec835 references,issue,50678,issue,50533,medium,issue.comments[0].body,This is the same issue as #50533,https://github.com/pytorch/pytorch/issues/50678,cdec7ad54d8a6b8c81737634971aae4a8aa91418bf405d5952dc98dd770d18bb references,issue,49372,issue,30774,medium,issue.comments[0].body,"see #30774. Now that we have manual_cpp_binding, that should be doable through native_functions.yaml",https://github.com/pytorch/pytorch/issues/49372,dca73a67e8b51888f3445f6b4e139f5b0cd8c65d39003ec5fbcd233eee37e83c references,issue,37766,issue,12449,medium,issue.body,"orch/share/cmake/Torch"") here as well as here The CMake file needs to be clear in the documentation. It is causing trouble for many people. #12449 #21976 #806 #12449 It's not a bug. But a doc improvement issue.",https://github.com/pytorch/pytorch/issues/37766,4ab7ebc8ef01c61d0edc1baf25ff8645f4c3626fcf85cda35b29e085266263b8 references,issue,39279,issue,12659,medium,issue.body,"The goal is to facilitate the implementation of differential optimizer for libraries such as higher (e.g. section 4) Related discussions in #12659 and #32005, and Internal document cc @albanD @mruberry @vincentqb @egrefen",https://github.com/pytorch/pytorch/issues/39279,50116f9d31e5c94772c40531a5cc48c0ee821ab64af2a2f215f2c499d36b628a references,issue,39279,issue,32005,medium,issue.body,"to facilitate the implementation of differential optimizer for libraries such as higher (e.g. section 4) Related discussions in #12659 and #32005, and Internal document cc @albanD @mruberry @vincentqb @egrefen",https://github.com/pytorch/pytorch/issues/39279,0de52e0a3b3634d1cf5fd343a05b639253c890eb65b858139e8184c8a9f3895d references,issue,42815,issue,41486,medium,issue.comments[0].body,@colesbury had some detailed suggestions for how to debug memory leaks in #41486 (comment) that would be great to include in such a guide.,https://github.com/pytorch/pytorch/issues/42815,2879417a956477587eb630e3fbe00c9abdbcaae599ddf74882ded354eb3a5642 references,issue,49459,issue,48945,medium,issue.comments[1].body,Any chance this is related to #48945 ?,https://github.com/pytorch/pytorch/issues/49459,1fcd849a4451b8095be9d7b5ef8ae5071271d23af2f94943be85a46c137dac78 references,issue,49716,issue,49477,medium,issue.comments[0].body,@ngimel This one might be relevant to #49477,https://github.com/pytorch/pytorch/issues/49716,63e681cf144bf8a37b4b5ef81d7f8e3ff194f24705dfec6a4db1108067739a57 references,issue,49321,issue,23425,medium,issue.comments[0].body,Found a relevant question here #23425 @mrshenli @chenchr Can you help? Thanks.,https://github.com/pytorch/pytorch/issues/49321,9c743762b97164e868bd5972957edb9840971496cb55eeb95219f35d962252ea references,issue,48878,issue,47669,medium,issue.comments[1].body,"Related: #47669 Will be resolved in the upcoming 1.7.1 release. Also, you could try the nightly binaries.",https://github.com/pytorch/pytorch/issues/48878,c31bbaf5d2fe644c705816beb726cc4ac0b1ce568c9539176eb90446154a08be references,issue,28245,issue,23110,medium,issue.body,"🚀 Feature Context for Model Parallel: #23110 Motivation When applications are using complex distributed primitives like RPC, RRef and Distributed Autograd, debugging issues can be cumb",https://github.com/pytorch/pytorch/issues/28245,3e137732dad282df59b9bb508be0a92764b8f7586963321adafbfc7f3785dd5d references,issue,46564,issue,39351,medium,issue.comments[1].body,Is this the same as #39351?,https://github.com/pytorch/pytorch/issues/46564,2d060117f5cfd328e736a44cd805f565152624493f717db331010580c55a6557 references,issue,48548,issue,44533,medium,issue.comments[1].body,"We have multiple tickets for such problem. eg #44533 is under investigating now. I think we need to use overload for these overloaded functions, but I do not see any automated way to do it at",https://github.com/pytorch/pytorch/issues/48548,41886fdaa083b5801da95a1e583d7ea9fb98e2bc901d740940d30fececa9aab8 references,issue,48088,issue,13246,medium,issue.body,"Usually, this indicates problems with leaking memory (#13246) or some problems with ulimit and number of open files. If possible, it would be good if they printed something before they die (if possibl",https://github.com/pytorch/pytorch/issues/48088,9cf30a0f8969d3213da1c876abc886cd517ca08191f0edb66cdadc5e60198361 references,issue,47649,issue,46902,medium,issue.comments[1].body,duplicate of #46902?,https://github.com/pytorch/pytorch/issues/47649,95773833161d781145360c58debba08a5ddcafd6a02639457fd7c92d683e82c2 references,issue,41142,issue,1249,medium,issue.comments[0].body,#1249 Dice Loss issue here #35882 Focal Loss issue here,https://github.com/pytorch/pytorch/issues/41142,885949b9720e1f8c2d483fc36b52e76900f40e3af041c9b1fb1eb2ec4436a3cb references,issue,46724,issue,25150,medium,issue.comments[0].body,Possibly similar to #25150 ?,https://github.com/pytorch/pytorch/issues/46724,3de74f5323249173083614d3eb2ab75c9059eaabbfcf8d89e040354945a4b5ba references,issue,44026,issue,40373,medium,issue.comments[0].body,Related: #40373,https://github.com/pytorch/pytorch/issues/44026,d13ce080f18394f34587d242b67dafff7d3e80b1b7b2d967bdbfde00861e4a98 references,issue,35363,issue,29093,medium,issue.body,"st PyTorch release is still 2.4.8, i think it can be upgraded to 2.5.6 now. I can see that there is another open ticket for the same issue: #29093 To Reproduce Steps to reproduce the behavior: Checkout PyTorch source code. Build on a linux system with kernel that is below 3.9...",https://github.com/pytorch/pytorch/issues/35363,eae31a2a6b9107cfe96b4cfa7cbeebc491f6d05097fc26e055997daf066876b5 references,issue,46184,issue,41243,medium,issue.comments[0].body,"Before selecting new name, also consider the discussion in #41243.",https://github.com/pytorch/pytorch/issues/46184,ad17cf8216e409eec65e0aec46b9e813b06a50971eea29f13af6c9e0e38d4ac6 references,issue,46176,issue,25743,medium,issue.body,"rward and can be done once and stored. Additional context My suggestion draws from several other methods of handling this use case such as: #25743, https://github.com/pytorch/vision/blob/d5379656a098f0d69ad4dbe3dd94f2701415824c/references/detection/group_by_aspect_ratio.py#L23...",https://github.com/pytorch/pytorch/issues/46176,cdfc57d1856bc39a08ba503fbae56e9c389a6f542d2e255cca8857dbe175f455 references,issue,46176,issue,41292,medium,issue.body,"/github.com/pytorch/vision/blob/d5379656a098f0d69ad4dbe3dd94f2701415824c/references/detection/group_by_aspect_ratio.py#L23, and comments in #41292. Since torchtext and torchvision already have similar classes, it'll be a good and very practical addition to the DataLoader's cap...",https://github.com/pytorch/pytorch/issues/46176,6b8ad717e22caee4cc8172c14a5bf1fad9d2b09e239b36c21f36e07fb1e471ac competes with,issue,19685,issue,19637,medium,issue.body,"is is causing a bug where tracing will ""constant-ify"" device instead of correctly deriving the device at runtime from the input tensor (see #19637). I tracked the overload sorting code to here but I don't understand it. My desired behavior is to swap the order of the overloads...",https://github.com/pytorch/pytorch/issues/19685,4a1de276d5179a337c40767e8c67593a0ed9165b364263c1713fe83fa23f0c6a references,issue,21457,issue,24870,medium,issue.comments[0].body,"If I understand correctly, @andreaskoepf is suggesting at #24870 (comment) that a better downsampling grid_sample result could be achieved by filtering the image in different ways before sampling. This co",https://github.com/pytorch/pytorch/issues/21457,96c5f8139463cf4a29e51349b6c254065383fcc281d54ac3aaf4d91909386911 references,issue,21688,issue,17774,medium,issue.comments[1].body,"mention, but have used. I find that having too large a batch size causes errors in conv2d. This is an open issue, and I am linking it here. #17774",https://github.com/pytorch/pytorch/issues/21688,c2c992c5bd78b9a35b62faa61bc200c58e772d9bc08e8a6733a596ecbfec6f27 references,issue,25039,issue,21457,medium,issue.body,"sampling points will in general not be a regular grid. This might require some research and definition design, and discussion is ongoing in #21457. lanczos: Another option for smoother interpolation. This is currently implemented in Torch, so a PyTorch implementation could pos...",https://github.com/pytorch/pytorch/issues/25039,56b10e4216fca56c50b9fe989cbc639dd1bb6136bb52e515689dbd5773a2f2b4 references,issue,25039,issue,24870,medium,issue.body,"Note: This issue is expanded out of #24870 to allow more room for discussion. 🚀 Feature Currently, torch.nn.functional.grid_sample() supports two interpolation modes: bilinear and ne",https://github.com/pytorch/pytorch/issues/25039,6641e08a58617859d167197d1e3f8ce9116f29fe2b22c0ef2e4f50bbe476e609 references,issue,31458,issue,18095,medium,issue.body,"xx. But this is not true currently. torch functions might have (surprisingly) different arg spec compared to Tensor functions. For example, #18095 indicates that Tensor.flip allows integer dimension while torch.flip only allows tuple dimension. There might be more inconsistenc...",https://github.com/pytorch/pytorch/issues/31458,517cff3571da78717e7a4ecdf6a1fd845c5c7a6c7eff332b57f7ee831a5ee390 references,issue,31474,issue,18634,medium,issue.comments[1].body,"sted on the CPU, it adds time over the original due to the call to torch.cat. (Apparently the speed of cat on the CPU is a known issue, see #18634.) I mention this as when the speed issue with cat is resolved, this version might be a better choice for computing the total norm....",https://github.com/pytorch/pytorch/issues/31474,903787abf6d2e4264a1802e18a976bf39b74cf56a66cf60d905d63e887e8dd66 references,issue,34544,issue,32918,medium,issue.comments[1].body,Maybe related to #32918. To make sure though: is content.png a 256x256 image?,https://github.com/pytorch/pytorch/issues/34544,4f67dc97ab137251ad80bc6488e828fcd7310fd4ef77281d425e6cde3de51c04 references,issue,34675,issue,28733,medium,issue.body,"Issue description @soumith, @VitalyFedyunin, @karimhasebou, I was investigating slow maxpool2d on cpu issue #28733. I thought I could use the conv2d implementation, but two important functions unfold2d_acc, unfold2d_copy (https://github.com/pytorch/pytor",https://github.com/pytorch/pytorch/issues/34675,afeeba8220cb346802bd0bd40ca3b75915df59a7077c74b7af413228bafb4ebe references,issue,31913,issue,30574,medium,issue.body,"ed tensor. In this case, the 1D index_add_ becomes useless and explicit advanced indexing will waste time doing copy. I see a similar issue #30574. This is kind of different compared to that because that is more-or-less an api that converts advanced indexing to proper language...",https://github.com/pytorch/pytorch/issues/31913,0baa3723085a6fe8013b2ee7c93a1d092af9921a37e24711c7530110c409b4a1 references,issue,45851,issue,44991,medium,issue.comments[0].body,Related about unfold: #44991,https://github.com/pytorch/pytorch/issues/45851,5133bb19305acffdd89cc6d26a2081aec62de36337ead9f05b98734be963c1bb references,issue,13218,issue,5580,medium,issue.comments[0].body,(was also tracked at #5580),https://github.com/pytorch/pytorch/issues/13218,eb23ae78876d3d1ab72a68544bf86e06e9d005053aab856cdd2c2ca8a150ae1b references,issue,29116,issue,29028,medium,issue.body,"Following up on #29028 (comment), there exists torch.masked_fill_ but cannot be accessed through out argument, which makes writing generic (inplace / out-of-place",https://github.com/pytorch/pytorch/issues/29116,3a1b4fba43cad8463e052fa55c1b30828316896aeee0085778bb6d107fb0904c references,issue,43125,issue,41625,medium,issue.body,"UPD: Original title: ""slice(None) and maybe slice in general is not supported in JIT"" Originally reported in #41625 (comment) def tensor_dim_slice(tensor, dim, dim_slice): return tensor[(dim if dim >= 0 else dim + tensor.dim()) * (slice(None), ) + (dim_sl",https://github.com/pytorch/pytorch/issues/43125,57fde397a6333cf845b7a84b023d3504eb40d1f7079515e72ca6f754db87b7ad references,issue,45901,issue,30387,medium,issue.comments[0].body,"Related: #30387, I used this to convert a byte array to a tensor, but this is super hacky",https://github.com/pytorch/pytorch/issues/45901,477149b4a28d57ad397d426bc39b150c63f89bebe0aa4f1d76ee0c7f0b658be4 references,issue,2129,issue,31945,medium,issue.comments[0].body,"ue should be reopened, since the PR was not merged in the end and the truncated normal is still not available in PyTorch. @soumith Related: #31945",https://github.com/pytorch/pytorch/issues/2129,09124e7cf67c4d4b96efc28a33679be330574d0fbcf1e81cc7ed563ba22f0af9 references,issue,45208,issue,44768,medium,issue.comments[0].body,"Using torch.no_grad is only supported in with statements, not as a decorator. There is an open issue I am working on (#44768) to fail better if it is used as a decorator. Can you use a with statement for whatever it is you're trying to do?",https://github.com/pytorch/pytorch/issues/45208,29a8aba73dbe788af2a6719f9a2982a0d3b448578450250991c7f4867a86ccaa references,issue,43502,issue,33296,medium,issue.comments[1].body,Related #33296,https://github.com/pytorch/pytorch/issues/43502,cdd1cf8dc637a4e76595b1f666c48405ee1057fbcc29e6b851bb6d823c6c3b64 references,issue,43012,issue,35642,medium,issue.body,This is follow-up of #35642 proposal and discussion under #39274 implementation. Proposed API: Option 1: c_map_dataset = CachingDataSet(map_dataset) Option 2: c_map_da,https://github.com/pytorch/pytorch/issues/43012,87be6e0353c042145471d3a0de2bba510fa049d2b6bfb2fc1bf3c4701ce37a28 references,issue,43453,issue,31528,medium,issue.body,d. That can be a problem for third party libraries that expect CUcontext exists. To Reproduce One issue with device guard was reported here #31528. Device guard in progress thread of mpi process group doesn't initialize cuda context because of described behavior and that leads...,https://github.com/pytorch/pytorch/issues/43453,08345d15b1387f5ebdaaa121fb1b6d22c99dcc0d775fe362997ce388e2d48749 references,issue,43459,issue,21700,medium,issue.comments[1].body,#21700,https://github.com/pytorch/pytorch/issues/43459,928848dacdb430225e3cf055b323e710815cff804b94c9cd66b0ee44d0710822 references,issue,43116,issue,29973,medium,issue.comments[0].body,This is a duplicate of #29973. Also interesting that a+b is 2.5x times slower in pytorch.,https://github.com/pytorch/pytorch/issues/43116,969653541fe6451ca641744083e3cc1651800ae3380e0fddcb7e65b372cd469c references,issue,43115,issue,2576,medium,issue.body,"llapse to be the same value. We disallow more than 2**24 categories for that reason, but large inaccuracies can start before that. Related: #2576 cc @vincentqb @fritzo @neerajprad @alicanb @vishwakftw",https://github.com/pytorch/pytorch/issues/43115,f654c05c7f373c958170268082de401f8fe79e6380fa39f8331c37a5bdf6da65 references,issue,41400,issue,41186,medium,issue.body,🐛 Bug Various tests from test_jit.py fail on Power. This might be related to #41186 as some tests have similar names to the ones failing there. The failing tests are: test_solve_batched_broadcast_A_cpu (main.TestJitGenerate,https://github.com/pytorch/pytorch/issues/41400,f3bd537df779a4568c1cbe347ba1c4ea27fc65f5987c15b2081e6738f68b328f references,issue,41400,issue,41186,medium,issue.comments[1].body,might be a good idea to have an x86 CI testing this configuration (but on x86). The reason is that the differences I see in this issue and #41186 are so big and especially with the test inter-dependencies (failures go away if tests are run individually) there might be a real b...,https://github.com/pytorch/pytorch/issues/41400,e1a7fa72c61d30afeb25bf5d008fc4fea4f3d65b8fd138d8739461bd1553a0bd references,issue,42350,issue,19053,medium,issue.body,"t every call, To compact weights again call flatten_parameters(). (_cudnn_impl at.. .. \ aten \ SRC \ aten \ native \ cudnn \ RNN CPP: 1266)#19053 I tried, but I don't work.And the model architecture of CRNN is the same as in the above questions, so is the program for PTH conv...",https://github.com/pytorch/pytorch/issues/42350,8c7d554403fdbe5a7ccd288a1fa66d03ff09c5d0bccb272995d3ded4ca346b05 references,issue,42295,issue,42316,medium,issue.comments[0].body,Related to #42316,https://github.com/pytorch/pytorch/issues/42295,4e0e2fcd92a523bd16ab56af664fb035a1b0f8536a4b06b0de952a4c7e7b1bab references,issue,41243,issue,41081,medium,issue.body,"I proposed it in #41081 (comment), but maybe a separate issue is a better place to discuss this: A more radical proposal: unite BatchNorm*d / SyncBatchNorm in one",https://github.com/pytorch/pytorch/issues/41243,a7c71486e65cf815317320c64c5a357fafebb2405368fda330d4cf3621b199d9 references,issue,41243,issue,2628,medium,issue.comments[0].body,Related to #2628 for BatchNorm,https://github.com/pytorch/pytorch/issues/41243,83426afc562fd1779fd0b56e8f6f5a5794cc6c1d01375dceb5701a349c2ec7c9 references,issue,13023,issue,7359,medium,issue.body,"s optional instead. Additionally, the max task number (currently 2 * num_workers) should also become configurable. Expose Sampler iterator (#7359) This would enable dynamic updates to the Sampler iterator states, e.g., dynamic reweighting of the samples. The API may be loader_...",https://github.com/pytorch/pytorch/issues/13023,11d0e10d7421b318da3b23ddbf387147cb7a24801e3302b83429fa8dc2f6174f references,issue,35446,issue,34361,medium,issue.body,"1.0): 1.5.0a0+efbd6b8 OS (e.g., Linux): Fedora 29 Python version: 3.7.6 CUDA/cuDNN version: 10.1 Additional context This may be related to #34361. The test failure may also be related to legacy profiling executer. By commenting out the following line, the legacy test will also...",https://github.com/pytorch/pytorch/issues/35446,ee535e6c4d29692c26aef0550e495090e978be56a0a157deda5388ebd1bb972a references,issue,4181,issue,4145,medium,issue.comments[1].body,May #4145 be a candidate for this refactoring process as well?,https://github.com/pytorch/pytorch/issues/4181,d447b4a41164a7b64db969eb7f42ddaaf81ed23564a66a0900ce6029babe8748 references,issue,39947,issue,15004,medium,issue.body,"re/init.h"" int main(int argc, char** argv) { caffe2::GlobalInit(&argc, &argv); return 0; } I used a cmake installation of PyTorch following #15004. I seem to have no problems creating the make file with cmake, but when i try to compile with make i get the following error messa...",https://github.com/pytorch/pytorch/issues/39947,0ca147bfca08d7cabd358f4f598cbc388c69cce8dc3183dc6289ec0db73eaf8d references,issue,39947,issue,11071,medium,issue.comments[0].body,"it's worth mentioning the environment you are using, e.g. how did you specify lib path? Similar issues are here #11071 (comment), could you try that?",https://github.com/pytorch/pytorch/issues/39947,7edfaf0835c04056cebc667cf8eca460bb3c297aa3c7c13634e28541f2d8bbb9 references,issue,32078,issue,19092,medium,issue.body,"In the memory layout work in #19092 we have a problem where strides undetermine the layout of a tensor when size = 1 or 0. In particular, if I have a tensor with size (1, 1, 1",https://github.com/pytorch/pytorch/issues/32078,87e9c9f070456d9eabaaf265ca80245f800d984ab381f6c56febd6cb43e74a50 references,issue,39836,issue,38010,medium,issue.body,is is done by adding a summary table in the main page that links to subpages for each function/class. More details can be found in this PR: #38010 Known Issue: New format generates new urls which would break existing documentation links. Proposed Fixes: #39086 #39032 Feedback...,https://github.com/pytorch/pytorch/issues/39836,3dc085923ebcdcc1e3e88e5cd53ebaef430fed9875075b1f2eb2f46fb9feac88 references,issue,30373,issue,27610,medium,issue.body,".. easier. While searching for possible earlier reports for this, I found the following issue to be related (although much more ambitious): #27610 Pitch A minimum viable refactoring answering the issue could consider the atomic steps: a. resolve github account name to local di...",https://github.com/pytorch/pytorch/issues/30373,6e5c12d7e6cb3c5549231cd3364d903bb913f0dc618d8e488af8a440d7dfa68f references,issue,38614,issue,24904,medium,issue.body,"only passed in one GPU, it would not throw that error. My original error when passing in multiple GPUs seems related to #30459, #28206, and #24904, although I'm not sure if the error in from this code snippet is related. cc @suo",https://github.com/pytorch/pytorch/issues/38614,2aa68486c09ff7a686d5785908ae03a8e107d24ac9487a55a90e92505d5a267e references,issue,38614,issue,28206,medium,issue.body,"t, but if I only passed in one GPU, it would not throw that error. My original error when passing in multiple GPUs seems related to #30459, #28206, and #24904, although I'm not sure if the error in from this code snippet is related. cc @suo",https://github.com/pytorch/pytorch/issues/38614,d6ab7cd389e0a4b2ed096b113950d6db84d77654fb17f84299306d4d785a2fbd references,issue,38614,issue,30459,medium,issue.body,"gradient, but if I only passed in one GPU, it would not throw that error. My original error when passing in multiple GPUs seems related to #30459, #28206, and #24904, although I'm not sure if the error in from this code snippet is related. cc @suo",https://github.com/pytorch/pytorch/issues/38614,0c46de27677ce053e8d86f7413f055149adf08ea5f6b10d76d06a79a0d91e678 references,issue,38975,issue,29548,medium,issue.comments[0].body,"Yeah, basically we should do the same thing as convolution_overrideable in this case. Or move to the glorious new world described in #29548",https://github.com/pytorch/pytorch/issues/38975,9acc0bb7a6fe6497b98b0366c5ff5a778ad624837b142228939b4eaabd24b45a references,issue,10119,issue,7807,medium,issue.body,"put - none of these combos work. I've seen other posts on this issue such as: facebookarchive/caffe2#1074 facebookarchive/caffe2#2282 #9736 #7807 facebookarchive/tutorials#7 But most of these issues were either older, posted by those who built from source or had different envi...",https://github.com/pytorch/pytorch/issues/10119,5a29f652a436a07eb663fb4f72fb243e6efe8d8086f09254d92de3b8b0fb58a4 references,issue,37837,issue,32078,medium,issue.body,"ate operators coverage tests Publish operators conversion issue to tracker, extend documentation if necessary . Add support fo permutations #32078 Write in details behaviour of contiguous(), IPC, Serialization Hide code behind #ifdef guards cc @VitalyFedyunin @jamesr66a",https://github.com/pytorch/pytorch/issues/37837,6160adaa0ababeb923752a4b0e9e2229916fb4f1af953ea3216dc627a3a61128 references,issue,38703,issue,13918,medium,issue.body,"Per title. Reported internally. See related (but distinct) #13918, which discusses conversion from lists of arrays. cc @mruberry @VitalyFedyunin @ngimel",https://github.com/pytorch/pytorch/issues/38703,1da8e8a7c69670c9892812b5f70ee28c77ccc3c8caffb7ccf218b3b5a0f9e3d0 references,issue,29816,issue,29814,medium,issue.body,"ta.to_dense()) # Exponential moving average of squared gradient values state['exp_avg_sq'] = torch.zeros_like(p.data.to_dense()) Similar to #29814, changing: p.data.add_(make_sparse(-step_size * numer.div_(denom))) to: with torch.no_grad(): p.add_(make_sparse(-step_size * nume...",https://github.com/pytorch/pytorch/issues/29816,f0d1f82cec7dfee45dd29b3bf25184ded600feec322b80de56ff67acb19605ce references,issue,33867,issue,13402,medium,issue.body,"cate a new tensor to make the contiguous, and thereby end up losing updates to the running mean/var entirely. Discovered while looking into #13402 cc @csarofeen @ptrblck",https://github.com/pytorch/pytorch/issues/33867,a251c724f7d96e7529979df05b2f8bbb83c5d0672edeb1d3fe990f40cfd2a73d references,issue,37529,issue,28245,medium,issue.body,"e also plan on having a metrics handler for RPC that reports metrics at a predefined interval, for which this would also be useful for (see #28245) cc @pietern @mrshenli @pritamdamania87 @zhaojuanmao @satgera @gqchen @aazzolini @rohan-varma @xush6528 @jjlilley @osalpekar",https://github.com/pytorch/pytorch/issues/37529,16bf93aeeea7f8c809595a265231da88a5861a5218d1fa5531742a313e5b53d3 references,issue,30929,issue,32407,medium,issue.comments[1].body,"to the somewhat different list in cmake/Modules/FindBLAS.cmake (from which you can choose with a different variable). Note that because of #32407, it is not currently possible to specify Eigen and keep another BLAS out of the build. It accepts the Eigan preference, adds the he...",https://github.com/pytorch/pytorch/issues/30929,701994a53a980395f7f63a8f6008df34ec1223412df34b647032b70174b57c94 references,issue,36999,issue,35678,medium,issue.comments[0].body,seen other build failures as available build flags or combinations of flags changed. I think this ticket was related to one of those cases: #35678,https://github.com/pytorch/pytorch/issues/36999,63955803d3d16b93cf86298e73fdbb96a518f56e0915e363aa0cb31a15206026 references,issue,23110,issue,26759,medium,issue.body,"inputs) bw_ctx_id = dist.autograd.backward(loss, timeout=60) # timeout of 60s optimizer.step(bw_ctx_id) RRef (more details are described in #26759) RRef is an important concept for building a distributed autograd graph. Each RRef is owned by a single worker (i.e., owner) and c...",https://github.com/pytorch/pytorch/issues/23110,4fa4bbaf7095fcda26b65588792142aa7b8f08d521e900afac08cf93c1ff164e references,issue,33227,issue,18173,medium,issue.body,tes/extending.html#extending-torch-autograd) How to use at::Tensor::register_hook How to compute higher-order gradients in C++ (e.g. issue: #18173) Blog post: C++ frontend revamp (ETA: 4/3) cc @yf225 @gchanan,https://github.com/pytorch/pytorch/issues/33227,ac6c85e80e97df3b93479d00e3da75aec014e437f10405c3b88fb2bff4a9c867 references,issue,33227,issue,25883,medium,issue.body,ill provide the following improvements to the PyTorch C++ frontend: Python/C++ API Parity: torch.nn modules and functional (tracking issue: #25883) RNN (#34322) LSTM (#34322) GRU (#34322) RNNCell (#34400) LSTMCell (#34400) GRUCell (#34400) AdaptiveLogSoftmaxWithLoss (#29076) P...,https://github.com/pytorch/pytorch/issues/33227,e7a91d86b55a45a1fdd8998c17e715cfc91415ed3827e201e4a3ebb23b05d87d references,issue,33227,issue,28440,medium,issue.body,"ossFuncOptions -> MultilabelSoftMarginLossFuncOptions (#35163) Turn on parity test (#35189, #35190) torch.optim optimizers (tracking issue: #28440, make sure to test that the new design doesn't break serialization BC) tensor multi-dim indexing API #32841 #30426 #30427 C++ tens...",https://github.com/pytorch/pytorch/issues/33227,1279678dc3eb57a468c5c6a9cbd854d74e9c180dcaf99bc3a532f7ecfe70ef0e references,issue,32851,issue,28743,medium,issue.body,"==0.2.2.post3 [conda] Could not collect Additional Context Probably that Sampler is not intended to be used in the way described above (see #28743), but since the code works and produce unexpected results, it was hard to identify what we did wrong when we stumble on this. cc @...",https://github.com/pytorch/pytorch/issues/32851,a336b9106709dc8227c1cc0de8dac1e0e16d3999e99adcfbb0107f7a893d74b8 references,issue,34649,issue,34646,medium,issue.comments[1].body,"My usecase is this: #34646 I tried a few things to simplify the code. One of them was creating a typed storage, and then trying some generic method to create a tensor",https://github.com/pytorch/pytorch/issues/34649,516a9111255c71564882631d4732242204c43bfa089edbb25a3636a52de7b55e references,issue,24915,issue,12672,medium,issue.body,5 wants to re-use worker processes FastDataLoader python 3.8 shared memory Internal: torchdata gil experiment DataLoader+Iterable Features: #12672 wants to move collate_fn functionality to datasets #26547 wants distributed random sampling #28743 for sampler for iterable datase...,https://github.com/pytorch/pytorch/issues/24915,938834d0e4f98af2648957dde5238091eba4fdf828eb6a453fdcd7c54c0b687d references,issue,24915,issue,28743,medium,issue.body,experiment DataLoader+Iterable Features: #12672 wants to move collate_fn functionality to datasets #26547 wants distributed random sampling #28743 for sampler for iterable datasets pytorch/vision#1315 wants to apply an instance of random transform sequence to many images cc @s...,https://github.com/pytorch/pytorch/issues/24915,803d1f7d5d609cb536cd532daa4aedc6b6b83cfe884bdbaf302f915a23ceed72 references,issue,30635,issue,15421,medium,issue.comments[1].body,"t-in-evaluation-mode-using-multi-gpu-learned-model/33280/4 -- is tracing inside of the module still the recommended route here? Referencing #15421 #16891 #17540 ... the last one in particular might be the most relevant to the eventual solution, but I'm not sure.",https://github.com/pytorch/pytorch/issues/30635,4a7a6a1b4cb9ad9daaaab3d8b5d8e23a62733dd95176e4730d79920ceb656f0c references,issue,32544,issue,30965,medium,issue.comments[1].body,This should also fix #30965,https://github.com/pytorch/pytorch/issues/32544,83c09201cacb83f4060b4510e304842f287e90ec9abe2c7c0900933dff2dfb0b references,issue,32463,issue,30421,medium,issue.body,"current implementation python objects are copied between on the boundary between the JIT and python. This results in issues like #31129 and #30421, and we've had internal reports of users passing a python dictionary into a mutating method. Example x = [1] def foo(x: List[int])...",https://github.com/pytorch/pytorch/issues/32463,b7024fed3b08c950b76167affe2fc5d3dd5b5ea559546b767dd5b91a32b92ff6 references,issue,32463,issue,31129,medium,issue.body,"cts In our current implementation python objects are copied between on the boundary between the JIT and python. This results in issues like #31129 and #30421, and we've had internal reports of users passing a python dictionary into a mutating method. Example x = [1] def foo(x:...",https://github.com/pytorch/pytorch/issues/32463,747bab5dd12797e2f3bac6434783606c29345ddbb311a9a71957fd880b63085d references,issue,30138,issue,24243,medium,issue.comments[0].body,Dup of #24243. cc @smessmer,https://github.com/pytorch/pytorch/issues/30138,a21d5288d3e6ac80b33c8caa35f7a012495ae6015dc4c4e518c2d85acc8c5038 references,issue,27479,issue,25267,medium,issue.body,e.g. #25267 And other internal reports cc @suo,https://github.com/pytorch/pytorch/issues/27479,b6cae8987e5130155569ed9058c12af5855724bd6dceae051d9dcf68c34d308c references,issue,27479,issue,1529,medium,issue.comments[0].body,also related to #1529,https://github.com/pytorch/pytorch/issues/27479,dee8027178530816d44b8dec2664e849d9a96a8a8360cfb35eaf9341a2c63429 references,issue,30049,issue,25267,medium,issue.comments[0].body,It's possible this is related to #25267 But it could also just be a DataParallel & JIT issue. cc: @gqchen,https://github.com/pytorch/pytorch/issues/30049,151ebae7fec8919a954630139453341f6ad18fd86ee523e97ebff454d16bc451 references,issue,30049,issue,25267,medium,issue.comments[1].body,@eellison My impression is that this is unrelated to #25267 given that I am not using MKL and I only encounter the CPU memory leak when the model is on GPU. Is there anything I can do to help you dia,https://github.com/pytorch/pytorch/issues/30049,511dd278dc79b90fd2090a798b8f5d1e33208ae309d2fc5a1866f0efcc0e2426 references,issue,29987,issue,25591,medium,issue.body,"de. Usually the clients only have CPUs but the servers have GPUs. As we haven't been able to directly serialize/deserialize the IValue (see #25591 for context), a workaround is to leverage the serialization methods of Module as follows: 1) create a container module; 2) pack th...",https://github.com/pytorch/pytorch/issues/29987,802dcadbbcb1dfd6408dcd4fe1a5633dc90026454beacd394047efe211152511 references,issue,27475,issue,27144,medium,issue.body,test the design develop a plan to deprecate the old methods on script::Module and implement them in terms of the new methods. For example: #27144 cc @suo,https://github.com/pytorch/pytorch/issues/27475,3c66d7d2f0dfca8249fa113c1da6b5d8999d66771e2e84ab1f0cc1e51b87a618 references,issue,33691,issue,33670,medium,issue.body,"rim::Constant"" } ], ""versions"": { ""producer"": 22 } } Environment Not important. Additional context Looks like this parse function is buggy. #33670",https://github.com/pytorch/pytorch/issues/33691,72eb074abb1cab8b27ad45cf6a481682e4e3251e629bc746197658df13b6ec70 references,issue,33670,issue,30812,medium,issue.body,"hich might be worth investigation. Or it's by design, in which case it should be fixed in support for tensorboard. This issue is similar to #30812, but I double-checked: it's a different issue.",https://github.com/pytorch/pytorch/issues/33670,2f2ca503a7582e6ff5d43b67c3e152ae39bfda17a3b899ea6e7da42c479f53aa competes with,issue,31708,issue,17897,medium,issue.body,"mory access. Additionally, if v is allocated separately rather than being sliced, it will not produce the error. This may be a duplicate of #17897 or just a similar underlying CUBLAS issue. This script can succeed if N*K > 2^31 - 1 (the commented N = ... line is bigger but sti...",https://github.com/pytorch/pytorch/issues/31708,1d2d35d31c86ff8c8d9e5f35e7772881df1f05ca1f660b5ba79157f47efe7020 references,issue,30291,issue,23756,medium,issue.body,"ng errors out. It would be nice to allow save/restore tensor versions when it's needed. Currently tensor._version is not writable. (related #23756, cc @ezyang @ssnl @albanD @zou3519 @gqchen)",https://github.com/pytorch/pytorch/issues/30291,bc6c416a7199d93a450e606efc20a9c8b9fad1da77010842e8807ae28237681c references,issue,33034,issue,30421,medium,issue.comments[0].body,Similar issue to #30421 and #31129. One fix here is everytime we invoke torch.jit.script to memoize the scripting of any mutable object.,https://github.com/pytorch/pytorch/issues/33034,107e2a8f74f58b59ed54979e67d1cc8a213c3feeae2281ca10b3c3a208101857 references,issue,33034,issue,31129,medium,issue.comments[0].body,Similar issue to #30421 and #31129. One fix here is everytime we invoke torch.jit.script to memoize the scripting of any mutable object.,https://github.com/pytorch/pytorch/issues/33034,615f5fdb1a35f1ef6d8535e93e7a5e09876b1645a6fd4fcce177355a147bd7fd references,issue,31907,issue,30903,medium,issue.comments[0].body,#30903 might be related,https://github.com/pytorch/pytorch/issues/31907,2928ad6e5302f17f8c27cc2f38a494c90d31a123e8833451b237a246e2454ae6 references,issue,31356,issue,30803,medium,issue.body,"when all the threads running for all parallel_for tasks(). Having some light operation running at few threads doesn’t help the performance. #30803 To further tune the performance, you may consider the following command line options and software configuration. Export KMP_BLOCKT...",https://github.com/pytorch/pytorch/issues/31356,52035fb9c8c18615ab4b34374c151684b5ba9558f19fe35f8fc96541c766ee22 references,issue,19668,issue,88,medium,issue.comments[1].body,000004ea137 in PyCFunction_Call () at ../Objects/methodobject.c:98 #87 0x00000000005c20e7 in PyObject_Call () at ../Objects/abstract.c:2165 #88 0x0000000000534870 in PyEval_CallObjectWithKeywords () at ../Python/ceval.c:4580 #89 0x0000000000539bfb in PyEval_EvalFrameEx () at ....,https://github.com/pytorch/pytorch/issues/19668,cb4f4609f43567da2959d317a86162166266462e9bd6356401a1aa8ce3a9f6e2 references,issue,31228,issue,30633,medium,issue.body,"uted.rpc, we will need to support packing share memory relating info while RPC send() pickles nn.Module. There is a hacky implementation in #30633, but it relies on the special reduce functions mentioned above that are supposed to work with Python's ForkingPickler. We will nee...",https://github.com/pytorch/pytorch/issues/31228,ef421640e4bd427262408f83656e7bd9d5d7634c68f79fdb4f5fa59b3f4e66d0 references,issue,31353,issue,31252,medium,issue.comments[1].body,"This is a very close duplicate of #31252, and there is a discussion going on already there. TL;DR it's complicated.",https://github.com/pytorch/pytorch/issues/31353,5fc336d2883080ea1c0c426a7a95adf4b23312b0021c83a25c93aedf5d0a74e8 references,issue,31007,issue,30532,medium,issue.comments[1].body,"Does pytorch itself work in your environment? GTX780 is compute capability 3.5 IIRC, and support for it was discontinued, see also #30532",https://github.com/pytorch/pytorch/issues/31007,fd038afa734ebbb2a4b5da224f914523ffaf9f5eb8be2093e5a01017fed08849 references,issue,26759,issue,23110,medium,issue.body,"With @pritamdamania87 @gqchen @aazzolini @satgera @xush6528 @zhaojuanmao Master Design Doc: Distributed Model Parallel Design: #23110 Main RRef PRs: #25499 #25169 Background RRef stands for Remote REFerence. Each RRef is owned by a single worker (i.e., owner) and can be us",https://github.com/pytorch/pytorch/issues/26759,62ded71fe76aa6f3efa09cf752e0aa2d806b159b6bc061f62ababd17cc06ee80 references,issue,29739,issue,29093,medium,issue.body,"🐛 Bug Since I can't use conda gcc 7.3 (#29093), I tried to build master with system gcc 7.4 and met CMake Error at third_party/fbgemm/third_party/asmjit/CMakeLists.txt:100 (target_compi",https://github.com/pytorch/pytorch/issues/29739,d04981a84438f017c8188834534ac6b4d0109af0743b2218c33aa2756a7f00e6 references,issue,27099,issue,23110,medium,issue.body,"urrently, dist.rpc and dist.remote only matches arguments with exact types. We should support implicit RRef type conversion as described in #23110 (i.e. T->RRef[T] and RRef[T] -> T). cc @pietern @mrshenli @pritamdamania87 @zhaojuanmao @satgera @rohan-varma @gqchen @aazzolini",https://github.com/pytorch/pytorch/issues/27099,b9ef4b287032ac03a6f3a0e1cccc1ea7e64e828645cb6a14d09ef04533302db9 references,issue,12983,issue,9484,medium,issue.body,"I tried deleting the previous ""build"" directory and regenerating all files, but the problem still exists. There are similar issues such as: #9484 #9604 facebookresearch/video-nonlocal-net#6 facebookresearch/Detectron#370 facebookarchive/caffe2#2513 The error messages reported...",https://github.com/pytorch/pytorch/issues/12983,8737823f7ec26eb24259315d094fb6b8743232f8a594482c6a4984e4a4950062 references,issue,12983,issue,9604,medium,issue.body,"d deleting the previous ""build"" directory and regenerating all files, but the problem still exists. There are similar issues such as: #9484 #9604 facebookresearch/video-nonlocal-net#6 facebookresearch/Detectron#370 facebookarchive/caffe2#2513 The error messages reported in the...",https://github.com/pytorch/pytorch/issues/12983,0b34482eb8e206b3aed9868bc4a81a04111eceb7382330264a9934739e09a4b2 references,issue,24478,issue,24498,medium,issue.body,"pare.cpp:18 pytorch/aten/src/ATen/native/TensorCompare.cpp Line 18 in eabfca3 at::CPU_tensor_apply4( #24498 Migrate CPU_tensor_apply to TensorIterator in aten/src/ATen/native/TensorCompare.cpp:30 pytorch/aten/src/ATen/native/TensorCompare...",https://github.com/pytorch/pytorch/issues/24478,22c7d04e0f0d7f6eed9af882ac9b8145dff0a6604100bcfb29e00efe0d22c83c references,issue,24090,issue,19092,medium,issue.comments[0].body,Keeping track of this by cc'ing #23403 #19092,https://github.com/pytorch/pytorch/issues/24090,7795c801e07985c9c4e922e2e0a33f130ebf24aab8e8d1a85396c12a9e7a6aec references,issue,9310,issue,9204,medium,issue.comments[0].body,similar to #9204,https://github.com/pytorch/pytorch/issues/9310,1f36f3e4d1044be90f2bf690008102b6d03d949c64fa0e37f210ac5fea72b537 references,issue,25014,issue,24470,medium,issue.body,"Similarly to #24470, the cuDNN affine_grid_generator should be benchmarked against the native CUDA version. The dispatch to cuDNN for affine_grid was disabled",https://github.com/pytorch/pytorch/issues/25014,075996986766be4b26d35a6d5c2cdb30cbeaa842592c5ab05153b8ecf1a4f879 references,issue,24498,issue,24478,medium,issue.body,", scalar_t, scalar_t>( How to use TensorIterator: https://github.com/pytorch/pytorch/wiki/How-to-use-TensorIterator Additional Instructions:#24478 Blocked by:#24472",https://github.com/pytorch/pytorch/issues/24498,9ead9dfc2e7eed346bd3009ea1c6107d0481c6b0d4b9766dfdd5c21e0a26e8b6 references,issue,23490,issue,10950,medium,issue.body,"🚀 Feature support python sub-interpreters and maintains all status of the torch library. Motivation as #10950 demonstrates, the current torch library cannot lives on multiple sub-interpreter simultaneously within the same process. But we do need to",https://github.com/pytorch/pytorch/issues/23490,3dc37e7160e41c407e92cad2e747bcce2838aca270b236a4e2dee9f9015671cc references,issue,23240,issue,88,medium,issue.body,ul 23 16:50:24 #87 0x55da327b053a in _PyFunction_FastCall /tmp/build/80754af9/python_1546130271559/work/Python/ceval.c:4933 Jul 23 16:50:24 #88 0x55da327b053a in fast_function /tmp/build/80754af9/python_1546130271559/work/Python/ceval.c:4968 Jul 23 16:50:24 #89 0x55da327b6504...,https://github.com/pytorch/pytorch/issues/23240,c908f71bf0502fcee97d5d3852c25ba3e9c7d88f661f7e3eb05631e56fa930c6 references,issue,22414,issue,21731,medium,issue.comments[1].body,"ahead is fast, we can increment state before launching a kernel, lock scope can be reduced and multithreading performance will be improved (#21731) CUDA kernels uses Philox, which means, we should be getting consistent distributions between both of these backends Cons: CPU dis...",https://github.com/pytorch/pytorch/issues/22414,fca57bb9adb54d49f100ae6cd8f3ba2bfd26a85dee51d87661a505b1eba2fb6c references,issue,12498,issue,8853,medium,issue.comments[0].body,ul on the way to make nn.Linear work for sparse. Can I ask what's your use cases? Please use #10043 for request on sparse. Ops in progress: #8853. Current state of sparse: #9674,https://github.com/pytorch/pytorch/issues/12498,d12bbe0215b8fbac59f204abd44d9aef7d4b642f510dc094aad5fc2fea37f5ac references,issue,12498,issue,9674,medium,issue.comments[0].body,work for sparse. Can I ask what's your use cases? Please use #10043 for request on sparse. Ops in progress: #8853. Current state of sparse: #9674,https://github.com/pytorch/pytorch/issues/12498,c037e55f86542c9aef5ae3247903bad38eaaedc59778911e19f6728a53b2dbc3 references,issue,12498,issue,10043,medium,issue.comments[0].body,kward() We are planing (#12308) to support matmul on the way to make nn.Linear work for sparse. Can I ask what's your use cases? Please use #10043 for request on sparse. Ops in progress: #8853. Current state of sparse: #9674,https://github.com/pytorch/pytorch/issues/12498,c45faf6e843777df4588bebeb5964a618a4f8fc4980cdd3d5b9e26ba3363d2b8 references,issue,12498,issue,12308,medium,issue.comments[0].body,"h.FloatTensor([1,1,1]) A = torch.sparse.FloatTensor(i, v, (3, 3, 3)) y = torch.matmul(A, x) loss = y.mean() loss.backward() We are planing (#12308) to support matmul on the way to make nn.Linear work for sparse. Can I ask what's your use cases? Please use #10043 for request on...",https://github.com/pytorch/pytorch/issues/12498,919a1803bc043c521e1d600353371dd17d9ecf59ba299b3b94ae0f90b3bdf6cb references,issue,19120,issue,17425,medium,issue.comments[0].body,"heck the index bounds as a recoverable Python error due to performance implications. But we can do a few things: Improve the error message (#17425) Add ""index out of bounds"" to the assertion like in Tensor indexing Document the behavior in https://pytorch.org/docs/stable/nn.ht...",https://github.com/pytorch/pytorch/issues/19120,d39589e8398349b0512b5b23d8f7896e64a898d72c68e004638ccc9d1eb60579 references,issue,15630,issue,12117,medium,issue.body,7 GPU models and configuration: Kepler + Pascal (build for Kepler+Maxwell+Pascal+Volta) Additional context Noticed several similar reports: #12117 #14872 #13808 #11203,https://github.com/pytorch/pytorch/issues/15630,06a3c3f9c9ee33abd2aaf23ad5eea2a098b16c9d3f56a8cc914be354fe7f5f48 references,issue,15630,issue,12117,medium,issue.comments[0].body,"#12117 has a pretty detailed description of how it gets triggered; is that applying to your case? If you only want to build PyTorch, a workaround",https://github.com/pytorch/pytorch/issues/15630,9abc42211fc7956883cb2d1e01877ce84053f1caef88257b783fdede7437769d references,issue,7127,issue,6273,medium,issue.comments[0].body,"it is same problem like #6273, but no answer",https://github.com/pytorch/pytorch/issues/7127,3d6552b0929013054da37d576a01f3e07c4293dab244ffecba54025713ce081d references,issue,10254,issue,10007,medium,issue.comments[0].body,Same problem? #10007,https://github.com/pytorch/pytorch/issues/10254,09e979f628d0c4691e74b7d987da0b45592096e9af18ed1f84c023768614565d references,issue,10007,issue,9853,medium,issue.body,"rtunately, the new version of Hypothesis has some behavior changes which cause previously stable tests to start flaking. Known cases: #9854 #9853 #9833 #9832 These need to be fixed before we can upgrade.",https://github.com/pytorch/pytorch/issues/10007,aa4573896510a14429f8a56d8c5728784149ed7b4106a8c62cc0586d26bdd965 references,issue,10007,issue,9854,medium,issue.body,". Unfortunately, the new version of Hypothesis has some behavior changes which cause previously stable tests to start flaking. Known cases: #9854 #9853 #9833 #9832 These need to be fixed before we can upgrade.",https://github.com/pytorch/pytorch/issues/10007,f6c161fdde2e4abaaa114996a46c9f4f21488d4ee5e59de42e6dc37c93fa9402 references,issue,10007,issue,10254,medium,issue.comments[0].body,"#10254 Do i have the same problem? My - hypothesis version is 3.66.26 , when i change it to 3.40.0, nothing happended. And nothing was printed. So",https://github.com/pytorch/pytorch/issues/10007,adfb531ccc2d1d7da4e0071b5b182eaff0cd96b089deb666c4f2afa020ece3b7 references,pr,189096,pr,189095,medium,pr.body,"Stack from ghstack (oldest at bottom): -> #189096 #189095 synchronize_stream / synchronize_device / synchronize_event block the CPU until a stream (or event) drains, so every subsequent kernel laun",https://github.com/pytorch/pytorch/pull/189096,0537134afb0d309fa15cd1ef91b578aa1f92eabab8725133834ca0c15259493a review guidance,pr,189096,pr,189095,high,pr.reviews[0].body,"(Reviewed by me, assisted by AI) [blocker] Same-stream sync ops aren't ordered relative to each other, which can silently corrupt CUDA event capture. This is the same ""bare side-effectful sync floats"" mechanism as bug #1 -- which this PR fixes for the CPU barriers -- left open for record/wait_eve...",https://github.com/pytorch/pytorch/pull/189096,47c657bc526a3dae7c5cde3a071b7d9ddd2aa4b1b3990d81a3bbfdbd8bbb5a45 review guidance,pr,189096,pr,189096,high,pr.reviews[0].body,"(Reviewed by me, assisted by AI) [blocker] Same-stream sync ops aren't ordered relative to each other, which can silently corrupt CUDA event capture. This is the same ""bare side-effectful sync floats"" mechanism as bug #1 -- which this PR fixes for the CPU barriers -- left open for record/wait_eve...",https://github.com/pytorch/pytorch/pull/189096,a7389fefb84bfae2caaee2d1aa4dbfb05856b52758e5bacbd9787636a50d2abe references,pr,189024,pr,157149,medium,pr.body,Stack from ghstack (oldest at bottom): -> #189024 #188825 #188834 #188824 #157149 #188639 #188638 #188004 #187744 #187690 Dynamo wrapped exceptions in a single ExceptionVariable that only tracked args/context/cause/traceb,https://github.com/pytorch/pytorch/pull/189024,d91279ce2e15bfc9ce940780a054d6dea24a3d97c732161a8f3372519e69cc3a references,pr,189024,pr,187690,medium,pr.body,Stack from ghstack (oldest at bottom): -> #189024 #188825 #188834 #188824 #157149 #188639 #188638 #188004 #187744 #187690 Dynamo wrapped exceptions in a single ExceptionVariable that only tracked args/context/cause/traceback. Reading exception-specific attribut,https://github.com/pytorch/pytorch/pull/189024,53deb6404fe641d0643c8ef01cb07e061abf206b4a190e1a935047630ebe65c2 references,pr,189024,pr,187744,medium,pr.body,Stack from ghstack (oldest at bottom): -> #189024 #188825 #188834 #188824 #157149 #188639 #188638 #188004 #187744 #187690 Dynamo wrapped exceptions in a single ExceptionVariable that only tracked args/context/cause/traceback. Reading exception-specific,https://github.com/pytorch/pytorch/pull/189024,23afcd3b4ac72e255e2c7c2bd89c15ea78f742656a46c99bc0d9d970df1a8a70 references,pr,189024,pr,188004,medium,pr.body,Stack from ghstack (oldest at bottom): -> #189024 #188825 #188834 #188824 #157149 #188639 #188638 #188004 #187744 #187690 Dynamo wrapped exceptions in a single ExceptionVariable that only tracked args/context/cause/traceback. Reading exception-s,https://github.com/pytorch/pytorch/pull/189024,c0168a7261e288d8c38a8e36b3f7ab1e969ccc5d3c3ce998e3572737b13b198e references,pr,189024,pr,188638,medium,pr.body,Stack from ghstack (oldest at bottom): -> #189024 #188825 #188834 #188824 #157149 #188639 #188638 #188004 #187744 #187690 Dynamo wrapped exceptions in a single ExceptionVariable that only tracked args/context/cause/traceback. Reading exc,https://github.com/pytorch/pytorch/pull/189024,d7594857563c04a67d5ad28600f60a3ba764e4fcd14550a1450c53777e865077 references,pr,189024,pr,188639,medium,pr.body,Stack from ghstack (oldest at bottom): -> #189024 #188825 #188834 #188824 #157149 #188639 #188638 #188004 #187744 #187690 Dynamo wrapped exceptions in a single ExceptionVariable that only tracked args/context/cause/traceback. Rea,https://github.com/pytorch/pytorch/pull/189024,e950d22eddb2db9623fbc6d684e39bec503c3b871de325baf828e02b5b855341 references,pr,189024,pr,188824,medium,pr.body,Stack from ghstack (oldest at bottom): -> #189024 #188825 #188834 #188824 #157149 #188639 #188638 #188004 #187744 #187690 Dynamo wrapped exceptions in a single ExceptionVariable that only tracked args/context/caus,https://github.com/pytorch/pytorch/pull/189024,244dc4824d32a53f3980c41d8be424b64d3f0cc03f4e416f173c0ecb071e1ef9 references,pr,189024,pr,188825,medium,pr.body,Stack from ghstack (oldest at bottom): -> #189024 #188825 #188834 #188824 #157149 #188639 #188638 #188004 #187744 #187690 Dynamo wrapped exceptions in a single ExceptionVariable that only tracked a,https://github.com/pytorch/pytorch/pull/189024,27131d5abe20e72ce9d8f991267cdd44e5db9ca8902cf8ea8a911ec6cc3b5bef references,pr,189024,pr,188834,medium,pr.body,Stack from ghstack (oldest at bottom): -> #189024 #188825 #188834 #188824 #157149 #188639 #188638 #188004 #187744 #187690 Dynamo wrapped exceptions in a single ExceptionVariable that only tracked args/cont,https://github.com/pytorch/pytorch/pull/189024,e6ab745cc0becffd45888149368aaaf1bbbac731aa5777b75bced041217f35d4 references,pr,189024,pr,187690,medium,pr.comments[5].body,## Merge failed **Reason**: Merge rule check failed for stacked PR #187690: 4 mandatory check(s) failed. The first few are: - [Lint / lintrunner-clang-all / lint](https://github.com/pytorch/pytorch/actions/runs/288,https://github.com/pytorch/pytorch/pull/189024,fcc247b7cedd321b85f274745adae1d604cb921b18d12c4e5c81b057a780377c review guidance,pr,189024,pr,157149,high,pr.reviews[0].body,"LGTM, thanks.",https://github.com/pytorch/pytorch/pull/189024,435576ca04384a0aca092373b217221adb2c0705d64e679b669475d16eaf16f5 review guidance,pr,189024,pr,187690,high,pr.reviews[0].body,"LGTM, thanks.",https://github.com/pytorch/pytorch/pull/189024,3a98488f90090eab7e449f5b21f971b485c8d2c9726cc790d98478e43e527c40 review guidance,pr,189024,pr,187744,high,pr.reviews[0].body,"LGTM, thanks.",https://github.com/pytorch/pytorch/pull/189024,2a21aaf476289b49ba272a3a841e7bb4fbe43d2f4ee661e7c1b33c35b723f420 review guidance,pr,189024,pr,188004,high,pr.reviews[0].body,"LGTM, thanks.",https://github.com/pytorch/pytorch/pull/189024,23e21aec819b7046bca710b0c253fc9f570c3358cf3f2da45d97e7b36754aa7f review guidance,pr,189024,pr,188638,high,pr.reviews[0].body,"LGTM, thanks.",https://github.com/pytorch/pytorch/pull/189024,4c907233db4d10e648919044207a2df17ae0f92143c36b998dcda4514c093dad review guidance,pr,189024,pr,188639,high,pr.reviews[0].body,"LGTM, thanks.",https://github.com/pytorch/pytorch/pull/189024,1eb2a070c10c766e2c9d8d5a2a29275f56e94937fa74ebd942b5c8c36be4fc4e review guidance,pr,189024,pr,188824,high,pr.reviews[0].body,"LGTM, thanks.",https://github.com/pytorch/pytorch/pull/189024,e4691394e112cb86d6b536a618a70a36aa003385f6967e5469d97c8638516e8c review guidance,pr,189024,pr,188825,high,pr.reviews[0].body,"LGTM, thanks.",https://github.com/pytorch/pytorch/pull/189024,d1f988cc8764d6b4b7675544b98203d368bb210769ebcb7cb2be08d5a4853f3b review guidance,pr,189024,pr,188834,high,pr.reviews[0].body,"LGTM, thanks.",https://github.com/pytorch/pytorch/pull/189024,e3cd736adfb6dfb8c77775e1563a2871ead6c5ab216d61083b9f649985bcbd31 review guidance,pr,189024,pr,189024,high,pr.reviews[0].body,"LGTM, thanks.",https://github.com/pytorch/pytorch/pull/189024,257d94e306258f1c8c70dce4058e89fd1f4486fe6ecb45a9a2084c9f7e7c27d8 review guidance,pr,189024,pr,157149,high,pr.reviews[2].body,"LGTM, thanks.",https://github.com/pytorch/pytorch/pull/189024#pullrequestreview-4644831542,f969be241a25385562189dd5688e2e614e93b5d0aabbcf0cbc3a8405e2c2ecc1 review guidance,pr,189024,pr,187690,high,pr.reviews[2].body,"LGTM, thanks.",https://github.com/pytorch/pytorch/pull/189024#pullrequestreview-4644831542,942d6d0d95b232fb710a5c4287acf847625e1892e1708bbd3a36f358b8e75379 review guidance,pr,189024,pr,187744,high,pr.reviews[2].body,"LGTM, thanks.",https://github.com/pytorch/pytorch/pull/189024#pullrequestreview-4644831542,c5bc23406e219d8aa1c68121dfdbd885062f214fb153d1cd39769a2dc2c6222c review guidance,pr,189024,pr,188004,high,pr.reviews[2].body,"LGTM, thanks.",https://github.com/pytorch/pytorch/pull/189024#pullrequestreview-4644831542,10adb85de5b076e8ed20329e29e16eb0412562bc086b09ccd3d089ab5b47b266 review guidance,pr,189024,pr,188638,high,pr.reviews[2].body,"LGTM, thanks.",https://github.com/pytorch/pytorch/pull/189024#pullrequestreview-4644831542,bcd712467a34ba2cfa9e6ef8a3e2aacd2367cee29d15a2b025cf8253f1e4c784 review guidance,pr,189024,pr,188639,high,pr.reviews[2].body,"LGTM, thanks.",https://github.com/pytorch/pytorch/pull/189024#pullrequestreview-4644831542,6bfab14af03456610f67e6e73f9af2b327defae156ce0eb25aa7eea53a917bcf review guidance,pr,189024,pr,188824,high,pr.reviews[2].body,"LGTM, thanks.",https://github.com/pytorch/pytorch/pull/189024#pullrequestreview-4644831542,29a64e948b60350b1a34e7655370840406cf64ab10ccd51153a262c2afd13b96 review guidance,pr,189024,pr,188825,high,pr.reviews[2].body,"LGTM, thanks.",https://github.com/pytorch/pytorch/pull/189024#pullrequestreview-4644831542,f8ea8cb7f2efe827ad8564032ec2e7b1bc8f4fda4cf1dc6d8e82aafccf5be9db review guidance,pr,189024,pr,188834,high,pr.reviews[2].body,"LGTM, thanks.",https://github.com/pytorch/pytorch/pull/189024#pullrequestreview-4644831542,a5b16849da8a74c2b9acd4a410370fbfd574f35f27902c655a7ae449f509e070 review guidance,pr,189024,pr,189024,high,pr.reviews[2].body,"LGTM, thanks.",https://github.com/pytorch/pytorch/pull/189024#pullrequestreview-4644831542,63e955e837c18354865aab7813b0ca443d21d28f44fdfdb7aa51271ed2d041ce references,pr,188597,pr,188242,medium,pr.body,"l (gfx942/gfx950/gfx1250). Enabled via the main merge: Flash attention and memory-efficient attention on gfx1250 via AOTriton 0.12.1b (from #188242), which this branch merges in. What is not enabled yet: CK SDPA on gfx1250: composable_kernel has no gfx1250 support; the CK SDPA...",https://github.com/pytorch/pytorch/pull/188597,c5e2e39f1a37d23af4595304a171d52f9035e57576f437990cd282bfd48c14ea references,pr,188597,pr,188242,medium,pr.reviews[0].body,Do not touch cmake/External/aotriton.cmake. This only creates conflicts with ongoing PR #188242,https://github.com/pytorch/pytorch/pull/188597,8a3d47d122be3a47d8beba39aa916ac246cdc8c9955431c35392bc97c74afe87 closes,pr,188980,issue,188790,high,pr.closingIssuesReferences,pr #188980 declares a closing reference to issue #188790.,https://github.com/pytorch/pytorch/pull/188980,c85c867cbed47ef9d25009a2a345cbf802fce2e75112b96d9b10da93f8006819 closes,pr,188980,issue,188790,high,pr.body,"Fixes #188790 Summary In test/distributed/checkpoint/_experimental/test_staging.py, the block that appends the async-staging and non-blocking-copy test c",https://github.com/pytorch/pytorch/pull/188980,16954e57ccb967fada8ca78f440bfaba274e9e17163091cb8c695536a27dee13 review guidance,pr,187940,pr,182619,high,pr.reviews[0].body,Hi @orrangetabby17 Could you please sign the EasyCLA. And show the performance difference data here.,https://github.com/pytorch/pytorch/pull/187940,3eb95fcc85f5013822489bf9c282f635f10469468dfc148d144df0fa32db7b90 review guidance,pr,187940,pr,182619,high,pr.reviews[1].body,Overall looks good. I think we need a test case for XPU to detect the generated kernel count is 0. Such as https://github.com/pytorch/pytorch/pull/127722/changes#diff-809c39aeafb3acc92289f42a63e670a8719d4ce5627d5f88820142d80edf8d2aR7882-R7904,https://github.com/pytorch/pytorch/pull/187940,95f3288a6a0bddc46d49671ef02be91d19cdd997dd8cd7e393ffb24c70e928d9 review guidance,pr,189284,pr,189284,high,pr.reviews[0].body,We should either say that the Node is not ready for inspection or actually trigger when ready!,https://github.com/pytorch/pytorch/pull/189284,4b3e8d2a6fb2686553b4e4095b38d71ac0515ece9c64f2a5adc14792408e0cb0 review guidance,pr,189284,pr,189284,high,pr.reviews[2].body,We should either say that the Node is not ready for inspection or actually trigger when ready!,https://github.com/pytorch/pytorch/pull/189284#pullrequestreview-4656593121,8c34fd694723e5a29c2a73be014f3762a297056fd9b1fc68bd57ce42a5ac256c references,pr,181680,issue,174929,medium,pr.comments[1].body,"ossibly due to flakiness on trunk: ⏳ inductor / inductor-cpu-test / test (cpu_inductor_torchbench, 1, 2, mt-l-x86iamx-8-64, unstable) (gh) (#174929) ⏳ inductor / inductor-cpu-test / test (cpu_inductor_torchbench, 2, 2, mt-l-x86iamx-8-64, unstable) (gh) (#174929) This comment w...",https://github.com/pytorch/pytorch/pull/181680,8c504f56ee90dcbd4f12a3c728b1187919e7f7d9820aa22f0f605cc408bd70bf review guidance,pr,187600,pr,149039,high,pr.reviews[0].body,Can you inline the new function rather than add to the API surface area? You only use the new function in one place if I read this right.,https://github.com/pytorch/pytorch/pull/187600,a9769beb6605e7ec5f843098368a8c6e95fc466e3b5b5ef24758295ff5d312a8 references,pr,189083,pr,188621,medium,pr.body,Stack from ghstack (oldest at bottom): #189185 -> #189083 #188621 #189109 The CUPTI monitor's background flush loop woke on a Python thread every background_flush_period_s just to call cuptiActivityFlushAl,https://github.com/pytorch/pytorch/pull/189083,6e03da635088df66cc72a0f49f41db68cc095ca9125eba587ae4b33922db2da1 references,pr,189083,pr,189109,medium,pr.body,Stack from ghstack (oldest at bottom): #189185 -> #189083 #188621 #189109 The CUPTI monitor's background flush loop woke on a Python thread every background_flush_period_s just to call cuptiActivityFlushAll -- a p,https://github.com/pytorch/pytorch/pull/189083,50e688d3091be6df633bf47c31383da1b99b81f1eadd1a0e36dc874f65e377a2 references,pr,189083,pr,189185,medium,pr.body,Stack from ghstack (oldest at bottom): #189185 -> #189083 #188621 #189109 The CUPTI monitor's background flush loop woke on a Python thread every background_flush_period_s just to call c,https://github.com/pytorch/pytorch/pull/189083,37443a1bb9adf973f201ed711a44a7ace8b82430ab3f5c190c18ff44532f81d9 references,pr,189185,pr,188621,medium,pr.body,Stack from ghstack (oldest at bottom): -> #189185 #189083 #188621 #189109 Ports the config-only change from commit 4a24d3b207 (drop the TORCH_CUPTI_MONITOR_* env vars) and turns CuptiMonitor into a process,https://github.com/pytorch/pytorch/pull/189185,86b5a8e186a02507ec08ea7070b6523ace814a89071ba2954d96be98ebc897eb references,pr,189185,pr,189083,medium,pr.body,Stack from ghstack (oldest at bottom): -> #189185 #189083 #188621 #189109 Ports the config-only change from commit 4a24d3b207 (drop the TORCH_CUPTI_MONITOR_* env vars) and turns CuptiMonitor into a,https://github.com/pytorch/pytorch/pull/189185,c60c54edf34a7b7c2311bcfa15a5cc1bdd5ebb6ec19032f0d1bb6a8608c4c85c references,pr,189185,pr,189109,medium,pr.body,Stack from ghstack (oldest at bottom): -> #189185 #189083 #188621 #189109 Ports the config-only change from commit 4a24d3b207 (drop the TORCH_CUPTI_MONITOR_* env vars) and turns CuptiMonitor into a process-wide si,https://github.com/pytorch/pytorch/pull/189185,423d670d1bfa3ad7372905cede0f7532e0ca441163b76c3e1cd20e733842dbd5 references,pr,188621,pr,189083,medium,pr.body,"Stack from ghstack (oldest at bottom): #189185 #189083 -> #188621 #189109 The v2 CUPTI user-defined-record path selects records by field id (CUpti_ActivityFieldIds), which cupti-python does not",https://github.com/pytorch/pytorch/pull/188621,a9248c2a533ad71023a89ac247849c819f62f71f8effd5d95569ae63a4257b71 competes with,pr,188621,pr,189109,medium,pr.body,"Stack from ghstack (oldest at bottom): #189185 #189083 -> #188621 #189109 The v2 CUPTI user-defined-record path selects records by field id (CUpti_ActivityFieldIds), which cupti-python does not expose. Instead of",https://github.com/pytorch/pytorch/pull/188621,6f7eb50642499c77eb970894e0ae6c32fde665a13795083d00df5477a0e1ab80 references,pr,188621,pr,189185,medium,pr.body,"Stack from ghstack (oldest at bottom): #189185 #189083 -> #188621 #189109 The v2 CUPTI user-defined-record path selects records by field id (CUpti_ActivityFieldIds), which cupti-python d",https://github.com/pytorch/pytorch/pull/188621,87c7a09963b02d8bf252f9aa0595780c081fbc8c8df39762d128ad233a7d2235 references,pr,188100,pr,188299,medium,pr.body,"Stack from ghstack (oldest at bottom): #189286 #188573 #188299 -> #188100 Allows for lazy building of error messages only when tests fail! Saves on overhead, especially when serialized tensors are invol",https://github.com/pytorch/pytorch/pull/188100,8387e966c73f7197f03132a17503fc7bedc25ecac4a5b17a29ceaf3557f5ce37 references,pr,188100,pr,188573,medium,pr.body,"Stack from ghstack (oldest at bottom): #189286 #188573 #188299 -> #188100 Allows for lazy building of error messages only when tests fail! Saves on overhead, especially when serialized tensors a",https://github.com/pytorch/pytorch/pull/188100,86b446380aeb2ccb8798fb041b34644148b48c22f3dc1a923550f14831f8c2ff references,pr,188100,pr,189286,medium,pr.body,"Stack from ghstack (oldest at bottom): #189286 #188573 #188299 -> #188100 Allows for lazy building of error messages only when tests fail! Saves on overhead, especially when serialized t",https://github.com/pytorch/pytorch/pull/188100,e0ba6df09d2b4ccc9280179487615cdeeb76d6640de4836bd283139e50407bf3 closes,pr,167224,issue,134385,high,pr.closingIssuesReferences,pr #167224 declares a closing reference to issue #134385.,https://github.com/pytorch/pytorch/pull/167224,d923d2bbdb665f76c6281e9b0a6015986ef4caa6d6e2f738cc735e919bc26c3d closes,pr,167224,issue,134385,high,pr.body,"Fixes #134385 FlopCounterMode returns NotImplemented when it encounters a Higher Order Operator it does not handle, which makes it hard to use for benchm",https://github.com/pytorch/pytorch/pull/167224,c72a75b0696c5c7625669d16b07436dd6a0094e9b90cf410ef322266ab937c8c closes,pr,167224,issue,134385,high,pr.reviews[1].body,"(Claude Review) [blocker] None of the 5 new tests exercise a Higher-Order Operator, but per #134385 the NotImplementedError crash only ever happens for HOPs — regular custom ops already execute and count 0 flops without raising (which is e",https://github.com/pytorch/pytorch/pull/167224,5a590f83b02ed185d94a19d8dc539ca2291c06a2d2c5db1b0bef19d1f15616e2 review guidance,pr,167224,issue,134385,high,pr.reviews[0].body,"I'm confused, the default behavior for custom ops is that we treat them as having zero flops, not raise an exception. Where are you getting an exception raised?",https://github.com/pytorch/pytorch/pull/167224,1288618e957ede62f9bb7b150d5e3438d8a686b4ac1a20bd3c984a2779af4c4a review guidance,pr,167224,issue,134385,high,pr.reviews[1].body,"(Claude Review) [blocker] None of the 5 new tests exercise a Higher-Order Operator, but per #134385 the NotImplementedError crash only ever happens for HOPs — regular custom ops already execute and count 0 flops without raising (which is exactly why test_skip_unsupported_default_behavior runs an...",https://github.com/pytorch/pytorch/pull/167224,a237ab9104aa5e9440168529ce1d4409e83977ae228879a537b30a490e0711ad references,pr,188299,pr,188100,medium,pr.body,Stack from ghstack (oldest at bottom): #189286 #188573 -> #188299 #188100 Apply new lazy error message feature to call-sites throughout test files. cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @Xia,https://github.com/pytorch/pytorch/pull/188299,884500ea577214459d2048e88b07cbd735f5f143704d378cf6b4dfa0f52c4e07 references,pr,188299,pr,188573,medium,pr.body,Stack from ghstack (oldest at bottom): #189286 #188573 -> #188299 #188100 Apply new lazy error message feature to call-sites throughout test files. cc @voznesenskym @penguinwu @EikanWang @jgong5,https://github.com/pytorch/pytorch/pull/188299,7becedc27cda8eb0990f8bf19f6cfd7b93751f2f49a568633a9f6ff87799e942 references,pr,188299,pr,189286,medium,pr.body,Stack from ghstack (oldest at bottom): #189286 #188573 -> #188299 #188100 Apply new lazy error message feature to call-sites throughout test files. cc @voznesenskym @penguinwu @EikanWang,https://github.com/pytorch/pytorch/pull/188299,58f0f0b1fec4725dda349b3e8bed1308e7591c4afe53c7c01575f4c80be38cb7 references,pr,188573,pr,188100,medium,pr.body,"Stack from ghstack (oldest at bottom): #189286 -> #188573 #188299 #188100 I don't expect huge perf improvements, but why not? cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe",https://github.com/pytorch/pytorch/pull/188573,87219b35df487d8011dd24cf97008af5ff396556fd9bd2fec1baca0184bc7193 references,pr,188573,pr,188299,medium,pr.body,"Stack from ghstack (oldest at bottom): #189286 -> #188573 #188299 #188100 I don't expect huge perf improvements, but why not? cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zh",https://github.com/pytorch/pytorch/pull/188573,57ca22aeb43a30d1907434720fa396df207b35337996296d189fa945edbad9b1 references,pr,188573,pr,189286,medium,pr.body,"Stack from ghstack (oldest at bottom): #189286 -> #188573 #188299 #188100 I don't expect huge perf improvements, but why not? cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen",https://github.com/pytorch/pytorch/pull/188573,c7294c16d4890a154b67db07cf92d768fb6075111b63981570f2580b6da55833 references,pr,187465,pr,189088,medium,pr.body,Stack from ghstack (oldest at bottom): -> #187465 #189088 This PR allows the NCCL symmetric-memory allocator to back its allocations with the CUDA caching allocator's expandable segments; when expa,https://github.com/pytorch/pytorch/pull/187465,8534a6c53eee6778b0326d33ec2c88676f3e99e55d781f9e3e1a0f29cb8baa58 references,pr,189314,pr,188112,medium,pr.body,"Stack from ghstack (oldest at bottom): #189316 #189315 -> #189314 #189190 #189188 #188739 #188112 #188470 #188469 #188468 Recognize structural 2-D block local-reduce aux outputs of the form acc.view(M // bm, bm, N // bn, bn) reduced over",https://github.com/pytorch/pytorch/pull/189314,120a3be1d88f35202ff9a0b84e73fb1d3ef508ffad0aab26600ff20bc0fb9a8e references,pr,189314,pr,188468,medium,pr.body,"Stack from ghstack (oldest at bottom): #189316 #189315 -> #189314 #189190 #189188 #188739 #188112 #188470 #188469 #188468 Recognize structural 2-D block local-reduce aux outputs of the form acc.view(M // bm, bm, N // bn, bn) reduced over both grouped dims. This",https://github.com/pytorch/pytorch/pull/189314,e5692c30a3ae8446eb9e3f94bf960d8dc4560e3985df3eb437c374bad57d2959 references,pr,189314,pr,188469,medium,pr.body,"Stack from ghstack (oldest at bottom): #189316 #189315 -> #189314 #189190 #189188 #188739 #188112 #188470 #188469 #188468 Recognize structural 2-D block local-reduce aux outputs of the form acc.view(M // bm, bm, N // bn, bn) reduced over both grouped di",https://github.com/pytorch/pytorch/pull/189314,8783bf66ca8d2c14168133809bb32842b1e55f55bbbf90126741d1859a4f4332 references,pr,189314,pr,188470,medium,pr.body,"Stack from ghstack (oldest at bottom): #189316 #189315 -> #189314 #189190 #189188 #188739 #188112 #188470 #188469 #188468 Recognize structural 2-D block local-reduce aux outputs of the form acc.view(M // bm, bm, N // bn, bn) reduced over both gr",https://github.com/pytorch/pytorch/pull/189314,86fa821d5a47a92676869d4867a90c2fec783b208812b428c3dc2b2c7afd8a75 references,pr,189314,pr,188739,medium,pr.body,"Stack from ghstack (oldest at bottom): #189316 #189315 -> #189314 #189190 #189188 #188739 #188112 #188470 #188469 #188468 Recognize structural 2-D block local-reduce aux outputs of the form acc.view(M // bm, bm, N // bn, bn) redu",https://github.com/pytorch/pytorch/pull/189314,05cfb4dc7ead8dca4ccca51997c5493262bbc5e973abc789caa846340e629543 references,pr,189314,pr,189188,medium,pr.body,"Stack from ghstack (oldest at bottom): #189316 #189315 -> #189314 #189190 #189188 #188739 #188112 #188470 #188469 #188468 Recognize structural 2-D block local-reduce aux outputs of the form acc.view(M // bm, bm, N // bn,",https://github.com/pytorch/pytorch/pull/189314,349b1cae62b393feb6b66da653678b35aa31f31521ff0df97af273339ae81b20 references,pr,189314,pr,189190,medium,pr.body,"Stack from ghstack (oldest at bottom): #189316 #189315 -> #189314 #189190 #189188 #188739 #188112 #188470 #188469 #188468 Recognize structural 2-D block local-reduce aux outputs of the form acc.view(M // bm, bm, N",https://github.com/pytorch/pytorch/pull/189314,01a61279d4bf609c7c8b7c456b791b0d9b706894fd6fa9266e0f6a32fd6ff26b references,pr,189314,pr,189315,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 #189315 -> #189314 #189190 #189188 #188739 #188112 #188470 #188469 #188468 Recognize structural 2-D block local-reduce aux outputs of the form acc.,https://github.com/pytorch/pytorch/pull/189314,33903e0746542f8e57eb821301fba3ac509b6569b1491e0893b7cb4369e3b5d0 references,pr,189314,pr,189316,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 #189315 -> #189314 #189190 #189188 #188739 #188112 #188470 #188469 #188468 Recognize structural 2-D block local-reduce aux outputs of the f,https://github.com/pytorch/pytorch/pull/189314,5547fd13dfefe38ea94bbec3300ec5d0c5aaa8f0b45cf5a9cc635f5b0cf67cc9 references,pr,189315,pr,188112,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 -> #189315 #189314 #189190 #189188 #188739 #188112 #188470 #188469 #188468 Enable the first QUACK-backed 2-D block local-reduce STORE path for 128x128 block-scale aux outputs. The compiler n,https://github.com/pytorch/pytorch/pull/189315,2c6b8c514a71d704742cb1aff9ebee3ed94a9e0d38366213196756f6282e5a89 references,pr,189315,pr,188468,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 -> #189315 #189314 #189190 #189188 #188739 #188112 #188470 #188469 #188468 Enable the first QUACK-backed 2-D block local-reduce STORE path for 128x128 block-scale aux outputs. The compiler now threads block geometr,https://github.com/pytorch/pytorch/pull/189315,654c2f23f1dbc2556f5f6a15293ced615bc8a4e22eba27a16e57c49b7d386b9d references,pr,189315,pr,188469,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 -> #189315 #189314 #189190 #189188 #188739 #188112 #188470 #188469 #188468 Enable the first QUACK-backed 2-D block local-reduce STORE path for 128x128 block-scale aux outputs. The compiler now threads block,https://github.com/pytorch/pytorch/pull/189315,a17e407986e52af1b18b5340ef7695391d998bd8f16cc451f45f3f1fe52c4cd4 references,pr,189315,pr,188470,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 -> #189315 #189314 #189190 #189188 #188739 #188112 #188470 #188469 #188468 Enable the first QUACK-backed 2-D block local-reduce STORE path for 128x128 block-scale aux outputs. The compiler now threa,https://github.com/pytorch/pytorch/pull/189315,215ee6d666490e2d2e7cfb5446353c3cd987efd28c0dccd6855f5be94d47a01b references,pr,189315,pr,188739,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 -> #189315 #189314 #189190 #189188 #188739 #188112 #188470 #188469 #188468 Enable the first QUACK-backed 2-D block local-reduce STORE path for 128x128 block-scale aux outputs. The co,https://github.com/pytorch/pytorch/pull/189315,198ba3f378ed003739c1784e71a2de890be2894ea656e0660ac663472576652f references,pr,189315,pr,189188,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 -> #189315 #189314 #189190 #189188 #188739 #188112 #188470 #188469 #188468 Enable the first QUACK-backed 2-D block local-reduce STORE path for 128x128 block-scale aux outputs,https://github.com/pytorch/pytorch/pull/189315,84f55f84dd3dbc43dc24cf46d12170bb14baff957c71c9482f2f77e490a3ec23 references,pr,189315,pr,189190,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 -> #189315 #189314 #189190 #189188 #188739 #188112 #188470 #188469 #188468 Enable the first QUACK-backed 2-D block local-reduce STORE path for 128x128 block-scale aux,https://github.com/pytorch/pytorch/pull/189315,e8ab5666488e71711185d5c532474e9174e57234a43b48ff60b8547c1b3fc893 references,pr,189315,pr,189314,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 -> #189315 #189314 #189190 #189188 #188739 #188112 #188470 #188469 #188468 Enable the first QUACK-backed 2-D block local-reduce STORE path for 128x128 block-s,https://github.com/pytorch/pytorch/pull/189315,aa342249cfdf8df3580f87f266f3855f045f51a192680a812a1a63d7a8e77839 references,pr,189315,pr,189316,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 -> #189315 #189314 #189190 #189188 #188739 #188112 #188470 #188469 #188468 Enable the first QUACK-backed 2-D block local-reduce STORE path,https://github.com/pytorch/pytorch/pull/189315,30edb2c4a2078ae6be629d01c44c87f36638ea3ea16453446a0f8fbef3162222 review guidance,pr,189315,pr,188112,high,pr.reviews[0].body,"💡 Codex Review Here are some automated review suggestions for this pull request. Reviewed commit: 0fe00ed181 ℹ️ About Codex in GitHub Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you Open a pull request for review Mark a draft as ready Comment ""@code...",https://github.com/pytorch/pytorch/pull/189315,e761df87c5124bfc599ed5d7dfa15238eea870e51bd06cca8494c2d2533dc08f review guidance,pr,189315,pr,188468,high,pr.reviews[0].body,"💡 Codex Review Here are some automated review suggestions for this pull request. Reviewed commit: 0fe00ed181 ℹ️ About Codex in GitHub Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you Open a pull request for review Mark a draft as ready Comment ""@code...",https://github.com/pytorch/pytorch/pull/189315,2697dcccb531cbe511196f5701be33ad18bde499c822fcbe46933de50f86eebd review guidance,pr,189315,pr,188469,high,pr.reviews[0].body,"💡 Codex Review Here are some automated review suggestions for this pull request. Reviewed commit: 0fe00ed181 ℹ️ About Codex in GitHub Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you Open a pull request for review Mark a draft as ready Comment ""@code...",https://github.com/pytorch/pytorch/pull/189315,cd65ff13e80eec5838875b28feee796e82265bbad1598f4e38071368f7aded93 review guidance,pr,189315,pr,188470,high,pr.reviews[0].body,"💡 Codex Review Here are some automated review suggestions for this pull request. Reviewed commit: 0fe00ed181 ℹ️ About Codex in GitHub Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you Open a pull request for review Mark a draft as ready Comment ""@code...",https://github.com/pytorch/pytorch/pull/189315,f39264758dbcb09d1925399bef7ba7ec1480694dba67742654fa12e325b5567d review guidance,pr,189315,pr,188739,high,pr.reviews[0].body,"💡 Codex Review Here are some automated review suggestions for this pull request. Reviewed commit: 0fe00ed181 ℹ️ About Codex in GitHub Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you Open a pull request for review Mark a draft as ready Comment ""@code...",https://github.com/pytorch/pytorch/pull/189315,8ca4a7611947255b18734895b9cbb338e67d3006c6f271c10bdc700ab1c6459b review guidance,pr,189315,pr,189188,high,pr.reviews[0].body,"💡 Codex Review Here are some automated review suggestions for this pull request. Reviewed commit: 0fe00ed181 ℹ️ About Codex in GitHub Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you Open a pull request for review Mark a draft as ready Comment ""@code...",https://github.com/pytorch/pytorch/pull/189315,373ff4375deba71914702da2621d944237e95576663cf9ad833b1b86758a8f59 review guidance,pr,189315,pr,189190,high,pr.reviews[0].body,"💡 Codex Review Here are some automated review suggestions for this pull request. Reviewed commit: 0fe00ed181 ℹ️ About Codex in GitHub Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you Open a pull request for review Mark a draft as ready Comment ""@code...",https://github.com/pytorch/pytorch/pull/189315,13677391b60e2ca9104f018515ae1be450ae21e0632eb39baed299937089bb77 review guidance,pr,189315,pr,189314,high,pr.reviews[0].body,"💡 Codex Review Here are some automated review suggestions for this pull request. Reviewed commit: 0fe00ed181 ℹ️ About Codex in GitHub Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you Open a pull request for review Mark a draft as ready Comment ""@code...",https://github.com/pytorch/pytorch/pull/189315,20f132f6fe0daec852a23b505714f84d999478421bab911c38bd84b7d42eafb6 review guidance,pr,189315,pr,189315,high,pr.reviews[0].body,"💡 Codex Review Here are some automated review suggestions for this pull request. Reviewed commit: 0fe00ed181 ℹ️ About Codex in GitHub Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you Open a pull request for review Mark a draft as ready Comment ""@code...",https://github.com/pytorch/pytorch/pull/189315,5f2c7e8c11f21776c1262bc1a4fda558220fe9f0452ffdcacfd4673866343b48 review guidance,pr,189315,pr,189316,high,pr.reviews[0].body,"💡 Codex Review Here are some automated review suggestions for this pull request. Reviewed commit: 0fe00ed181 ℹ️ About Codex in GitHub Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you Open a pull request for review Mark a draft as ready Comment ""@code...",https://github.com/pytorch/pytorch/pull/189315,3f72bd44913a446fc9dc36cbdffa9a92b6abf4c42357d3491b79d8be7e50551f references,pr,189316,pr,188112,medium,pr.body,Stack from ghstack (oldest at bottom): -> #189316 #189315 #189314 #189190 #189188 #188739 #188112 #188470 #188469 #188468 Compose 2-D block local-reduce stores with physical-only finalizers such as mx_e8m0_scale. Block aux plan classific,https://github.com/pytorch/pytorch/pull/189316,a3799adb0be4f045c8eb15f3a754975bb5aa21ff368bf6a1506559c409404e19 references,pr,189316,pr,188468,medium,pr.body,Stack from ghstack (oldest at bottom): -> #189316 #189315 #189314 #189190 #189188 #188739 #188112 #188470 #188469 #188468 Compose 2-D block local-reduce stores with physical-only finalizers such as mx_e8m0_scale. Block aux plan classification now traces a shape,https://github.com/pytorch/pytorch/pull/189316,a796b799ac11ae51526ad5fbd83147b62f1df06fa7828ae627c5ef26fd7d78f3 references,pr,189316,pr,188469,medium,pr.body,Stack from ghstack (oldest at bottom): -> #189316 #189315 #189314 #189190 #189188 #188739 #188112 #188470 #188469 #188468 Compose 2-D block local-reduce stores with physical-only finalizers such as mx_e8m0_scale. Block aux plan classification now traces,https://github.com/pytorch/pytorch/pull/189316,4b076fd8edffdbf12ee4c6800fe6d3442c3acb94fc86189b430eb29adee9343a references,pr,189316,pr,188470,medium,pr.body,Stack from ghstack (oldest at bottom): -> #189316 #189315 #189314 #189190 #189188 #188739 #188112 #188470 #188469 #188468 Compose 2-D block local-reduce stores with physical-only finalizers such as mx_e8m0_scale. Block aux plan classification no,https://github.com/pytorch/pytorch/pull/189316,64766ea290f3668819e5b75fd5d8be03b03e66de0e51ff2b29344e113201ffc1 references,pr,189316,pr,188739,medium,pr.body,Stack from ghstack (oldest at bottom): -> #189316 #189315 #189314 #189190 #189188 #188739 #188112 #188470 #188469 #188468 Compose 2-D block local-reduce stores with physical-only finalizers such as mx_e8m0_scale. Block aux plan c,https://github.com/pytorch/pytorch/pull/189316,4dafdb3460ab390ea98763d818aca05d1de71e01c7027ddef299a870b0fe5046 references,pr,189316,pr,189188,medium,pr.body,Stack from ghstack (oldest at bottom): -> #189316 #189315 #189314 #189190 #189188 #188739 #188112 #188470 #188469 #188468 Compose 2-D block local-reduce stores with physical-only finalizers such as mx_e8m0_scale. Block au,https://github.com/pytorch/pytorch/pull/189316,359f88109b5121b1dffeb02ff958dd53f31aab8a75b87ab90143cf2dfcc9c03a references,pr,189316,pr,189190,medium,pr.body,Stack from ghstack (oldest at bottom): -> #189316 #189315 #189314 #189190 #189188 #188739 #188112 #188470 #188469 #188468 Compose 2-D block local-reduce stores with physical-only finalizers such as mx_e8m0_scale.,https://github.com/pytorch/pytorch/pull/189316,0c3a41aef17e8eba3e11a2c3c7c690fe983755058d530936e13cbd77508bce27 references,pr,189316,pr,189314,medium,pr.body,Stack from ghstack (oldest at bottom): -> #189316 #189315 #189314 #189190 #189188 #188739 #188112 #188470 #188469 #188468 Compose 2-D block local-reduce stores with physical-only finalizers such as mx_e8m0,https://github.com/pytorch/pytorch/pull/189316,b6d156dda232f3dc70eeadf950b428bc7c733a719b98bc7f7ad7589f8668c787 references,pr,189316,pr,189315,medium,pr.body,Stack from ghstack (oldest at bottom): -> #189316 #189315 #189314 #189190 #189188 #188739 #188112 #188470 #188469 #188468 Compose 2-D block local-reduce stores with physical-only finalizers such as,https://github.com/pytorch/pytorch/pull/189316,f49c0298364ed0cfb74d177e13002d344cd05da23c0264507577c65c099dd93c references,pr,189320,pr,189321,medium,pr.body,"Stack from ghstack (oldest at bottom): #189321 -> #189320 The normal Inductor invoke_subgraph path compiled each nested region under the surrounding graph's Inductor config, so per-regio",https://github.com/pytorch/pytorch/pull/189320,e0f3bd855d16bde19e867e5f57121b6bbfc5f20107189dfd82048903763fb83c references,pr,189324,issue,174929,medium,pr.comments[0].body,"ossibly due to flakiness on trunk: ⏳ inductor / inductor-cpu-test / test (cpu_inductor_torchbench, 1, 2, mt-l-x86iamx-8-64, unstable) (gh) (#174929) ⏳ inductor / inductor-cpu-test / test (cpu_inductor_torchbench, 2, 2, mt-l-x86iamx-8-64, unstable) (gh) (#174929) This comment w...",https://github.com/pytorch/pytorch/pull/189324,fd4f31a6228194bf084a4feffd109d26d513d06a2266ed9679ef2a9321e478a5 references,pr,189321,pr,189320,medium,pr.body,"Stack from ghstack (oldest at bottom): -> #189321 #189320 Building on nested Inductor configs, a region can set triton.cudagraphs to opt into or out of cudagraphs independently of the enclosing gra",https://github.com/pytorch/pytorch/pull/189321,86d0f6805a2756c5e98c5ee095a87dfc374d17e79990046aee61f964980811cd references,pr,189291,pr,189234,medium,pr.body,"Splits the MPS backend portion out of #189234 (gh-187806) per review. Problem c10::metal::erfc is 1.0 - erf(x): once erf saturates in fp32 (x ~ 3.9), erfc returns 0.0, i.e. 100% relativ",https://github.com/pytorch/pytorch/pull/189291,3e2888e299562b8015a992dd779af87daa06badfe8d402c1aac55d82d75cb94c references,pr,189291,issue,174929,medium,pr.comments[0].body,"ossibly due to flakiness on trunk: ⏳ inductor / inductor-cpu-test / test (cpu_inductor_torchbench, 1, 2, mt-l-x86iamx-8-64, unstable) (gh) (#174929) This comment was automatically generated by Dr. CI and updates every 15 minutes.",https://github.com/pytorch/pytorch/pull/189291,d1775e50ff46a4dabe1ee026fa8f807dde323a196ec4fc23e4d38bf066cd7a74 references,pr,188862,issue,188230,medium,pr.body,"crash under torch.compile(backend=""inductor"") with ValueError: The argument 'False' is not comparable., while eager mode works. Reported in #188230 (repro: threshold -> torch.eq -> torch.empty_like -> torch.minimum). The root cause is that inductor feeds boolean sympy values i...",https://github.com/pytorch/pytorch/pull/188862,c881a84fa000ea00091660491fda9eac2e9d88615d30b89bc7dd45fb8fea8157 review guidance,pr,188643,pr,188643,high,pr.reviews[0].body,Please fix lint errors,https://github.com/pytorch/pytorch/pull/188643,07252c34e9abf4213be0e497ebd7511fb18c46568d6fc6390e43d144870fd146 review guidance,pr,188643,pr,188643,high,pr.reviews[1].body,Please fix lint errors,https://github.com/pytorch/pytorch/pull/188643#pullrequestreview-4620614811,63a80a1bd960de4ae684e56f631c39bb8a85a1bdd66079341c655ff8560376bc references,pr,189109,pr,188621,medium,pr.body,"Stack from ghstack (oldest at bottom): #189185 #189083 #188621 -> #189109 The CUPTI field-id codegen (tools/gen_cupti_stubs.py, stacked on top) parses the CUPTI ABI header cupti_activity.h with libclang",https://github.com/pytorch/pytorch/pull/189109,4ac6cccbffe5142f6b18c448a822b98de10a04adade6c596859f49ecce5a5d31 references,pr,189109,pr,189083,medium,pr.body,"Stack from ghstack (oldest at bottom): #189185 #189083 #188621 -> #189109 The CUPTI field-id codegen (tools/gen_cupti_stubs.py, stacked on top) parses the CUPTI ABI header cupti_activity.h with",https://github.com/pytorch/pytorch/pull/189109,723f49df4b2808e26a96f042e9b016fd9549eeec9cdb7714994b6d18828fdcec references,pr,189109,pr,189185,medium,pr.body,"Stack from ghstack (oldest at bottom): #189185 #189083 #188621 -> #189109 The CUPTI field-id codegen (tools/gen_cupti_stubs.py, stacked on top) parses the CUPTI ABI header cupti_activity",https://github.com/pytorch/pytorch/pull/189109,795210d7b629154549f8d3a9ae39df8df2eff4fd958bf88ec3c5a63d0a9d4a35 review guidance,pr,189109,pr,188621,high,pr.reviews[4].body,"IMO if you need a single header, that might be missing in the toolkit, the right solution is to just vendor it in the third_party folder",https://github.com/pytorch/pytorch/pull/189109#pullrequestreview-4649327188,00f8aa298f2dfb3feee443ad0a150ae06310b566c0132b8073a26c71f68c3bd2 review guidance,pr,189109,pr,189083,high,pr.reviews[4].body,"IMO if you need a single header, that might be missing in the toolkit, the right solution is to just vendor it in the third_party folder",https://github.com/pytorch/pytorch/pull/189109#pullrequestreview-4649327188,996a8d4a10d853fe92ff76ace53cdcb100330a76729d88f1a04e0de658b4c698 review guidance,pr,189109,pr,189109,high,pr.reviews[4].body,"IMO if you need a single header, that might be missing in the toolkit, the right solution is to just vendor it in the third_party folder",https://github.com/pytorch/pytorch/pull/189109#pullrequestreview-4649327188,c1c88e0d7f8d051e162c3b993a4e52f7fca4d38474ba0357496d366d00544e12 review guidance,pr,189109,pr,189185,high,pr.reviews[4].body,"IMO if you need a single header, that might be missing in the toolkit, the right solution is to just vendor it in the third_party folder",https://github.com/pytorch/pytorch/pull/189109#pullrequestreview-4649327188,0cf8fee63a178f7bfaa8543c85dbc7912cf3a9c384eb033c7b5cffb45b593193 review guidance,pr,189109,pr,188621,high,pr.reviews[10].body,"Let's try to avoid including real cupti headers, taking care of versions would be hard",https://github.com/pytorch/pytorch/pull/189109#pullrequestreview-4658342017,32eff2b79822312496b24292bbefb9bc5215078dfb94f460fb7721700e807bef review guidance,pr,189109,pr,189083,high,pr.reviews[10].body,"Let's try to avoid including real cupti headers, taking care of versions would be hard",https://github.com/pytorch/pytorch/pull/189109#pullrequestreview-4658342017,4efde0c7d29de420b293e76a4760b584f1b8e3cb269c576d0a908da2927103f6 review guidance,pr,189109,pr,189109,high,pr.reviews[10].body,"Let's try to avoid including real cupti headers, taking care of versions would be hard",https://github.com/pytorch/pytorch/pull/189109#pullrequestreview-4658342017,e88885f11ad91c230e0e507d4d99fa9bc989ca733a897522055540ff03a13d44 review guidance,pr,189109,pr,189185,high,pr.reviews[10].body,"Let's try to avoid including real cupti headers, taking care of versions would be hard",https://github.com/pytorch/pytorch/pull/189109#pullrequestreview-4658342017,fca0967c7da54d8550a2807e69cd02cb2b8b31ac0a29152855861b7fbb0f4a57 references,pr,181726,pr,181727,medium,pr.body,d PRs for review. I also created a combined PR #187318 to test on CI. #181726 [xpu][1/4]Implement scaled_mm_v2 for MXFP8/MXFP4/NVFP4 on XPU #181727 [xpu][2/4]Implement scaled_mm_v1 for MXFP8/MXFP4/NVFP4 on XPU #187315 [xpu][3/4] inductor: route MX scaled_mm_v2 fallback through...,https://github.com/pytorch/pytorch/pull/181726,8b983dee6c433d51ffd2b10ef71517ff66fd92fbbda5643e63aec5d387e8941c references,pr,181726,pr,181728,medium,pr.body,[2/4]Implement scaled_mm_v1 for MXFP8/MXFP4/NVFP4 on XPU #187315 [xpu][3/4] inductor: route MX scaled_mm_v2 fallback through v2 aten kernel #181728 [xpu][4/4]Enable MXFP8/MXFP4/NVFP4 tests for XPU Authored with Claude. cc @jgong5 @mingfeima @XiaobingSuper @sanchitintel @ashoke...,https://github.com/pytorch/pytorch/pull/181726,dac9ffd6ec008ef770a55544d637dba34132e1198cf13f3a4d492cd2eb81ec4a references,pr,181726,pr,187315,medium,pr.body,#181726 [xpu][1/4]Implement scaled_mm_v2 for MXFP8/MXFP4/NVFP4 on XPU #181727 [xpu][2/4]Implement scaled_mm_v1 for MXFP8/MXFP4/NVFP4 on XPU #187315 [xpu][3/4] inductor: route MX scaled_mm_v2 fallback through v2 aten kernel #181728 [xpu][4/4]Enable MXFP8/MXFP4/NVFP4 tests for X...,https://github.com/pytorch/pytorch/pull/181726,2ca831d8f7dfcb0af3a9eca32f2fb3bcec8d366c2618af3d9b2b991748bd4faa references,pr,181726,pr,187318,medium,pr.body,"ged. PR Stack: Since I don't have ghstack permission, I manually created the following stacked PRs for review. I also created a combined PR #187318 to test on CI. #181726 [xpu][1/4]Implement scaled_mm_v2 for MXFP8/MXFP4/NVFP4 on XPU #181727 [xpu][2/4]Implement scaled_mm_v1 for...",https://github.com/pytorch/pytorch/pull/181726,efa0446e57422907f48d5b4abed675b89fad0ad7d1531926cbc3e67d8f601674 competes with,pr,189051,pr,185424,medium,pr.body,Stack from ghstack (oldest at bottom): #185424 #189053 #189052 #189022 -> #189051 #189021 set(a=1) / set().init(a=1) silently returned an empty set under Dynamo instead of raising TypeEr,https://github.com/pytorch/pytorch/pull/189051,76a6ced105822c391d7ceac5c59ceaa1adb6437c3a548a8c830caa3eca0adb4c competes with,pr,189051,pr,189021,medium,pr.body,Stack from ghstack (oldest at bottom): #185424 #189053 #189052 #189022 -> #189051 #189021 set(a=1) / set().init(a=1) silently returned an empty set under Dynamo instead of raising TypeError. BuiltinVariable.call_set / call_frozen,https://github.com/pytorch/pytorch/pull/189051,5bfb615923884d0b0b26fd44178fa540d4dc922b42aeddfeb59935b8a2d2e93d competes with,pr,189051,pr,189022,medium,pr.body,Stack from ghstack (oldest at bottom): #185424 #189053 #189052 #189022 -> #189051 #189021 set(a=1) / set().init(a=1) silently returned an empty set under Dynamo instead of raising TypeError. BuiltinVariable.cal,https://github.com/pytorch/pytorch/pull/189051,3e70660ae9213bccf8cdbae4a2fc2b3184578daaeb5fc8b57d88fdb7e890abd9 competes with,pr,189051,pr,189052,medium,pr.body,Stack from ghstack (oldest at bottom): #185424 #189053 #189052 #189022 -> #189051 #189021 set(a=1) / set().init(a=1) silently returned an empty set under Dynamo instead of raising TypeError. BuiltinVari,https://github.com/pytorch/pytorch/pull/189051,2ef2e165111634d8c83d3660c8392d7421a8a18d89417862cc3ed4dcbcf20c6e competes with,pr,189051,pr,189053,medium,pr.body,Stack from ghstack (oldest at bottom): #185424 #189053 #189052 #189022 -> #189051 #189021 set(a=1) / set().init(a=1) silently returned an empty set under Dynamo instead of raising TypeError. Bui,https://github.com/pytorch/pytorch/pull/189051,0f5d64aa1fd38de85e5dbf669c7ca6eeda4cfb97b2fd93f20f45c71e06d8601c closes,pr,189317,issue,181474,high,pr.closingIssuesReferences,pr #189317 declares a closing reference to issue #181474.,https://github.com/pytorch/pytorch/pull/189317,5d3db8785fc5d30f65a8cdcb6157868688e952f10a1dbc251dc545c5efb185bf closes,pr,189317,issue,181474,high,pr.body,(#177248) op_skips / op_decorators for operator-level skipping (#177256) skipped_testcases for test class/method skipping (#180820) Closes #181474,https://github.com/pytorch/pytorch/pull/189317,adce277bb9f03b265fee9a65fe8ee68acea7ce7a89c4ba6c30d6ba8825a46b1e closes,pr,188632,issue,188544,high,pr.body,"ethod guard sources rooted at the class descriptor, where .func is valid, instead of adding a narrow special case for opaque objects. Fixes #188544 Generated by my agent Test Plan: Reproduced the original InternalTorchDynamoError before the fix with the issue repro. python tes...",https://github.com/pytorch/pytorch/pull/188632,f4e3df3254cab793bd9ae6e403173df7e0f8594abaecba6fd6176b04a0e885fb review guidance,pr,188632,issue,188544,high,pr.reviews[0].body,"LGTM, thanks!",https://github.com/pytorch/pytorch/pull/188632,13d9ded749ec6ad1efefa4145b588add2ec2a70c36025bc813546a3e9aab9734 review guidance,pr,188632,pr,188632,high,pr.reviews[0].body,"LGTM, thanks!",https://github.com/pytorch/pytorch/pull/188632,1ccbb116ca3491737b0089fd047c254965205d8a24339c76b43438e6c3f1fa7c references,pr,179286,issue,176662,medium,pr.body,"implify some template code. In this PR, I have changed some instances of enable_if with concepts where it's not hard to verify correctness. #176662 cc @EikanWang @jgong5",https://github.com/pytorch/pytorch/pull/179286,c7edc518efea462315377acd6dc641e55375ec124230133880b5bc3be330cbfc references,pr,178393,pr,179094,medium,pr.body,Stack from ghstack (oldest at bottom): -> #178393 #179094 #188029 #187631 #189306,https://github.com/pytorch/pytorch/pull/178393,73b4c6aff8257776692d5bba92ead4ba95654094fda45c86da8ff71e4557ee63 references,pr,178393,pr,187631,medium,pr.body,Stack from ghstack (oldest at bottom): -> #178393 #179094 #188029 #187631 #189306,https://github.com/pytorch/pytorch/pull/178393,6230277ef2325cc31249d0f5d4ab58707c75962e0729bf652d78711b927dfdbc references,pr,178393,pr,188029,medium,pr.body,Stack from ghstack (oldest at bottom): -> #178393 #179094 #188029 #187631 #189306,https://github.com/pytorch/pytorch/pull/178393,0c81acfd02088c5f436c91668e1f524a6f521ee2edf35a91f6ee7b127f3e7b20 references,pr,178393,pr,189306,medium,pr.body,Stack from ghstack (oldest at bottom): -> #178393 #179094 #188029 #187631 #189306,https://github.com/pytorch/pytorch/pull/178393,307ac2616c91846e09f2cebe499f3f7c13a58883e94e8ba490dc42017f20a0fe review guidance,pr,188742,pr,188742,high,pr.reviews[0].body,address automated feedback from claude,https://github.com/pytorch/pytorch/pull/188742,4cdb42e30de7d90b0605b1da5f4bda0a0760f0529c100be2ffb383d51e9b3d6d references,pr,186055,pr,186056,medium,pr.body,Stack from ghstack (oldest at bottom): #186056 -> #186055 Motivation is to allow us to handle some forms of dynamic shapes inside of a cuda graph. This does not necessarily improve perfo,https://github.com/pytorch/pytorch/pull/186055,144c19db0d4e286672c029c98857901df67b828497fb09ed3cf67ec20e2affd9 review guidance,pr,186055,pr,186055,high,pr.reviews[0].body,"minimal changes from previously accepted pr - #140979 , accepting",https://github.com/pytorch/pytorch/pull/186055,f49e1236ecd5b78d8e1ac4bec6368367d65afee8d61d5f885d1fe65c41521108 review guidance,pr,186055,pr,186056,high,pr.reviews[0].body,"minimal changes from previously accepted pr - #140979 , accepting",https://github.com/pytorch/pytorch/pull/186055,0468ffc0be1ad45ae5444d7351081137a41a652ac815664ae1836a48a4aa178b references,pr,189088,pr,187465,medium,pr.body,Stack from ghstack (oldest at bottom): #187465 -> #189088 We place the signal pad at the front of symmetric memory allocations in this PR. The purpose is to fix the potential signal pad,https://github.com/pytorch/pytorch/pull/189088,974ac5098372d16d2a152de137a69c83918fb0c63870c2dcea9a2ff356d18c2a review guidance,pr,189088,pr,187465,high,pr.reviews[0].body,I just reviewed the CUDA part. Will continue later.,https://github.com/pytorch/pytorch/pull/189088,bd7d2a52e9868b0adc5ebf780be1f52f2614720d3fd2d90f01b24911b2107fe7 review guidance,pr,189088,pr,189088,high,pr.reviews[0].body,I just reviewed the CUDA part. Will continue later.,https://github.com/pytorch/pytorch/pull/189088,ef4e292549edad75600b615034dc2c31377d3ab7e74e64d747e346f99f7afbab review guidance,pr,189088,pr,187465,high,pr.reviews[2].body,I just reviewed the CUDA part. Will continue later.,https://github.com/pytorch/pytorch/pull/189088#pullrequestreview-4640495139,44b52d9c46efe027737d6c0de4ec550bd70a3d269eb104ad3b708c36bb414f92 review guidance,pr,189088,pr,189088,high,pr.reviews[2].body,I just reviewed the CUDA part. Will continue later.,https://github.com/pytorch/pytorch/pull/189088#pullrequestreview-4640495139,6853c294b7acc670bbab5666f9d16c2aa0bceba6165d3caccd59595efe884e6d references,pr,187631,pr,178393,medium,pr.body,Stack from ghstack (oldest at bottom): #178393 #179094 #188029 -> #187631 #189306 Reusing the name of the iterator variable in a list comprehension that graph breaks to store the result,https://github.com/pytorch/pytorch/pull/187631,2c1ac7dce71f1c9a7bac2890368210cc7997134709466e29fd2a4ad74fb7f534 references,pr,187631,pr,179094,medium,pr.body,Stack from ghstack (oldest at bottom): #178393 #179094 #188029 -> #187631 #189306 Reusing the name of the iterator variable in a list comprehension that graph breaks to store the result causes a,https://github.com/pytorch/pytorch/pull/187631,85b89c8a839916c3cc6e23e11d78f179b6b91ae34805432e9f8cd3c1f07b3a10 references,pr,187631,pr,188029,medium,pr.body,Stack from ghstack (oldest at bottom): #178393 #179094 #188029 -> #187631 #189306 Reusing the name of the iterator variable in a list comprehension that graph breaks to store the result causes a segfaul,https://github.com/pytorch/pytorch/pull/187631,4b15524da96b67f0ad75089580ff55eb5db15e4ddd452cfa47147d3038c265f7 references,pr,187631,pr,189306,medium,pr.body,Stack from ghstack (oldest at bottom): #178393 #179094 #188029 -> #187631 #189306 Reusing the name of the iterator variable in a list comprehension that graph breaks to store the result causes a segfault (the stack underf,https://github.com/pytorch/pytorch/pull/187631,5420621bcac2660cbf6f008e7ba4d65564c9d4afb0af6827f842dca7ee2dd591 review guidance,pr,187631,pr,178393,high,pr.reviews[0].body,There's one test failure. Can you check?,https://github.com/pytorch/pytorch/pull/187631,f673d0ec37031074bd7f385f9f4981f5c3df7117c257c8b8c8a8929fd27cf3d7 review guidance,pr,187631,pr,179094,high,pr.reviews[0].body,There's one test failure. Can you check?,https://github.com/pytorch/pytorch/pull/187631,dfc8dbe946536ba9186fe45bcc67b0c9097a6c121d8f94a391d14d9aa38dc6eb review guidance,pr,187631,pr,187631,high,pr.reviews[0].body,There's one test failure. Can you check?,https://github.com/pytorch/pytorch/pull/187631,3e28eb8a302103e94a4b2521629e065dd5996e83915b32e5a50c3dab3b888f56 review guidance,pr,187631,pr,188029,high,pr.reviews[0].body,There's one test failure. Can you check?,https://github.com/pytorch/pytorch/pull/187631,89dd682e59c298d637f35da3bd34775e4751e69c90691f5dbc477d69d48fbaef review guidance,pr,187631,pr,189306,high,pr.reviews[0].body,There's one test failure. Can you check?,https://github.com/pytorch/pytorch/pull/187631,3eeb888404684fef9bab9efcb540e94d6def0baa69070d25eea1c9b7c209c192 references,pr,188029,pr,178393,medium,pr.body,Stack from ghstack (oldest at bottom): #178393 #179094 -> #188029 #187631 #189306 Cell variables that share a name with a local variable are stored separately in localsplus in cpython. T,https://github.com/pytorch/pytorch/pull/188029,13fcbf3bb16234dc4eeeaa7779e6bc3268373574aa5151cc23c72b763a8fb21e references,pr,188029,pr,179094,medium,pr.body,Stack from ghstack (oldest at bottom): #178393 #179094 -> #188029 #187631 #189306 Cell variables that share a name with a local variable are stored separately in localsplus in cpython. This dist,https://github.com/pytorch/pytorch/pull/188029,a600bc8f193ba981a62b078d0e44e9da002be279c7ed503bc0eb0e059a80a5ce references,pr,188029,pr,187631,medium,pr.body,Stack from ghstack (oldest at bottom): #178393 #179094 -> #188029 #187631 #189306 Cell variables that share a name with a local variable are stored separately in localsplus in cpython. This distinction is lost in,https://github.com/pytorch/pytorch/pull/188029,30d63c88d2df8cd009b0621207fa4e2f9214848a3f8f03faf0edfdb873193366 references,pr,188029,pr,189306,medium,pr.body,"Stack from ghstack (oldest at bottom): #178393 #179094 -> #188029 #187631 #189306 Cell variables that share a name with a local variable are stored separately in localsplus in cpython. This distinction is lost in dynamo,",https://github.com/pytorch/pytorch/pull/188029,130ae2d004570891404c1caf259bf4fb7626308c8c3c1dc0e021c2d9a75ef211 references,pr,183328,issue,76324,medium,pr.body,"Issue Partially addresses #76324. Summary Adds arbitrary-order modified Bessel functions: torch.special.modified_bessel_i(x, nu) torch.special.modified_bessel_k(x, nu) This",https://github.com/pytorch/pytorch/pull/183328,3088dde11ef6dc925ee5e5cda09474a3de33eba18183a4aa3dbe31bf859b84bc references,pr,179094,pr,178393,medium,pr.body,Stack from ghstack (oldest at bottom): #178393 -> #179094 #188029 #187631 #189306 Authored with Claude.,https://github.com/pytorch/pytorch/pull/179094,be7b4ca137c72844a903c5466040f726600b86faaa343aeaa846cc8762e3cb41 references,pr,179094,pr,187631,medium,pr.body,Stack from ghstack (oldest at bottom): #178393 -> #179094 #188029 #187631 #189306 Authored with Claude.,https://github.com/pytorch/pytorch/pull/179094,2a666960908321ec7637ec6ac268a3c9e92123aa3ef0bee2b30e05fd20a6822e references,pr,179094,pr,188029,medium,pr.body,Stack from ghstack (oldest at bottom): #178393 -> #179094 #188029 #187631 #189306 Authored with Claude.,https://github.com/pytorch/pytorch/pull/179094,a20c1f0c686a2a576e5673c0e7e15217970b47eb59744377f57f02f17bb4645e references,pr,179094,pr,189306,medium,pr.body,Stack from ghstack (oldest at bottom): #178393 -> #179094 #188029 #187631 #189306 Authored with Claude.,https://github.com/pytorch/pytorch/pull/179094,32b8a35b4b3a387938ccb32d9a05ddc757b8395135e5db911b80e4a83d6d82ee closes,pr,186252,issue,184408,high,pr.closingIssuesReferences,pr #186252 declares a closing reference to issue #184408.,https://github.com/pytorch/pytorch/pull/186252,0c6abc4b8f3283dacd74c037c16542fb8a128f170c51d94560ffe3542c22f0dd closes,pr,186252,issue,184408,high,pr.body,"Fixes #184408 Triton's approximate fp32 division can produce results slightly below the true quotient, causing trunc(a / b) to be off by one when the quo",https://github.com/pytorch/pytorch/pull/186252,8c3ebafb5a10489a5d7535872aa866048e2198cd61fc1f594b086f872e56e94b closes,pr,189313,issue,102948,high,pr.closingIssuesReferences,pr #189313 declares a closing reference to issue #102948.,https://github.com/pytorch/pytorch/pull/189313,55efa1aafd2bef59e2dd1106c234d46abf4c818a6f82189053f737f09c027f07 references,pr,189286,pr,188100,medium,pr.body,Stack from ghstack (oldest at bottom): -> #189286 #188573 #188299 #188100 Fixes #189133 This PR addresses flaky tests by fixing some DeviceContext mode leaks that may randomly affect other tests (e.g. the vmap tes,https://github.com/pytorch/pytorch/pull/189286,0fbdd19e1f1cb771d5009f4f83b7d8b1bcabe2e379f744b9f52458049d48a79c references,pr,189286,pr,188299,medium,pr.body,Stack from ghstack (oldest at bottom): -> #189286 #188573 #188299 #188100 Fixes #189133 This PR addresses flaky tests by fixing some DeviceContext mode leaks that may randomly affect other tests (e.g. the,https://github.com/pytorch/pytorch/pull/189286,553511142840e258b04350dbbc573b30bc27c18dea98460230481de7f4d2c308 references,pr,189286,pr,188573,medium,pr.body,Stack from ghstack (oldest at bottom): -> #189286 #188573 #188299 #188100 Fixes #189133 This PR addresses flaky tests by fixing some DeviceContext mode leaks that may randomly affect other tests (e,https://github.com/pytorch/pytorch/pull/189286,ecd4948475c05bafc816687326d7d01774dcfa25910243159edae3ae5d02f85e closes,pr,189286,issue,189133,high,pr.body,Stack from ghstack (oldest at bottom): -> #189286 #188573 #188299 #188100 Fixes #189133 This PR addresses flaky tests by fixing some DeviceContext mode leaks that may randomly affect other tests (e.g. the vmap test disabled by,https://github.com/pytorch/pytorch/pull/189286,72eb56af4af98cfdd94d23755221d8521124d50010da6ebdcefbb401d7348a23 review guidance,pr,189286,pr,188100,high,pr.reviews[1].body,Thanks!,https://github.com/pytorch/pytorch/pull/189286,18e4a65324ed97d75c91a366e508c11bb4d3cdd6818b08f1aeb10cc21f85a7f4 review guidance,pr,189286,pr,188299,high,pr.reviews[1].body,Thanks!,https://github.com/pytorch/pytorch/pull/189286,b0452b2aede1d71135cc8572d6a94dd01e0777b9afc54c0387e1c923435c7a52 review guidance,pr,189286,pr,188573,high,pr.reviews[1].body,Thanks!,https://github.com/pytorch/pytorch/pull/189286,78f914fbdfb9cff8c8f338709f6d0f4f00229e263f0e99456ca9c5e6680bd9e8 review guidance,pr,189286,issue,189133,high,pr.reviews[1].body,Thanks!,https://github.com/pytorch/pytorch/pull/189286,b3400615745923eb8d64af30a5ffdcdb7cef6dbe3bfaad605ddaa435a3664d72 review guidance,pr,189286,pr,189286,high,pr.reviews[1].body,Thanks!,https://github.com/pytorch/pytorch/pull/189286,97995fa2f0376b929c2a888803273dba548d259498734899ef050afb08c7322b references,pr,189306,pr,178393,medium,pr.body,"Stack from ghstack (oldest at bottom): #178393 #179094 #188029 #187631 -> #189306 Two changes to update this for 3.15 co_lnotab has been deprecated since 3.10, and was removed in 3.15. I",https://github.com/pytorch/pytorch/pull/189306,c2ee162b2ef1bdb683a75590ba05579091c5d9f4d74df7219ac485ae20c61124 references,pr,189306,pr,179094,medium,pr.body,"Stack from ghstack (oldest at bottom): #178393 #179094 #188029 #187631 -> #189306 Two changes to update this for 3.15 co_lnotab has been deprecated since 3.10, and was removed in 3.15. It was re",https://github.com/pytorch/pytorch/pull/189306,b09153da5265d97a96617696f8c741428b736059cdba41cb656189d501801a25 references,pr,189306,pr,187631,medium,pr.body,"Stack from ghstack (oldest at bottom): #178393 #179094 #188029 #187631 -> #189306 Two changes to update this for 3.15 co_lnotab has been deprecated since 3.10, and was removed in 3.15. It was replaced by co_lin",https://github.com/pytorch/pytorch/pull/189306,1411851df0cbd78be2e3cafcec8f01d79175868890141517f17ed58ff691959e references,pr,189306,pr,188029,medium,pr.body,"Stack from ghstack (oldest at bottom): #178393 #179094 #188029 #187631 -> #189306 Two changes to update this for 3.15 co_lnotab has been deprecated since 3.10, and was removed in 3.15. It was replaced b",https://github.com/pytorch/pytorch/pull/189306,96b8e2c2b388f5baa2574e8a1e2de5bbd75184761b6fcb5c082ababc623aa117 references,pr,189297,issue,188477,medium,pr.body,"cutlass._mlir (present in 4.5.2), so it is known-compatible with the gated cutlass-dsl. This is the same cutlass-dsl version-skew family as #188477. Test Plan: CI only (B200). The ""B200 Smoke Tests / ...sm100 / test (smoke_b200)"" job should go green; test_flex_flash no longer...",https://github.com/pytorch/pytorch/pull/189297,00a4739ea61e20e18aec7f2006e255a39e7c98635dbcec1adb90bf8f3702cf12 references,pr,189319,issue,106571,medium,pr.body,"bear B007 (unused loop control variable), one of the remaining bugbear codes still suppressed in the ignore list of pyproject.toml. Part of #106571. This removes B007 from the ignore list and fixes all 169 violations across 110 files so the lint is enforced going forward. Appr...",https://github.com/pytorch/pytorch/pull/189319,a3d5d0e35eb231be0c2c9c3fa4e2bc0342a326d6a907996e32b83cfc415c4f86 review guidance,pr,188600,pr,188600,high,pr.reviews[0].body,address automated feedback from claude,https://github.com/pytorch/pytorch/pull/188600,ee1efd4361d60613c3ffdfb3c9f8eab130eed08e375090ce45ff66b769c39e04 references,pr,181781,pr,181780,medium,pr.body,Stack from ghstack (oldest at bottom): -> #181781 #188176 #188175 #181780 cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv @jiayisunx @kadeng @chauhang @a,https://github.com/pytorch/pytorch/pull/181781,7714e2059eb050a3c1e288984595aaf54bc0cf726892158718824634617e4e89 references,pr,181781,pr,188175,medium,pr.body,Stack from ghstack (oldest at bottom): -> #181781 #188176 #188175 #181780 cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv @jiayisunx @kadeng @cha,https://github.com/pytorch/pytorch/pull/181781,80dfaf405a2e7f3aa30b6f0d00dc83784d3d6333ef52c0963b7607682a67c031 references,pr,181781,pr,188176,medium,pr.body,Stack from ghstack (oldest at bottom): -> #181781 #188176 #188175 #181780 cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv @jiayisunx @kad,https://github.com/pytorch/pytorch/pull/181781,b64a6953f25864a70e9b7ccae446ca6511bb892b6fc998780b1369b2216b131e review guidance,pr,188137,pr,188137,high,pr.reviews[0].body,Can you explain more about why this is needed? What is the use case? Where/who will use this? What other options were considered? What does this enable or make better?,https://github.com/pytorch/pytorch/pull/188137,3d1b99dfdfe0a159f2ce7cf5aabc51f8866e728cc17c0d4919cee953b8dac3da review guidance,pr,188137,pr,188137,high,pr.reviews[1].body,"(Reviewed by me, assisted by AI) [blocker] The description no longer matches the diff, and closing that gap is exactly what jansel's open CHANGES_REQUESTED is asking for: It describes ""an opt-in Inductor post-grad pass,"" but the current implementation is a regular @register_decomposition in torch...",https://github.com/pytorch/pytorch/pull/188137,db3846058d28a7763456605d1df5d2d99206a0c3d817facd2b0159cb64101c88 references,pr,189022,pr,185424,medium,pr.body,"Stack from ghstack (oldest at bottom): #185424 #189053 #189052 -> #189022 #189051 #189021 Two itertools.count object-protocol gaps, both mirroring CPython Modules/itertoolsmodule.c: Cons",https://github.com/pytorch/pytorch/pull/189022,691260dbc8f10f9bf73cf9a5c51d30658f4bbce6b5511dd68b627cedb4178071 references,pr,189022,pr,189021,medium,pr.body,"Stack from ghstack (oldest at bottom): #185424 #189053 #189052 -> #189022 #189051 #189021 Two itertools.count object-protocol gaps, both mirroring CPython Modules/itertoolsmodule.c: Construction: the count branch had a not kwargs",https://github.com/pytorch/pytorch/pull/189022,3c6b66c877f0f6ddcf2d5559c6809b9a5d8a94b64130a25749737c15afa045e9 references,pr,189022,pr,189051,medium,pr.body,"Stack from ghstack (oldest at bottom): #185424 #189053 #189052 -> #189022 #189051 #189021 Two itertools.count object-protocol gaps, both mirroring CPython Modules/itertoolsmodule.c: Construction: the count branch had a no",https://github.com/pytorch/pytorch/pull/189022,8d3379b864ed6dceac9c2e465dd31244ba8fe7b12b99f1fbde43bc0176902f00 references,pr,189022,pr,189052,medium,pr.body,"Stack from ghstack (oldest at bottom): #185424 #189053 #189052 -> #189022 #189051 #189021 Two itertools.count object-protocol gaps, both mirroring CPython Modules/itertoolsmodule.c: Construction: the co",https://github.com/pytorch/pytorch/pull/189022,18f1e66dd00cb87f201abd75affdbebc942100d58f83a4331dd03f59a5cbab06 references,pr,189022,pr,189053,medium,pr.body,"Stack from ghstack (oldest at bottom): #185424 #189053 #189052 -> #189022 #189051 #189021 Two itertools.count object-protocol gaps, both mirroring CPython Modules/itertoolsmodule.c: Construction",https://github.com/pytorch/pytorch/pull/189022,5a773fe9ed4996f40e30a7726cb4adca0b21120bb2d4687a9927bde5b23559d8 closes,pr,188006,issue,168868,high,pr.closingIssuesReferences,pr #188006 declares a closing reference to issue #168868.,https://github.com/pytorch/pytorch/pull/188006,7d38cc63ffcddfbe29ef142935e20633b3950bec0ace41bf4081323bbb859781 closes,pr,188006,issue,168868,high,pr.body,Fixes #168868 Investigation summary Test environment PyTorch version: 2.13.0a0+git75d18bb Hip version: 7.2.53211 GPU 0 name: AMD Instinct MI355X Command,https://github.com/pytorch/pytorch/pull/188006,35d68377e9fb453e6956c2eab42e4b5cad4b1bd1fa02c6d1982acc3bcbdaf145 review guidance,pr,188006,issue,168868,high,pr.reviews[0].body,"Nice investigation, but a few things before this lands: Keep the DeviceSqrt.cuh cleanup, but stop crediting it as a fix. I diffed the GCN ISA of the old ::sqrt/::sqrtf specializations vs the new std::sqrt template (HIP 7.2, gfx90a/942/950): instruction streams are byte-identical. device_sqrt #189092 #189091 #187707 #187532 #187531 #187469 #187468 Fix object_generic_getattr step 7 to raise ObservedAttributeError instead of _UnhandledDescriptorError when an attribute genuinely doesn't,https://github.com/pytorch/pytorch/pull/189092,3923e719c3862332c8a1fd5ad572a68f2bad51f9327b8a58a46c39fa551d9747 competes with,pr,189092,pr,187469,medium,pr.body,Stack from ghstack (oldest at bottom): -> #189092 #189091 #187707 #187532 #187531 #187469 #187468 Fix object_generic_getattr step 7 to raise ObservedAttributeError instead of _UnhandledDescriptorError when an attribute genuinely,https://github.com/pytorch/pytorch/pull/189092,36d8e29c5fcfbde4a9198ce00bbd0ea1287146993588e29d1009e45acd81c0d2 competes with,pr,189092,pr,187531,medium,pr.body,Stack from ghstack (oldest at bottom): -> #189092 #189091 #187707 #187532 #187531 #187469 #187468 Fix object_generic_getattr step 7 to raise ObservedAttributeError instead of _UnhandledDescriptorError when an attribute ge,https://github.com/pytorch/pytorch/pull/189092,e9d9e90c83b36f29e10b032bc16c4984ed9bbf22e2f05fe11a4f4511cb369ccf competes with,pr,189092,pr,187532,medium,pr.body,Stack from ghstack (oldest at bottom): -> #189092 #189091 #187707 #187532 #187531 #187469 #187468 Fix object_generic_getattr step 7 to raise ObservedAttributeError instead of _UnhandledDescriptorError when an attr,https://github.com/pytorch/pytorch/pull/189092,e265f19ed2efb4bfe3db005c44d905866cfc44e8d4b7b801103f19768cac70d0 competes with,pr,189092,pr,187707,medium,pr.body,Stack from ghstack (oldest at bottom): -> #189092 #189091 #187707 #187532 #187531 #187469 #187468 Fix object_generic_getattr step 7 to raise ObservedAttributeError instead of _UnhandledDescriptorError when,https://github.com/pytorch/pytorch/pull/189092,2b19edc418ea596ff55782a11f176b0b53d17026df834499ddb10c39a87a2a4e competes with,pr,189092,pr,189091,medium,pr.body,Stack from ghstack (oldest at bottom): -> #189092 #189091 #187707 #187532 #187531 #187469 #187468 Fix object_generic_getattr step 7 to raise ObservedAttributeError instead of _UnhandledDescriptorEr,https://github.com/pytorch/pytorch/pull/189092,bc25d1b1ff281e98ed9af0a9439a3fb5a4398a2e0791129f1673729078b24d39 references,pr,189092,issue,174929,medium,pr.comments[0].body,"possibly due to flakiness on trunk: inductor / inductor-cpu-test / test (cpu_inductor_torchbench, 1, 2, mt-l-x86iamx-8-64, unstable) (gh) (#174929) doctr_reco_predictor inductor / inductor-cpu-test / test (cpu_inductor_torchbench, 2, 2, mt-l-x86iamx-8-64, unstable) (gh) (#1749...",https://github.com/pytorch/pytorch/pull/189092,5ab5222f9dc9733eac3fa5962c9d96391bc27391026ab48f8de23b635f213d27 references,pr,187707,pr,187468,medium,pr.body,"Stack from ghstack (oldest at bottom): #189092 #189091 -> #187707 #187532 #187531 #187469 #187468 GetAttrVariable was a catch-all fallback that deferred attribute access by wrapping it in a proxy. It made it impossible to distinguish ""at",https://github.com/pytorch/pytorch/pull/187707,eb670f60443348b32a9f8161ca6a09731f48bb5dd837ed6d128c7f0882d56438 references,pr,187707,pr,187469,medium,pr.body,Stack from ghstack (oldest at bottom): #189092 #189091 -> #187707 #187532 #187531 #187469 #187468 GetAttrVariable was a catch-all fallback that deferred attribute access by wrapping it in a proxy. It made it impossible to disting,https://github.com/pytorch/pytorch/pull/187707,3bb2d67b867805e8fe9d9bfb675d52c2cc00067d09d8dd063154dcb97f5b1f1f references,pr,187707,pr,187531,medium,pr.body,Stack from ghstack (oldest at bottom): #189092 #189091 -> #187707 #187532 #187531 #187469 #187468 GetAttrVariable was a catch-all fallback that deferred attribute access by wrapping it in a proxy. It made it impossible to,https://github.com/pytorch/pytorch/pull/187707,bed6543f410ea3ad11c2ea23d716f9e24491846ce69e7f77099003e346666706 references,pr,187707,pr,187532,medium,pr.body,Stack from ghstack (oldest at bottom): #189092 #189091 -> #187707 #187532 #187531 #187469 #187468 GetAttrVariable was a catch-all fallback that deferred attribute access by wrapping it in a proxy. It made it impos,https://github.com/pytorch/pytorch/pull/187707,73584a2027a2f1cbc56129aa86570d31dab41aa726629dd6629f6497a1201e31 references,pr,187707,pr,189091,medium,pr.body,Stack from ghstack (oldest at bottom): #189092 #189091 -> #187707 #187532 #187531 #187469 #187468 GetAttrVariable was a catch-all fallback that deferred attribute access by wrapping it in a prox,https://github.com/pytorch/pytorch/pull/187707,8025a628ecfb027620b3b83656966f9b1c2e626dc3b3a717e3fff57b0fdc1036 references,pr,187707,pr,189092,medium,pr.body,Stack from ghstack (oldest at bottom): #189092 #189091 -> #187707 #187532 #187531 #187469 #187468 GetAttrVariable was a catch-all fallback that deferred attribute access by wrapping it i,https://github.com/pytorch/pytorch/pull/187707,6d3c754ed06ddeee0eba0ee44066f33d3d44cde1b0b655e88db656a840aa1397 references,pr,187707,issue,174929,medium,pr.comments[0].body,"possibly due to flakiness on trunk: inductor / inductor-cpu-test / test (cpu_inductor_torchbench, 1, 2, mt-l-x86iamx-8-64, unstable) (gh) (#174929) doctr_reco_predictor inductor / inductor-cpu-test / test (cpu_inductor_torchbench, 2, 2, mt-l-x86iamx-8-64, unstable) (gh) (#1749...",https://github.com/pytorch/pytorch/pull/187707,647c41c72953e5ae0c7b1db921ca64085eba872d7ba59245b840be2f5b746bee references,pr,189091,pr,187468,medium,pr.body,Stack from ghstack (oldest at bottom): #189092 -> #189091 #187707 #187532 #187531 #187469 #187468 GetAttrBuiltinVariable.call_function has a constant-fold fallback that fires when getattro_impl raises Unsupported and all args are python,https://github.com/pytorch/pytorch/pull/189091,2408d622025d9c12746cb58a9081ca01c03e29be23527ec54655728b773bb487 references,pr,189091,pr,187469,medium,pr.body,Stack from ghstack (oldest at bottom): #189092 -> #189091 #187707 #187532 #187531 #187469 #187468 GetAttrBuiltinVariable.call_function has a constant-fold fallback that fires when getattro_impl raises Unsupported and all args are,https://github.com/pytorch/pytorch/pull/189091,77b77b7892db6f11b3296975d48416706d321ec62445a2f0a3ad2e72ee425c78 references,pr,189091,pr,187531,medium,pr.body,Stack from ghstack (oldest at bottom): #189092 -> #189091 #187707 #187532 #187531 #187469 #187468 GetAttrBuiltinVariable.call_function has a constant-fold fallback that fires when getattro_impl raises Unsupported and all,https://github.com/pytorch/pytorch/pull/189091,6ab0ad07ec53b0312c4085f7a243360072182d1a47b377813deb90113270ed44 references,pr,189091,pr,187532,medium,pr.body,Stack from ghstack (oldest at bottom): #189092 -> #189091 #187707 #187532 #187531 #187469 #187468 GetAttrBuiltinVariable.call_function has a constant-fold fallback that fires when getattro_impl raises Unsupported,https://github.com/pytorch/pytorch/pull/189091,f76ed610babe94d0e8473fb9a7675661241b5fc8085250802da139e28b4cb22b references,pr,189091,pr,187707,medium,pr.body,Stack from ghstack (oldest at bottom): #189092 -> #189091 #187707 #187532 #187531 #187469 #187468 GetAttrBuiltinVariable.call_function has a constant-fold fallback that fires when getattro_impl raises Unsu,https://github.com/pytorch/pytorch/pull/189091,b6106eedf5b8db9117ab3f6f3a72f5b9ca5a869454d167469bc6dbc3b5515b41 references,pr,189091,pr,189092,medium,pr.body,Stack from ghstack (oldest at bottom): #189092 -> #189091 #187707 #187532 #187531 #187469 #187468 GetAttrBuiltinVariable.call_function has a constant-fold fallback that fires when getatt,https://github.com/pytorch/pytorch/pull/189091,ece3376eaf031ba5bec88831368113df2754f8053d5311807a7ba1abf5c30e05 references,pr,189091,issue,174929,medium,pr.comments[0].body,"possibly due to flakiness on trunk: inductor / inductor-cpu-test / test (cpu_inductor_torchbench, 1, 2, mt-l-x86iamx-8-64, unstable) (gh) (#174929) doctr_reco_predictor inductor / inductor-cpu-test / test (cpu_inductor_torchbench, 2, 2, mt-l-x86iamx-8-64, unstable) (gh) (#1749...",https://github.com/pytorch/pytorch/pull/189091,771f33519dca47b7c695f7ac48e67011aaffd9a1e4a61578be04e86646cda085 references,pr,188978,pr,188979,medium,pr.body,specifically (ie in fake_tensor.py) but this should be removed after C++ FakeTensor fully takes over Stack from ghstack (oldest at bottom): #188979 -> #188978 cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv @jiayisunx...,https://github.com/pytorch/pytorch/pull/188978,cd4348ee668780b686aac8a8884ff4ed3a0d36a88e211133eb2dd94a106a60a2 references,pr,187531,pr,187468,medium,pr.body,Stack from ghstack (oldest at bottom): #189092 #189091 #187707 #187532 -> #187531 #187469 #187468 Adds a get_value_for_setattr() opt-in hook on the base VariableTracker that enables attribute mutation (setattr/getattr/hasattr) for VTs be,https://github.com/pytorch/pytorch/pull/187531,29960ec43e875f7f4af8ae09b7dcc151c7d22f85dad725e2131468411bb57580 references,pr,187531,pr,187469,medium,pr.body,Stack from ghstack (oldest at bottom): #189092 #189091 #187707 #187532 -> #187531 #187469 #187468 Adds a get_value_for_setattr() opt-in hook on the base VariableTracker that enables attribute mutation (setattr/getattr/hasattr) fo,https://github.com/pytorch/pytorch/pull/187531,bd47c1e486896e3695c83cfbf5f8ef73e31bd902c42a1f00408cca0e5d354157 references,pr,187531,pr,187532,medium,pr.body,Stack from ghstack (oldest at bottom): #189092 #189091 #187707 #187532 -> #187531 #187469 #187468 Adds a get_value_for_setattr() opt-in hook on the base VariableTracker that enables attribute mutation (setattr/,https://github.com/pytorch/pytorch/pull/187531,1d3bc1bbf14d70f60617be1b97be3ae43edd48c070861fcb785610d838514d14 references,pr,187531,pr,187707,medium,pr.body,Stack from ghstack (oldest at bottom): #189092 #189091 #187707 #187532 -> #187531 #187469 #187468 Adds a get_value_for_setattr() opt-in hook on the base VariableTracker that enables attribute mutation (,https://github.com/pytorch/pytorch/pull/187531,68b42ce8fb70259dba4908d575721933cf4c8fece52eedebf99d416d250741bf references,pr,187531,pr,189091,medium,pr.body,Stack from ghstack (oldest at bottom): #189092 #189091 #187707 #187532 -> #187531 #187469 #187468 Adds a get_value_for_setattr() opt-in hook on the base VariableTracker that enables attribute mu,https://github.com/pytorch/pytorch/pull/187531,67c9f240231ea7de9222650c0743fa76e12477a1d0acab8e1046699515d0072d references,pr,187531,pr,189092,medium,pr.body,Stack from ghstack (oldest at bottom): #189092 #189091 #187707 #187532 -> #187531 #187469 #187468 Adds a get_value_for_setattr() opt-in hook on the base VariableTracker that enables attr,https://github.com/pytorch/pytorch/pull/187531,93b8356246dbbfb6426b7c0c1dc53742d7a7f62c1e47bc00290cc93c1cbbdc35 references,pr,187531,issue,174929,medium,pr.comments[0].body,"possibly due to flakiness on trunk: inductor / inductor-cpu-test / test (cpu_inductor_torchbench, 1, 2, mt-l-x86iamx-8-64, unstable) (gh) (#174929) doctr_reco_predictor inductor / inductor-cpu-test / test (cpu_inductor_torchbench, 2, 2, mt-l-x86iamx-8-64, unstable) (gh) (#1749...",https://github.com/pytorch/pytorch/pull/187531,1981492d2865ba4234d69cca555dfb07f3cb9ad050181062d4598eda1072c7cc references,pr,187532,pr,187468,medium,pr.body,"Stack from ghstack (oldest at bottom): #189092 #189091 #187707 -> #187532 #187531 #187469 #187468 Previously, both obj.__getattr__(""x"") and obj.__getattribute__(""x"") in call_method dispatched through getattro_impl, which does full Generi",https://github.com/pytorch/pytorch/pull/187532,ff4599e029abd3377bdd14e8cabcb68544ae30c0645640c154bb900fda9054a9 references,pr,187532,pr,187469,medium,pr.body,"Stack from ghstack (oldest at bottom): #189092 #189091 #187707 -> #187532 #187531 #187469 #187468 Previously, both obj.__getattr__(""x"") and obj.__getattribute__(""x"") in call_method dispatched through getattro_impl, which does ful",https://github.com/pytorch/pytorch/pull/187532,72ba02f6c8db7c33133a1203e3938ec6537838712b6426499689d80f276f8a34 references,pr,187532,pr,187531,medium,pr.body,"Stack from ghstack (oldest at bottom): #189092 #189091 #187707 -> #187532 #187531 #187469 #187468 Previously, both obj.__getattr__(""x"") and obj.__getattribute__(""x"") in call_method dispatched through getattro_impl, which",https://github.com/pytorch/pytorch/pull/187532,7e289ef72012b40d05dbbeb95ef97780387c6f4589e3437df3b50035f3181f80 references,pr,187532,pr,187707,medium,pr.body,"Stack from ghstack (oldest at bottom): #189092 #189091 #187707 -> #187532 #187531 #187469 #187468 Previously, both obj.__getattr__(""x"") and obj.__getattribute__(""x"") in call_method dispatched through ge",https://github.com/pytorch/pytorch/pull/187532,eaa3121e72ae1c0a5547526ed3e3aede48359586804631436ca73d1b321beae4 references,pr,187532,pr,189091,medium,pr.body,"Stack from ghstack (oldest at bottom): #189092 #189091 #187707 -> #187532 #187531 #187469 #187468 Previously, both obj.__getattr__(""x"") and obj.__getattribute__(""x"") in call_method dispatched th",https://github.com/pytorch/pytorch/pull/187532,f800e8c35ae84acad3ee8ee2a98aab86c75eea486c96ba0eaea5a30b5d506b95 references,pr,187532,pr,189092,medium,pr.body,"Stack from ghstack (oldest at bottom): #189092 #189091 #187707 -> #187532 #187531 #187469 #187468 Previously, both obj.__getattr__(""x"") and obj.__getattribute__(""x"") in call_method dispa",https://github.com/pytorch/pytorch/pull/187532,a90a315df3633cf3a4980e96343943c01bb21e1e7e957e2fa07bc6054705334e references,pr,187532,issue,174929,medium,pr.comments[0].body,"possibly due to flakiness on trunk: inductor / inductor-cpu-test / test (cpu_inductor_torchbench, 1, 2, mt-l-x86iamx-8-64, unstable) (gh) (#174929) doctr_reco_predictor inductor / inductor-cpu-test / test (cpu_inductor_torchbench, 2, 2, mt-l-x86iamx-8-64, unstable) (gh) (#1749...",https://github.com/pytorch/pytorch/pull/187532,0c1af67ac0c7d13e55e79866db2bc1a831d2b3abb31ef195480ee82535c21cbd references,pr,187468,pr,187469,medium,pr.body,"Stack from ghstack (oldest at bottom): #189092 #189091 #187707 #187532 #187531 #187469 -> #187468 LOAD_ATTR previously relied on the step() fallback for graph breaks, which restarts from the most recent empty-stack checkpoint.",https://github.com/pytorch/pytorch/pull/187468,cb9e50a763cda1d0674b533d38290f15cc97635f18396ebef1dcee1e5230812b references,pr,187468,pr,187531,medium,pr.body,"Stack from ghstack (oldest at bottom): #189092 #189091 #187707 #187532 #187531 #187469 -> #187468 LOAD_ATTR previously relied on the step() fallback for graph breaks, which restarts from the most recent empty-stack che",https://github.com/pytorch/pytorch/pull/187468,b4ce98ef6a8d4ebe21f4836685a4de8bbe35205a6aaaea6945e2f3bb57282688 references,pr,187468,pr,187532,medium,pr.body,"Stack from ghstack (oldest at bottom): #189092 #189091 #187707 #187532 #187531 #187469 -> #187468 LOAD_ATTR previously relied on the step() fallback for graph breaks, which restarts from the most recent empty-s",https://github.com/pytorch/pytorch/pull/187468,144b5f97054be9ce2bc4549439d9b3ee540ad7e84e42aa19c9cc3d33091e5f40 references,pr,187468,pr,187707,medium,pr.body,"Stack from ghstack (oldest at bottom): #189092 #189091 #187707 #187532 #187531 #187469 -> #187468 LOAD_ATTR previously relied on the step() fallback for graph breaks, which restarts from the most recent",https://github.com/pytorch/pytorch/pull/187468,f3c5e46a6a61cb393b361bea25beb4a59dbd5ddcf64160f7b041fcc41c3b66de references,pr,187468,pr,189091,medium,pr.body,"Stack from ghstack (oldest at bottom): #189092 #189091 #187707 #187532 #187531 #187469 -> #187468 LOAD_ATTR previously relied on the step() fallback for graph breaks, which restarts from the mos",https://github.com/pytorch/pytorch/pull/187468,2334ea2a0f4618abf31a55a2c0dac9fb1e9413a97ef02d0b21028d03d335dd12 references,pr,187468,pr,189092,medium,pr.body,"Stack from ghstack (oldest at bottom): #189092 #189091 #187707 #187532 #187531 #187469 -> #187468 LOAD_ATTR previously relied on the step() fallback for graph breaks, which restarts from",https://github.com/pytorch/pytorch/pull/187468,d550fb104eae96cb5031ad9b19701b529bf1adcf9c70da2e6daf484b7a169551 references,pr,187468,issue,174929,medium,pr.comments[0].body,"possibly due to flakiness on trunk: inductor / inductor-cpu-test / test (cpu_inductor_torchbench, 1, 2, mt-l-x86iamx-8-64, unstable) (gh) (#174929) doctr_reco_predictor inductor / inductor-cpu-test / test (cpu_inductor_torchbench, 2, 2, mt-l-x86iamx-8-64, unstable) (gh) (#1749...",https://github.com/pytorch/pytorch/pull/187468,d765d9d986467532807713f8efe3302c895cc061ac63fcf02ac59b3b5496655b references,pr,187469,pr,187468,medium,pr.body,"Stack from ghstack (oldest at bottom): #189092 #189091 #187707 #187532 #187531 -> #187469 #187468 Extract two chunks of UDOV's generic_getattr into hook overrides, matching the hook interface on base VariableTracker: lookup_instance_dict",https://github.com/pytorch/pytorch/pull/187469,41ad5d9b915785e26f44fc85689cba34de1b51519317f79c87be63c11d11d668 references,pr,187469,pr,187531,medium,pr.body,"Stack from ghstack (oldest at bottom): #189092 #189091 #187707 #187532 #187531 -> #187469 #187468 Extract two chunks of UDOV's generic_getattr into hook overrides, matching the hook interface on base VariableTracker: l",https://github.com/pytorch/pytorch/pull/187469,d01116ac2ec938444518dd89970c29fad5b461eed67a525e5eb08576161bf6dc references,pr,187469,pr,187532,medium,pr.body,"Stack from ghstack (oldest at bottom): #189092 #189091 #187707 #187532 #187531 -> #187469 #187468 Extract two chunks of UDOV's generic_getattr into hook overrides, matching the hook interface on base VariableTr",https://github.com/pytorch/pytorch/pull/187469,eac40e0904308ecf271d6e4f3616046b116a39d68a02c9f95ef11f70c5e96816 references,pr,187469,pr,187707,medium,pr.body,"Stack from ghstack (oldest at bottom): #189092 #189091 #187707 #187532 #187531 -> #187469 #187468 Extract two chunks of UDOV's generic_getattr into hook overrides, matching the hook interface on base Va",https://github.com/pytorch/pytorch/pull/187469,f9af337e4fcc89c14e13975f3e759b007652465740b2d3f41008feb189270b19 references,pr,187469,pr,189091,medium,pr.body,"Stack from ghstack (oldest at bottom): #189092 #189091 #187707 #187532 #187531 -> #187469 #187468 Extract two chunks of UDOV's generic_getattr into hook overrides, matching the hook interface on",https://github.com/pytorch/pytorch/pull/187469,d77c6bb2f26e80bd57a0907edb1d85d33422ec78fdd6ca11fd0792bf3eb9c770 references,pr,187469,pr,189092,medium,pr.body,"Stack from ghstack (oldest at bottom): #189092 #189091 #187707 #187532 #187531 -> #187469 #187468 Extract two chunks of UDOV's generic_getattr into hook overrides, matching the hook inte",https://github.com/pytorch/pytorch/pull/187469,6967e6aa2815b527fe2440fa834e158e815c628af356ad60ca60dee468b1c58e references,pr,187469,issue,174929,medium,pr.comments[0].body,"possibly due to flakiness on trunk: inductor / inductor-cpu-test / test (cpu_inductor_torchbench, 1, 2, mt-l-x86iamx-8-64, unstable) (gh) (#174929) RuntimeError: nms_kernel_impl, /tmp/pip-req-build-1iwz2feu/torchvision/csrc/ops/cpu/nms_kernel.cpp:24, dets should have the same...",https://github.com/pytorch/pytorch/pull/187469,b83f25d88728473c02ba2944a1d05f2b9963a558336d94524894ee81b6568331 references,pr,181780,pr,181781,medium,pr.body,Stack from ghstack (oldest at bottom): #181781 #188176 #188175 -> #181780 cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv @jia,https://github.com/pytorch/pytorch/pull/181780,43574d85a0890e03c20dbe8dcad8f7b74863944e37f11f0aef2ea154cfb5ad43 references,pr,181780,pr,188175,medium,pr.body,Stack from ghstack (oldest at bottom): #181781 #188176 #188175 -> #181780 cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv @jiayisunx @kadeng @,https://github.com/pytorch/pytorch/pull/181780,ee052658bf1d5af1a32e3312b0aec805608d5cf76c83360f98cd11f921c3de0d references,pr,181780,pr,188176,medium,pr.body,Stack from ghstack (oldest at bottom): #181781 #188176 #188175 -> #181780 cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv @jiayisunx @,https://github.com/pytorch/pytorch/pull/181780,fcb1301651b246e3f8b9b709a268cf117694b0675b4c446921ef7a282e6785ac references,pr,188175,pr,181780,medium,pr.body,Stack from ghstack (oldest at bottom): #181781 #188176 -> #188175 #181780 Two tests opted out of canonicalize_output_graph_node_order because they were sensitive to placeholder ordering. This PR fixes both so the,https://github.com/pytorch/pytorch/pull/188175,779f3c3d89380ef2873c780f6780a1ae770d069ad1183ed6d1fe5ef3a1286018 references,pr,188175,pr,181781,medium,pr.body,Stack from ghstack (oldest at bottom): #181781 #188176 -> #188175 #181780 Two tests opted out of canonicalize_output_graph_node_order because they were sensitive to placeholder ordering.,https://github.com/pytorch/pytorch/pull/188175,bd3d3f24009f8882cb9c09aae00211867531dd6e98f1b7421ec4d39b5618274d references,pr,188175,pr,188176,medium,pr.body,Stack from ghstack (oldest at bottom): #181781 #188176 -> #188175 #181780 Two tests opted out of canonicalize_output_graph_node_order because they were sensitive to placeholder ordering. This PR,https://github.com/pytorch/pytorch/pull/188175,0fc27a90cbe5cb475e1682ded36c34fe2b06b1b8bb28305b90914107a6d245ea references,pr,188176,pr,181780,medium,pr.body,Stack from ghstack (oldest at bottom): #181781 -> #188176 #188175 #181780 Two tests in test_perf.py opted out of canonicalize_output_graph_node_order unnecessarily. This removes both opt-outs. test_cat_pointwise p,https://github.com/pytorch/pytorch/pull/188176,2fa7a8f80165fb23a452bac6fc4d38b061be834f2e86a6d11ee7c88f89b10da5 references,pr,188176,pr,181781,medium,pr.body,Stack from ghstack (oldest at bottom): #181781 -> #188176 #188175 #181780 Two tests in test_perf.py opted out of canonicalize_output_graph_node_order unnecessarily. This removes both opt,https://github.com/pytorch/pytorch/pull/188176,ce80c9f8af3f2867e70ca237b7c6723951a83e2083328f9c0981ff33a1367e9a references,pr,188176,pr,188175,medium,pr.body,Stack from ghstack (oldest at bottom): #181781 -> #188176 #188175 #181780 Two tests in test_perf.py opted out of canonicalize_output_graph_node_order unnecessarily. This removes both opt-outs. test_cat_poi,https://github.com/pytorch/pytorch/pull/188176,0608caf9730fc85cb4cd25820706483ee61cfbe3852bfae9e2156cfeb590d2d8 references,pr,189190,pr,188112,medium,pr.body,"Stack from ghstack (oldest at bottom): #189316 #189315 #189314 -> #189190 #189188 #188739 #188112 #188470 #188469 #188468 The feed-main matcher now recurses through trailing shape-preserving pointwise nodes after the un-grouping view, so",https://github.com/pytorch/pytorch/pull/189190,5f14c683761cfb2cbc15f5e156d8f8053d53bcb6698127e8f44e44648d0dff28 references,pr,189190,pr,188468,medium,pr.body,"Stack from ghstack (oldest at bottom): #189316 #189315 #189314 -> #189190 #189188 #188739 #188112 #188470 #188469 #188468 The feed-main matcher now recurses through trailing shape-preserving pointwise nodes after the un-grouping view, so quant casts like .to(fl",https://github.com/pytorch/pytorch/pull/189190,bc1bb804f7f41d84b97d67c0b59a602ca3eef643de6e841bbc69ad375e7b9420 references,pr,189190,pr,188469,medium,pr.body,"Stack from ghstack (oldest at bottom): #189316 #189315 #189314 -> #189190 #189188 #188739 #188112 #188470 #188469 #188468 The feed-main matcher now recurses through trailing shape-preserving pointwise nodes after the un-grouping view, so quant casts lik",https://github.com/pytorch/pytorch/pull/189190,343585362e8c8795bf13a5015e369971253ea5a87080e4e268d90ac646118a52 references,pr,189190,pr,188470,medium,pr.body,"Stack from ghstack (oldest at bottom): #189316 #189315 #189314 -> #189190 #189188 #188739 #188112 #188470 #188469 #188468 The feed-main matcher now recurses through trailing shape-preserving pointwise nodes after the un-grouping view, so quant c",https://github.com/pytorch/pytorch/pull/189190,70b1dc6c66433e6fcfd1c7691b8ada38fbc4bacf1db1dad8338f4a7a569f88d4 references,pr,189190,pr,188739,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 #189315 #189314 -> #189190 #189188 #188739 #188112 #188470 #188469 #188468 The feed-main matcher now recurses through trailing shape-preserving pointwise nodes after the un-grouping,https://github.com/pytorch/pytorch/pull/189190,b45f4f9459b3b04fd0629c6549149523fe21e8b35ebe5f11213ac869b6e4b1bb references,pr,189190,pr,189188,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 #189315 #189314 -> #189190 #189188 #188739 #188112 #188470 #188469 #188468 The feed-main matcher now recurses through trailing shape-preserving pointwise nodes after the un-g,https://github.com/pytorch/pytorch/pull/189190,6cf2f51698f75d57472f0cd373214b8e9529f4ec40d173426571400e993f0a4a references,pr,189190,pr,189314,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 #189315 #189314 -> #189190 #189188 #188739 #188112 #188470 #188469 #188468 The feed-main matcher now recurses through trailing shape-preserving pointwise n,https://github.com/pytorch/pytorch/pull/189190,ac8f15683961773110e706c3bb067bbc40dc98d00354c9aed8db0a86dced5a44 references,pr,189190,pr,189315,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 #189315 #189314 -> #189190 #189188 #188739 #188112 #188470 #188469 #188468 The feed-main matcher now recurses through trailing shape-preserving poi,https://github.com/pytorch/pytorch/pull/189190,07426de69236137eec0569235ef4d7715df658aa06b11f0f09739b975f0025db references,pr,189190,pr,189316,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 #189315 #189314 -> #189190 #189188 #188739 #188112 #188470 #188469 #188468 The feed-main matcher now recurses through trailing shape-preser,https://github.com/pytorch/pytorch/pull/189190,62969931c59c3aad1ea863ba29d130db4b67dd43c9e515fe8cc7a9ed79a78e06 references,pr,189188,pr,188112,medium,pr.body,"Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 -> #189188 #188739 #188112 #188470 #188469 #188468 Add a fourth captured-epilogue-arg kind ""scalar"" beside tile/row/col for [1, 1] read-only tensors. This unlocks nat",https://github.com/pytorch/pytorch/pull/189188,fecf9e1ba94c661af1ce41d8f9470d9f9af668ada6e5a654b2cda052434bf829 references,pr,189188,pr,188468,medium,pr.body,"Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 -> #189188 #188739 #188112 #188470 #188469 #188468 Add a fourth captured-epilogue-arg kind ""scalar"" beside tile/row/col for [1, 1] read-only tensors. This unlocks native float8 tensorwise qu",https://github.com/pytorch/pytorch/pull/189188,9253e52ddc1f3f4be34c277cdbb459fea4b2f5ecc733f5bad2d366cf793d61bc references,pr,189188,pr,188469,medium,pr.body,"Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 -> #189188 #188739 #188112 #188470 #188469 #188468 Add a fourth captured-epilogue-arg kind ""scalar"" beside tile/row/col for [1, 1] read-only tensors. This unlocks native float8 tenso",https://github.com/pytorch/pytorch/pull/189188,1bdc17173c56ecbf1926cadf552cc6b34e65738e786fae59a02d76d80a3aeea5 references,pr,189188,pr,188470,medium,pr.body,"Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 -> #189188 #188739 #188112 #188470 #188469 #188468 Add a fourth captured-epilogue-arg kind ""scalar"" beside tile/row/col for [1, 1] read-only tensors. This unlocks native floa",https://github.com/pytorch/pytorch/pull/189188,ea1fe0ef347c1cea50fe1486fc96d18b27991c02d71400184141555ef0404196 references,pr,189188,pr,188739,medium,pr.body,"Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 -> #189188 #188739 #188112 #188470 #188469 #188468 Add a fourth captured-epilogue-arg kind ""scalar"" beside tile/row/col for [1, 1] read-only tensors. This unl",https://github.com/pytorch/pytorch/pull/189188,038ab7e94218a6ced61a461fcabc255c44c2710db51f06c9c43512e8a950d569 references,pr,189188,pr,189190,medium,pr.body,"Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 -> #189188 #188739 #188112 #188470 #188469 #188468 Add a fourth captured-epilogue-arg kind ""scalar"" beside tile/row/col for [1, 1] read-onl",https://github.com/pytorch/pytorch/pull/189188,0dc09c0656b2b9af0082d55e050cf1f4e58545b1c68b8a15eb92c9eed5af4581 references,pr,189188,pr,189314,medium,pr.body,"Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 -> #189188 #188739 #188112 #188470 #188469 #188468 Add a fourth captured-epilogue-arg kind ""scalar"" beside tile/row/col for [1, 1]",https://github.com/pytorch/pytorch/pull/189188,5e72902e53a018bb8e70e3db62ab456e7abf89702e9206ab74129c9125cd61f0 references,pr,189188,pr,189315,medium,pr.body,"Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 -> #189188 #188739 #188112 #188470 #188469 #188468 Add a fourth captured-epilogue-arg kind ""scalar"" beside tile/row/col for",https://github.com/pytorch/pytorch/pull/189188,d3264215e0ecead6ab09323ce58c375e92ce1ee3ae6297e7b155c58c8d2a584b references,pr,189188,pr,189316,medium,pr.body,"Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 -> #189188 #188739 #188112 #188470 #188469 #188468 Add a fourth captured-epilogue-arg kind ""scalar"" beside tile/row",https://github.com/pytorch/pytorch/pull/189188,d7e8b596ad50d0c6adb2b249aff4bafc00b16f2043aa43503e52d8a7899fba31 review guidance,pr,189188,pr,188112,high,pr.reviews[0].body,"💡 Codex Review Here are some automated review suggestions for this pull request. Reviewed commit: 70f1a957fb ℹ️ About Codex in GitHub Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you Open a pull request for review Mark a draft as ready Comment ""@code...",https://github.com/pytorch/pytorch/pull/189188,d6604fc41794dfc3d7c44d7fecdbcccc63669894e894de81af85149bfa24a6a2 review guidance,pr,189188,pr,188468,high,pr.reviews[0].body,"💡 Codex Review Here are some automated review suggestions for this pull request. Reviewed commit: 70f1a957fb ℹ️ About Codex in GitHub Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you Open a pull request for review Mark a draft as ready Comment ""@code...",https://github.com/pytorch/pytorch/pull/189188,b6d2ed7f9c85017cd046ceb8af36b431ac012faa92c132a95faaddfcbe5f9d1e review guidance,pr,189188,pr,188469,high,pr.reviews[0].body,"💡 Codex Review Here are some automated review suggestions for this pull request. Reviewed commit: 70f1a957fb ℹ️ About Codex in GitHub Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you Open a pull request for review Mark a draft as ready Comment ""@code...",https://github.com/pytorch/pytorch/pull/189188,8c3f4bb1552385e1bb694c2a665568a3981e6b81ca284ec2d19f39325030429e review guidance,pr,189188,pr,188470,high,pr.reviews[0].body,"💡 Codex Review Here are some automated review suggestions for this pull request. Reviewed commit: 70f1a957fb ℹ️ About Codex in GitHub Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you Open a pull request for review Mark a draft as ready Comment ""@code...",https://github.com/pytorch/pytorch/pull/189188,76bddaf56ff4a42ed6f94a489d066d37b652f88d492f02d7bd1b0a7a75b85bd8 review guidance,pr,189188,pr,188739,high,pr.reviews[0].body,"💡 Codex Review Here are some automated review suggestions for this pull request. Reviewed commit: 70f1a957fb ℹ️ About Codex in GitHub Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you Open a pull request for review Mark a draft as ready Comment ""@code...",https://github.com/pytorch/pytorch/pull/189188,cbb2381257da650e626c00f9c6fbc7f2690fc2d68241e821087ed511a7151151 review guidance,pr,189188,pr,189188,high,pr.reviews[0].body,"💡 Codex Review Here are some automated review suggestions for this pull request. Reviewed commit: 70f1a957fb ℹ️ About Codex in GitHub Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you Open a pull request for review Mark a draft as ready Comment ""@code...",https://github.com/pytorch/pytorch/pull/189188,c9002004e439cd5d07ccb01d30e6063d1e7a70a3af5751414a875cf04b3ab8f6 review guidance,pr,189188,pr,189190,high,pr.reviews[0].body,"💡 Codex Review Here are some automated review suggestions for this pull request. Reviewed commit: 70f1a957fb ℹ️ About Codex in GitHub Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you Open a pull request for review Mark a draft as ready Comment ""@code...",https://github.com/pytorch/pytorch/pull/189188,3cd63eb727aea6554fdae97c592704ac2afc572f58b7b9c4bd4ba7a64e84efe4 review guidance,pr,189188,pr,189314,high,pr.reviews[0].body,"💡 Codex Review Here are some automated review suggestions for this pull request. Reviewed commit: 70f1a957fb ℹ️ About Codex in GitHub Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you Open a pull request for review Mark a draft as ready Comment ""@code...",https://github.com/pytorch/pytorch/pull/189188,d5efa9f328241a146952c9d78f75cf04a3ca041d981149fc9533d57d94abe9d4 review guidance,pr,189188,pr,189315,high,pr.reviews[0].body,"💡 Codex Review Here are some automated review suggestions for this pull request. Reviewed commit: 70f1a957fb ℹ️ About Codex in GitHub Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you Open a pull request for review Mark a draft as ready Comment ""@code...",https://github.com/pytorch/pytorch/pull/189188,59f974e19c731b46fb368dbb649c6360298d2094d0094b010aeff8a92a4f7f34 review guidance,pr,189188,pr,189316,high,pr.reviews[0].body,"💡 Codex Review Here are some automated review suggestions for this pull request. Reviewed commit: 70f1a957fb ℹ️ About Codex in GitHub Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you Open a pull request for review Mark a draft as ready Comment ""@code...",https://github.com/pytorch/pytorch/pull/189188,a1f2f77969c519db7f53f4e8a5bfd1e277f309e802fc30fdbba63f793fd254cc references,pr,188739,pr,188112,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 -> #188739 #188112 #188470 #188469 #188468,https://github.com/pytorch/pytorch/pull/188739,68889581c4c2231b9eb926a0cb980de18d4308b3731a487a7f937f35c880b14b references,pr,188739,pr,188468,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 -> #188739 #188112 #188470 #188469 #188468,https://github.com/pytorch/pytorch/pull/188739,d84852bad30be4d51ff592806a40ae34c5b01484ab528999339df3d8cc4af764 references,pr,188739,pr,188469,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 -> #188739 #188112 #188470 #188469 #188468,https://github.com/pytorch/pytorch/pull/188739,b64119180513fa97744d39bdf781cf52ed71713290f8f1ffce0650e7e575db4c references,pr,188739,pr,188470,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 -> #188739 #188112 #188470 #188469 #188468,https://github.com/pytorch/pytorch/pull/188739,70d8613d080277468e8fab9dc2a6f08feee4feb4809db934326f77e041de10c2 references,pr,188739,pr,189188,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 -> #188739 #188112 #188470 #188469 #188468,https://github.com/pytorch/pytorch/pull/188739,b65760a8c23fd445bc0b8e061ba0fc8ed1d2439c7c2052f650d8b70415b89a72 references,pr,188739,pr,189190,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 -> #188739 #188112 #188470 #188469 #188468,https://github.com/pytorch/pytorch/pull/188739,4f9c8abb92e011339d04ad5d69d1f22efadfa22048fb90f9b055898ec2887621 references,pr,188739,pr,189314,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 -> #188739 #188112 #188470 #188469 #188468,https://github.com/pytorch/pytorch/pull/188739,6cba1ac2652a9288315331b523953697ea48d7941959f701e58973ac52b74354 references,pr,188739,pr,189315,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 -> #188739 #188112 #188470 #188469 #188468,https://github.com/pytorch/pytorch/pull/188739,cb3e46f3f63a81017065c59d722536da4d40933125b34d51d039fcd87a107ded references,pr,188739,pr,189316,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 -> #188739 #188112 #188470 #188469 #188468,https://github.com/pytorch/pytorch/pull/188739,ebd59cedbb6253d5b21ac9e1eda0282bf162d0c89970fef911628529c9bf0db5 references,pr,188112,pr,188468,medium,pr.body,"Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 #188739 -> #188112 #188470 #188469 #188468 Let dense aten.mm FlexGEMM epilogues feed a grouped local reduction back into the main output expression (and optionally store it), e.g. ac",https://github.com/pytorch/pytorch/pull/188112,208917867f2893bdf0e724e56d82baf03f161ea57dbd6b162f896d9b06bbc9c5 references,pr,188112,pr,188469,medium,pr.body,"Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 #188739 -> #188112 #188470 #188469 #188468 Let dense aten.mm FlexGEMM epilogues feed a grouped local reduction back into the main output expression (and optionally store it),",https://github.com/pytorch/pytorch/pull/188112,56e8b91f73010656050e478e61ce5df9b1413104aba8151652d5f44b21bf99e0 references,pr,188112,pr,188470,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 #188739 -> #188112 #188470 #188469 #188468 Let dense aten.mm FlexGEMM epilogues feed a grouped local reduction back into the main output expression (and optionally st,https://github.com/pytorch/pytorch/pull/188112,3c39213ac0604ba942c2ecfb510d44859aa607fdfe6e0c93ddf1c73f9c40a900 references,pr,188112,pr,188739,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 #188739 -> #188112 #188470 #188469 #188468 Let dense aten.mm FlexGEMM epilogues feed a grouped local reduction back into the main output expression,https://github.com/pytorch/pytorch/pull/188112,9df18252f7d8bb1757f1995df77d912eb43f5b245ac4fb2bb9778d2e4020c1d5 references,pr,188112,pr,189188,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 #188739 -> #188112 #188470 #188469 #188468 Let dense aten.mm FlexGEMM epilogues feed a grouped local reduction back into the main output ex,https://github.com/pytorch/pytorch/pull/188112,f09649bac76ec12c5c1aecf2af669d994a5eb0e0a4416a5346f5b638f1dea4ba references,pr,188112,pr,189190,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 #188739 -> #188112 #188470 #188469 #188468 Let dense aten.mm FlexGEMM epilogues feed a grouped local reduction back into the main o,https://github.com/pytorch/pytorch/pull/188112,4c007627a3532c940f7c9e6f09604233e19e1f8a362c031fd58b8e5347822e9f references,pr,188112,pr,189314,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 #188739 -> #188112 #188470 #188469 #188468 Let dense aten.mm FlexGEMM epilogues feed a grouped local reduction back into th,https://github.com/pytorch/pytorch/pull/188112,116931a5ff57b54b5bdfce766d8145275f2c106fb319faa9cc5b40fe76fc03b0 references,pr,188112,pr,189315,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 #188739 -> #188112 #188470 #188469 #188468 Let dense aten.mm FlexGEMM epilogues feed a grouped local reduction back,https://github.com/pytorch/pytorch/pull/188112,1ca2f2c00a6a28bd2f22af41c9ba00e1cddea637aa70355e07114b16e759a5b9 references,pr,188112,pr,189316,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 #188739 -> #188112 #188470 #188469 #188468 Let dense aten.mm FlexGEMM epilogues feed a grouped local reduct,https://github.com/pytorch/pytorch/pull/188112,ed7848d197bce6cbf322f28b4e6e471262274c2c63fe8aa42ace68773c78c536 references,pr,188470,pr,188112,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 #188739 #188112 -> #188470 #188469 #188468 Extend compressed local-reduce aux outputs beyond one 32-lane TensorSSA fragment: axis-1 groups above the fragme,https://github.com/pytorch/pytorch/pull/188470,298dbca38caa1f8aeb243550e37e4ed6302cf09510860ac811b2999e87c03459 references,pr,188470,pr,188468,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 #188739 #188112 -> #188470 #188469 #188468 Extend compressed local-reduce aux outputs beyond one 32-lane TensorSSA fragment: axis-1 groups above the fragment width (CTA-subtile group,https://github.com/pytorch/pytorch/pull/188470,5af4848af5ae95ea41c516ec831d9cf7443edf2e94fd91c8759ae64bdd9f088a references,pr,188470,pr,188469,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 #188739 #188112 -> #188470 #188469 #188468 Extend compressed local-reduce aux outputs beyond one 32-lane TensorSSA fragment: axis-1 groups above the fragment width (CTA-subti,https://github.com/pytorch/pytorch/pull/188470,50d0eaa2c6069268e1d7938537a3df65f092d98d80528fc7269d4e943fae5d64 references,pr,188470,pr,188739,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 #188739 #188112 -> #188470 #188469 #188468 Extend compressed local-reduce aux outputs beyond one 32-lane TensorSSA fragment: axis-1 groups above th,https://github.com/pytorch/pytorch/pull/188470,52746094686e3c0df7832af8ac055f5c0023db43ac933de31c9f7249d9af5396 references,pr,188470,pr,189188,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 #188739 #188112 -> #188470 #188469 #188468 Extend compressed local-reduce aux outputs beyond one 32-lane TensorSSA fragment: axis-1 groups,https://github.com/pytorch/pytorch/pull/188470,bb5f24c773a2253a748366dcd59ccbbd454680c3272add80b0ef189b00a42c80 references,pr,188470,pr,189190,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 #188739 #188112 -> #188470 #188469 #188468 Extend compressed local-reduce aux outputs beyond one 32-lane TensorSSA fragment: axis-1,https://github.com/pytorch/pytorch/pull/188470,3aa35f6fc9661c9cbd44d81d250b96ef3d44daafeeee303478bf8374e7d97243 references,pr,188470,pr,189314,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 #188739 #188112 -> #188470 #188469 #188468 Extend compressed local-reduce aux outputs beyond one 32-lane TensorSSA fragment,https://github.com/pytorch/pytorch/pull/188470,363bd1a5e08fb33cc764ae6e4ccdf7ae8b179865503d0bd41e4f54ede6faae0d references,pr,188470,pr,189315,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 #188739 #188112 -> #188470 #188469 #188468 Extend compressed local-reduce aux outputs beyond one 32-lane TensorSSA,https://github.com/pytorch/pytorch/pull/188470,9f332057e43fcb2e16ef2087a0f4ef7a5c5f6858537e265cfdbe1afe81844347 references,pr,188470,pr,189316,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 #188739 #188112 -> #188470 #188469 #188468 Extend compressed local-reduce aux outputs beyond one 32-lane Te,https://github.com/pytorch/pytorch/pull/188470,3f14b9ecf5221f5fec6c67cc03bf9f473b504352bf2091ad407684e2269d6ba9 references,pr,188469,pr,188112,medium,pr.body,"Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 #188739 #188112 #188470 -> #188469 #188468 Support dense non-batched aten.mm FlexGEMM epilogues that return one compressed local-reduce aux output, e.g. ac",https://github.com/pytorch/pytorch/pull/188469,b5819e45323880d2cd79c58d9b073450d18e40fc427c9b45623dee9760d2d83a references,pr,188469,pr,188468,medium,pr.body,"Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 #188739 #188112 #188470 -> #188469 #188468 Support dense non-batched aten.mm FlexGEMM epilogues that return one compressed local-reduce aux output, e.g. acc.relu(), acc.view(M, -1, g",https://github.com/pytorch/pytorch/pull/188469,23ab9de626c7e6fd8a6ad3e1ed6b42f2514eae40038a1609145285e3ff9eec6d references,pr,188469,pr,188470,medium,pr.body,"Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 #188739 #188112 #188470 -> #188469 #188468 Support dense non-batched aten.mm FlexGEMM epilogues that return one compressed local-reduce aux output, e.g. acc.relu()",https://github.com/pytorch/pytorch/pull/188469,a279bab04b82aaa4329b263b80483bcc7383b5e0878619d6c44c44153f0681c3 references,pr,188469,pr,188739,medium,pr.body,"Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 #188739 #188112 #188470 -> #188469 #188468 Support dense non-batched aten.mm FlexGEMM epilogues that return one compressed local-reduce aux output,",https://github.com/pytorch/pytorch/pull/188469,a480002f4d73a4b36880d31136d1e443b2745eb904cad040c18f0abce891f4f0 references,pr,188469,pr,189188,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 #188739 #188112 #188470 -> #188469 #188468 Support dense non-batched aten.mm FlexGEMM epilogues that return one compressed local-reduce aux,https://github.com/pytorch/pytorch/pull/188469,9e58204eec5848f6dce2ea235a60deb2a334758fd11311cbf568f0ee7ccf6ba3 references,pr,188469,pr,189190,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 #188739 #188112 #188470 -> #188469 #188468 Support dense non-batched aten.mm FlexGEMM epilogues that return one compressed local-re,https://github.com/pytorch/pytorch/pull/188469,bd9dbbd7df4c4dece2d58feb377fe38883d1346de5a405cc0aa16b8439d9eb26 references,pr,188469,pr,189314,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 #188739 #188112 #188470 -> #188469 #188468 Support dense non-batched aten.mm FlexGEMM epilogues that return one compressed,https://github.com/pytorch/pytorch/pull/188469,f899c27f19a535ccb05a6731f266acbdc0283d3439c58fe53990cff0665104f4 references,pr,188469,pr,189315,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 #188739 #188112 #188470 -> #188469 #188468 Support dense non-batched aten.mm FlexGEMM epilogues that return one com,https://github.com/pytorch/pytorch/pull/188469,b4fbb6e1ebc7066b796f6262c20a009ee754e8a42b83ddfa75aaaf67e1b8b979 references,pr,188469,pr,189316,medium,pr.body,Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 #188739 #188112 #188470 -> #188469 #188468 Support dense non-batched aten.mm FlexGEMM epilogues that return,https://github.com/pytorch/pytorch/pull/188469,858c6d0a3e97b26b71be694a2f2a33f36274b795ffa963fb56a230c01f8b8638 references,pr,188468,pr,188112,medium,pr.body,"Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 #188739 #188112 #188470 #188469 -> #188468 Standalone analysis vocabulary for FlexGEMM local reductions, consumed by the compressed local-reduce aux-output",https://github.com/pytorch/pytorch/pull/188468,853bc3d3c535c31bb351942aa5fa2e28ca6a5607cbcdb0891d3fc90327ea20e3 references,pr,188468,pr,188469,medium,pr.body,"Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 #188739 #188112 #188470 #188469 -> #188468 Standalone analysis vocabulary for FlexGEMM local reductions, consumed by the compressed local-reduce aux-output support stacked",https://github.com/pytorch/pytorch/pull/188468,9b7f0a33265ce3b89aadc2baaa55863979ace7350793fbdda53bf57e9cc3e365 references,pr,188468,pr,188470,medium,pr.body,"Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 #188739 #188112 #188470 #188469 -> #188468 Standalone analysis vocabulary for FlexGEMM local reductions, consumed by the compressed local-reduce aux-output support",https://github.com/pytorch/pytorch/pull/188468,451c3a4ab349048a4c1d3ac3cdba042cf9dd76c57e85a6e1541ccdebe60588db references,pr,188468,pr,188739,medium,pr.body,"Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 #188739 #188112 #188470 #188469 -> #188468 Standalone analysis vocabulary for FlexGEMM local reductions, consumed by the compressed local-reduce au",https://github.com/pytorch/pytorch/pull/188468,f4417c2ca2dca65e738f3aefa07763035566c858c232c801c9bc1f2e900b3645 references,pr,188468,pr,189188,medium,pr.body,"Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 #188739 #188112 #188470 #188469 -> #188468 Standalone analysis vocabulary for FlexGEMM local reductions, consumed by the compressed local-r",https://github.com/pytorch/pytorch/pull/188468,bc926372ee8d3164d1a7500fbb87ef5f88231fe27a35440971d17e7192a0e296 references,pr,188468,pr,189190,medium,pr.body,"Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 #188739 #188112 #188470 #188469 -> #188468 Standalone analysis vocabulary for FlexGEMM local reductions, consumed by the compressed",https://github.com/pytorch/pytorch/pull/188468,d314e026b9ae0e49ca0b9225db2d4a8010d861d5314708cc27ea6b2d4641ef46 references,pr,188468,pr,189314,medium,pr.body,"Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 #188739 #188112 #188470 #188469 -> #188468 Standalone analysis vocabulary for FlexGEMM local reductions, consumed by the co",https://github.com/pytorch/pytorch/pull/188468,cccc6422faad0fa0175401efc000f0393cf849de7dba39d9001d044f1413de32 references,pr,188468,pr,189315,medium,pr.body,"Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 #188739 #188112 #188470 #188469 -> #188468 Standalone analysis vocabulary for FlexGEMM local reductions, consumed b",https://github.com/pytorch/pytorch/pull/188468,4f05c92ea4268142e9a5d2e361c17041f1824c80ec9d1ca2c230136ac90949f8 references,pr,188468,pr,189316,medium,pr.body,"Stack from ghstack (oldest at bottom): #189316 #189315 #189314 #189190 #189188 #188739 #188112 #188470 #188469 -> #188468 Standalone analysis vocabulary for FlexGEMM local reductions, co",https://github.com/pytorch/pytorch/pull/188468,b4191e0c348bafeec5254e322b83dbedba57b09f465d4198bb7ab027cfcb2345 review guidance,pr,188468,pr,188112,high,pr.reviews[0].body,"💡 Codex Review Here are some automated review suggestions for this pull request. Reviewed commit: 6c808f7e38 ℹ️ About Codex in GitHub Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you Open a pull request for review Mark a draft as ready Comment ""@code...",https://github.com/pytorch/pytorch/pull/188468,bc8c982204e715e294e6664ecfd973a6599cf5bc9b19771da389ef25f4458543 review guidance,pr,188468,pr,188468,high,pr.reviews[0].body,"💡 Codex Review Here are some automated review suggestions for this pull request. Reviewed commit: 6c808f7e38 ℹ️ About Codex in GitHub Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you Open a pull request for review Mark a draft as ready Comment ""@code...",https://github.com/pytorch/pytorch/pull/188468,149c079008d2695d22a2c938571dab61e3768aa7eabaabc805bbbe36f147564c review guidance,pr,188468,pr,188469,high,pr.reviews[0].body,"💡 Codex Review Here are some automated review suggestions for this pull request. Reviewed commit: 6c808f7e38 ℹ️ About Codex in GitHub Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you Open a pull request for review Mark a draft as ready Comment ""@code...",https://github.com/pytorch/pytorch/pull/188468,256d7a620328aafd643d11c7c6665482534c82274b2fe1557c6ca991d1703379 review guidance,pr,188468,pr,188470,high,pr.reviews[0].body,"💡 Codex Review Here are some automated review suggestions for this pull request. Reviewed commit: 6c808f7e38 ℹ️ About Codex in GitHub Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you Open a pull request for review Mark a draft as ready Comment ""@code...",https://github.com/pytorch/pytorch/pull/188468,6a02d49f8a72dbc40419837a7dc5ca037b6873b965ebced01061a9404a295cb7 review guidance,pr,188468,pr,188739,high,pr.reviews[0].body,"💡 Codex Review Here are some automated review suggestions for this pull request. Reviewed commit: 6c808f7e38 ℹ️ About Codex in GitHub Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you Open a pull request for review Mark a draft as ready Comment ""@code...",https://github.com/pytorch/pytorch/pull/188468,b59fe6bb622cbb24794253517d951092540429134cc037b6019e985a7a139069 review guidance,pr,188468,pr,189188,high,pr.reviews[0].body,"💡 Codex Review Here are some automated review suggestions for this pull request. Reviewed commit: 6c808f7e38 ℹ️ About Codex in GitHub Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you Open a pull request for review Mark a draft as ready Comment ""@code...",https://github.com/pytorch/pytorch/pull/188468,28cd6d1bf911bfdb73bc21cbacefa556ca0d89f06e12b5feb3a64e91bab2fb87 review guidance,pr,188468,pr,189190,high,pr.reviews[0].body,"💡 Codex Review Here are some automated review suggestions for this pull request. Reviewed commit: 6c808f7e38 ℹ️ About Codex in GitHub Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you Open a pull request for review Mark a draft as ready Comment ""@code...",https://github.com/pytorch/pytorch/pull/188468,31e17f150ec0baa66217cdb1e65c8ee4009fae20c21dfff9594a3612a8510f82 review guidance,pr,188468,pr,189314,high,pr.reviews[0].body,"💡 Codex Review Here are some automated review suggestions for this pull request. Reviewed commit: 6c808f7e38 ℹ️ About Codex in GitHub Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you Open a pull request for review Mark a draft as ready Comment ""@code...",https://github.com/pytorch/pytorch/pull/188468,043b01f0a7ace7b33de005cc097971eb7545008e3473a08816ff4f925691da69 review guidance,pr,188468,pr,189315,high,pr.reviews[0].body,"💡 Codex Review Here are some automated review suggestions for this pull request. Reviewed commit: 6c808f7e38 ℹ️ About Codex in GitHub Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you Open a pull request for review Mark a draft as ready Comment ""@code...",https://github.com/pytorch/pytorch/pull/188468,7afcf748325eaacc8f4aed1fda2f039398240857a08808b520863324f3f94b90 review guidance,pr,188468,pr,189316,high,pr.reviews[0].body,"💡 Codex Review Here are some automated review suggestions for this pull request. Reviewed commit: 6c808f7e38 ℹ️ About Codex in GitHub Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you Open a pull request for review Mark a draft as ready Comment ""@code...",https://github.com/pytorch/pytorch/pull/188468,a89f06d3ce7755cea9839794b19bbad5336de810f6f36548ad5d98f25f8cce90 closes,pr,186358,issue,185484,high,pr.closingIssuesReferences,pr #186358 declares a closing reference to issue #185484.,https://github.com/pytorch/pytorch/pull/186358,9ccbe29a4567df0872eae0473db7c0c08fcd6f2970b3bbe136f9be5efbfd3264 closes,pr,186358,issue,185484,high,pr.body,Fixes #185484 The softshrink decomposition and CPU eager kernel did not cast the scalar lambd to the input tensor's dtype before arithmetic. For reduced,https://github.com/pytorch/pytorch/pull/186358,4ae89df12dc95d8fefe10ffc2e1455eb6c10c7dc00363e50316d3b52743486f3 references,pr,187458,pr,187457,medium,pr.body,"d leaf (cumulative-leaf fork on main), so the GitHub diff shows the whole softmax stack + this one. To see only what this PR adds on top of #187457, use this fork compare (renders as a normal diff of just this delta): anagnorisis2peripeteia/pytorch@mps-logsoftmax-pr2...mps-cro...",https://github.com/pytorch/pytorch/pull/187458,54aaaa017d4e36a1046fb6b9f51c8b608c7938f39759137df0ed25090a68386d references,pr,189295,pr,189294,medium,pr.comments[1].body,", which breaks scipy 1.10.1 (ValueError: numpy.dtype size changed, may indicate binary incompatibility). That's tracked/fixed separately in #189294 and will clear once that lands or this is rebased past it. The target job for this PR, linux-jammy-cuda13_0-py3_10-gcc11-sm90-FA3...",https://github.com/pytorch/pytorch/pull/189295,4e110e0dd80dd2adcdc05ef68f6fdbd3800c3174d88e2c08b0f40a55b44da5c7 references,pr,186965,issue,186449,medium,pr.body,ompilation on value change. .set() and .reset() graph-break with a SUPPORTABLE hint. This is a draft PR for the Phase - 1 implementation of #186449 cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv @jiayisunx @kadeng @c...,https://github.com/pytorch/pytorch/pull/186965,16eec6bb123f040269f98fd40750c7d00746f29f7919ff5891afe6b76273ff19 review guidance,pr,186965,issue,186449,high,pr.reviews[0].body,"Overall design looks good, just some comments on the implementation",https://github.com/pytorch/pytorch/pull/186965,da86102675d8461ed9ab388b16334ac733becda11a28566d7f728e7b89511cd0 references,pr,180020,issue,167885,medium,pr.body,#167885 Implements a context manager that warns when any work is enqueued on the NULL stream. This can help detect unintended use of the NULL strea,https://github.com/pytorch/pytorch/pull/180020,277458ed12c2cd6ef26b39c9f094c765fc1abbad694bb9e5db5075bfed8a9c87 closes,pr,189129,issue,188890,high,pr.closingIssuesReferences,pr #189129 declares a closing reference to issue #188890.,https://github.com/pytorch/pytorch/pull/189129,8f0b0cc1919e4888266bc08f59206507d149ef988833062b5a3513bb9714240d closes,pr,189129,issue,188890,high,pr.body,"Fixes #188890 Description When torch.use_deterministic_algorithms(True) is enabled, the AOTAutograd decomposition of repeat_interleave combined with slic",https://github.com/pytorch/pytorch/pull/189129,c8b909ba76c67a70ebd78af5c61a3485e334637ab0e6f47d8a761b64debb4df5 review guidance,pr,186790,pr,186790,high,pr.reviews[0].body,merge conflicts,https://github.com/pytorch/pytorch/pull/186790,88cfb9e97c5f92947223750a385f93f595e95641ea13543fb505c255dc38918c review guidance,pr,186790,pr,186790,high,pr.reviews[1].body,.,https://github.com/pytorch/pytorch/pull/186790,1f438f6e3dabc82f981fffb9f088d2e28168c173b2f763aa13376c0ad87a531b review guidance,pr,185617,pr,185617,high,pr.reviews[0].body,address automated feedback from claude,https://github.com/pytorch/pytorch/pull/185617,4fdf81ba74186dd97b3135e8344275c3504954a1e4bd3b6536c93c59942a9dca references,pr,189256,pr,181720,medium,pr.body,Stack from ghstack (oldest at bottom): #181720 -> #189256 copy_from_mps_ and copy_to_mps_stride_contig both wrapped the host side of a CPU<->MPS copy identically: page-align the storage,https://github.com/pytorch/pytorch/pull/189256,b57a061da0d5833bc7244099c4ac770b82bbc68fb328a30a0790f61639a3fe13 references,pr,189230,pr,185057,medium,pr.body,"This is a WIP approach to supporting multiple MemPools in cudagraph_trees and thus mode=""reduce-overhead"". It's built on top of #185057. Putting this up in draft to get some eyes on it, but it should get rebased and will be easier to review once #185057, #188755, #189123, #1",https://github.com/pytorch/pytorch/pull/189230,66cd246106778d31478ed18d88975621fec7f1748c8947069e204cabc003c802 references,pr,189230,pr,188755,medium,pr.body,"lt on top of #185057. Putting this up in draft to get some eyes on it, but it should get rebased and will be easier to review once #185057, #188755, #189123, #189173, and #189175 go in. cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blz...",https://github.com/pytorch/pytorch/pull/189230,f4df01950fe2eb10bc30360e62547359f18c98488fd680d84cf2d54c8bd74233 references,pr,189230,pr,189123,medium,pr.body,"of #185057. Putting this up in draft to get some eyes on it, but it should get rebased and will be easier to review once #185057, #188755, #189123, #189173, and #189175 go in. cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenz...",https://github.com/pytorch/pytorch/pull/189230,71f5b34f5ad62e54ae3461c9db1816a080e70f5a8183843a817c9fa6d99792c0 references,pr,189230,pr,189173,medium,pr.body,"57. Putting this up in draft to get some eyes on it, but it should get rebased and will be easier to review once #185057, #188755, #189123, #189173, and #189175 go in. cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv @...",https://github.com/pytorch/pytorch/pull/189230,3a67ae1540624063191948a0a6398129ba866c0cc6b82bcb0cf09c896091f6be references,pr,189230,pr,189175,medium,pr.body,"his up in draft to get some eyes on it, but it should get rebased and will be easier to review once #185057, #188755, #189123, #189173, and #189175 go in. cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv @jiayisunx @ip...",https://github.com/pytorch/pytorch/pull/189230,80503d351e01113f46be47e03c960ba3d3357167a822b1c1c135bc1fff49d47d closes,pr,186082,issue,140960,high,pr.body,lso made explicit after the size mismatch return. This is behavior-preserving and keeps clang-tidy happy when this header is touched. Fixes #140960 Generated by my agent Test Plan: ninja -C build torch_python ninja -C build c10_SymInt_test build/bin/c10_SymInt_test --gtest_fil...,https://github.com/pytorch/pytorch/pull/186082,d69e2a8c5e2a316a48f01d64c8cbaa59cdf4834f9c7094e6b952f232b6612d0a closes,pr,189294,issue,189034,high,pr.body,"o the numpy/scipy pair is consistent. Mirrors the compatible numpy==2.0.2 + scipy==1.13.1 set already used in test.sh's numpy_2 path. Fixes #189034. Test Plan: CI: the ""Limited CI on H100 / ...sm90 / test (smoke)"" job should go green; test_foreach no longer crashes importing s...",https://github.com/pytorch/pytorch/pull/189294,4e3c8c84bfa4ef9e09101c3d31a3663b75fe52f6f5b54c001c86aedba3f26277 competes with,pr,189294,issue,189034,medium,pr.comments[1].body,"rted crash. test_foreach now imports cleanly and its TestForeachMM suite runs (18 passed) instead of dying on the numpy/scipy ABI error, so #189034 is resolved. The smoke job is still red, but on a separate, pre-existing issue this unmasked: test_foreach_mm_cuda_mixed_muon_bfl...",https://github.com/pytorch/pytorch/pull/189294,04258b8230af8f388411d751f77faa11ab306e74830fc5acd74d9e579d9439b9 closes,pr,189288,issue,189271,high,pr.body,"Stack from ghstack (oldest at bottom): -> #189288 Fixes #189271. Root cause: GraphLowering.propagate_mutation emits copy_(old_arg, new_arg) to reflect an in-place op's mutation back onto its original arg",https://github.com/pytorch/pytorch/pull/189288,404fb4fa46e34150937b03423f294737b562ae4781d4d0939406db83192f2bf5 review guidance,pr,189187,pr,189187,high,pr.reviews[0].body,"💡 Codex Review Here are some automated review suggestions for this pull request. Reviewed commit: b6975c3b4c ℹ️ About Codex in GitHub Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you Open a pull request for review Mark a draft as ready Comment ""@code...",https://github.com/pytorch/pytorch/pull/189187,d0d99b8c735e69c05fc62029065663c2d70687b0a90034ae72d34a51f660ab46 references,pr,189021,pr,185424,medium,pr.body,Stack from ghstack (oldest at bottom): #185424 #189053 #189052 #189022 #189051 -> #189021 Constructing a tuple subclass with no arguments -- e.g. class MyTuple(tuple): pass; MyTuple() --,https://github.com/pytorch/pytorch/pull/189021,a7b2ddd595f285cf39b687cac055bd3aa3ad4d5bc2f7fd5e81feb48b7e4f759d references,pr,189021,pr,189022,medium,pr.body,Stack from ghstack (oldest at bottom): #185424 #189053 #189052 #189022 #189051 -> #189021 Constructing a tuple subclass with no arguments -- e.g. class MyTuple(tuple): pass; MyTuple() -- crashed under Dynamo wi,https://github.com/pytorch/pytorch/pull/189021,7ab252246bfe1739f9cac8b9d5beb39219de8166456727604094e4207708afae references,pr,189021,pr,189051,medium,pr.body,Stack from ghstack (oldest at bottom): #185424 #189053 #189052 #189022 #189051 -> #189021 Constructing a tuple subclass with no arguments -- e.g. class MyTuple(tuple): pass; MyTuple() -- crashed under Dynamo with Asser,https://github.com/pytorch/pytorch/pull/189021,a6f0921ac01c1d6f0d4628ecc128409de3bb53147cd6aec639c39f8eacf2df52 references,pr,189021,pr,189052,medium,pr.body,Stack from ghstack (oldest at bottom): #185424 #189053 #189052 #189022 #189051 -> #189021 Constructing a tuple subclass with no arguments -- e.g. class MyTuple(tuple): pass; MyTuple() -- crashed under D,https://github.com/pytorch/pytorch/pull/189021,2aa8197f36c3dcf47721c3ea5eb76db55285349a30ac0dc64303c5944ac5ffe9 references,pr,189021,pr,189053,medium,pr.body,Stack from ghstack (oldest at bottom): #185424 #189053 #189052 #189022 #189051 -> #189021 Constructing a tuple subclass with no arguments -- e.g. class MyTuple(tuple): pass; MyTuple() -- crashed,https://github.com/pytorch/pytorch/pull/189021,0c302d1f9987a28b7d81ab0dca8b1789b34ff17e171d9db4b1272e340b817d8c references,pr,189021,pr,189051,medium,pr.comments[1].body,Starting merge as part of PR stack under #189051,https://github.com/pytorch/pytorch/pull/189021,12a542a111d78a264dfc1548413232a7b0726074e51dc0f506b1bcf5af29a663 review guidance,pr,189021,pr,185424,high,pr.reviews[0].body,"Thanks, looks good.",https://github.com/pytorch/pytorch/pull/189021,8ae810fa3c947d6cfeb942c635ced710fa99bb77712f6bf55bca06da34bfada9 review guidance,pr,189021,pr,189021,high,pr.reviews[0].body,"Thanks, looks good.",https://github.com/pytorch/pytorch/pull/189021,953974bff2546afea7dc88c012717e47031dd3889a450115fe75d018cc217cbb review guidance,pr,189021,pr,189022,high,pr.reviews[0].body,"Thanks, looks good.",https://github.com/pytorch/pytorch/pull/189021,10dfaf936c7c4a3c5ef2ad84bcb8a639e2c2d6bbe2f159929343cf0e879407dd review guidance,pr,189021,pr,189051,high,pr.reviews[0].body,"Thanks, looks good.",https://github.com/pytorch/pytorch/pull/189021,c724cbf9fe276ea996053e6a187ee15e5aeb2fa4b5d786b7827d99980f3c5f64 review guidance,pr,189021,pr,189052,high,pr.reviews[0].body,"Thanks, looks good.",https://github.com/pytorch/pytorch/pull/189021,3635d6add390ac48a26c34c64fd268a00b540a3207ccf53ebf2d63f04db53833 review guidance,pr,189021,pr,189053,high,pr.reviews[0].body,"Thanks, looks good.",https://github.com/pytorch/pytorch/pull/189021,1ce0e9d253213306f0970b7139082ff3fcbe37a465e88f6a6970b12e13d5a798 closes,pr,188931,issue,188891,high,pr.closingIssuesReferences,pr #188931 declares a closing reference to issue #188891.,https://github.com/pytorch/pytorch/pull/188931,71b9f23a409e70213ffdd4176a0ebe31ac66e563e8a78bfbac39600c22895b11 closes,pr,188931,issue,188891,high,pr.body,"essfully: .venv/bin/python test/test_nn.py -k test_linear_1d_weight_bias .venv/bin/python test/test_ops.py -k ""nn_functional_linear"" fixes: #188891 This PR was authored with the help of an AI assistant.",https://github.com/pytorch/pytorch/pull/188931,2fb232ad522e994ce2ac19705868c23ce52d814c38fafcf1b0b8bd0cbe7d2c30 closes,pr,188913,issue,188150,high,pr.closingIssuesReferences,pr #188913 declares a closing reference to issue #188150.,https://github.com/pytorch/pytorch/pull/188913,104683071572f4648c194eb7654ac87c42aa0dda332f76cc919948b234925d23 closes,pr,188913,issue,188150,high,pr.body,"Fixes #188150 Problem With mode=""reduce-overhead"", cudagraph trees record one graph per distinct dynamic-shape key: deferred_cudagraphify keys fn_cache o",https://github.com/pytorch/pytorch/pull/188913,8a52db1a90f2ff67aba3765b79d9d3df0686e0df40a1f0570280398ee785d104 closes,pr,188998,issue,188492,high,pr.closingIssuesReferences,pr #188998 declares a closing reference to issue #188492.,https://github.com/pytorch/pytorch/pull/188998,195413ec4eb8a13a071370d20bdffbe7526f5a960d00c3c35772cb59871d071e closes,pr,188998,issue,188492,high,pr.body,"Fixes #188492 Under Triton 3.8 on Hopper (SM90), bf16 Swin-Transformer and similar bf16 workloads compiled with torch.compile(mode=""max-autotune"") showed",https://github.com/pytorch/pytorch/pull/188998,be7fbf11d2b429072da4b6aa824a07eab176c1c99c401d0994f001a485f6b0fd closes,pr,189119,issue,189118,high,pr.closingIssuesReferences,pr #189119 declares a closing reference to issue #189118.,https://github.com/pytorch/pytorch/pull/189119,4df33b52b50c8207ad3df110409c6e6f6d3d91383dfa929c1d70669e44b90296 closes,pr,189119,issue,189118,high,pr.body,Issue Fixes #189118 Summary What Problem This Solves upload_to_s3_artifacts() creates three separate intermediate artifacts for CI test output: test-reports-*.,https://github.com/pytorch/pytorch/pull/189119,53df107561c9211912c3f4f0f91671cde026d2388eba86a738e95f8cce8d3672 closes,pr,188996,issue,188711,high,pr.closingIssuesReferences,pr #188996 declares a closing reference to issue #188711.,https://github.com/pytorch/pytorch/pull/188996,5b80c54a7977141cbdb829a0becea8f602a1cdc37982d124a25422f0c13c2cca closes,pr,188996,issue,188711,high,pr.body,Fixes #188711 The test_main_loop_scaling FP8 test was failing because Inductor's scale-shape validator (get_scaling_options/is_desired_scaling) rejected,https://github.com/pytorch/pytorch/pull/188996,b8842d3fc7577d0c59ef43e02dc74c8f1a1638f3211ebbb35ac342af58443be1 references,pr,188996,issue,188721,medium,pr.comments[0].body,"akiness on trunk: B200 Smoke Tests / linux-jammy-cuda13.0-py3.12-gcc11-sm100 / test-osdc (smoke_b200, 1, 1, mt-l-x86iamx-22-225-b200) (gh) (#188721) test/inductor/test_flex_flash.py::TestFlexFlashCUDA::test_captured_table_int64_index_cuda This comment was automatically generat...",https://github.com/pytorch/pytorch/pull/188996,e09fc55814f0331a60e2008732c7c78d151bb302880e530edaad814957e32610 closes,pr,189010,issue,188866,high,pr.closingIssuesReferences,pr #189010 declares a closing reference to issue #188866.,https://github.com/pytorch/pytorch/pull/189010,b45a0f5a57c3c6c12497b3e4a19e2e5fffab1be720cf699ebbd855b9709c0102 closes,pr,189010,issue,188866,high,pr.body,"Related Isue closes #188866 Why The tvm backend depends on tvm.relay and tvm.contrib.graph_executor, removed upstream in TVM 0.20, so every pip-installable TVM (apache",https://github.com/pytorch/pytorch/pull/189010,fafb55a7ac42dbf06ba489a9e80a56afcb04a4bff9f43bb8a74858c9f6a77bbe closes,pr,189158,issue,189157,high,pr.closingIssuesReferences,pr #189158 declares a closing reference to issue #189157.,https://github.com/pytorch/pytorch/pull/189158,7a52b98d12d18c48e34d6ad294f880b6189ca85480f8fde2c883210a4f7003a7 closes,pr,189158,issue,189157,high,pr.body,Fixing an Issue Issue Fixes #189157 Summary What Problem This Solves _download_artifact() already detected when an artifact name's runattemptN did not match the requested work,https://github.com/pytorch/pytorch/pull/189158,3dfccb4da74c270ef000c05b62b56e7834efea92536a49a9503c9a7f907fabe8 closes,pr,181895,issue,181807,high,pr.closingIssuesReferences,pr #181895 declares a closing reference to issue #181807.,https://github.com/pytorch/pytorch/pull/181895,748beec50f884ea022a6e1a7e2667753c9e282201a04b79484da5367bd8fe113 closes,pr,181895,issue,181807,high,pr.body,"sagree on int64 linspace for the same start, end, and steps. Use double for the integral CUDA path to match RangeFactoriesKernel.cpp. Fixes #181807",https://github.com/pytorch/pytorch/pull/181895,d7e93ef6682b83d3501df72ecaaa5774958986bc106ecf2d518d3e9bd422a32f review guidance,pr,181895,issue,181807,high,pr.reviews[0].body,Needs GB300 bench across variety of sizes showing no impact,https://github.com/pytorch/pytorch/pull/181895,68171ec599c933024eaef1e1276df4aa07ee75a040cc5d132c4461e9368c5d06 references,pr,188825,pr,157149,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 -> #188825 #188834 #188824 #157149 #188639 #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188825,2c99004d9fc5c418c130fa0b0283568cb2aa79c40b7bf64bf672d8d6df2d808a references,pr,188825,pr,187690,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 -> #188825 #188834 #188824 #157149 #188639 #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188825,23680afe8ea9ef1bdae44b79aac73a72d1e7fb489eb75d89b2131133d0f7d9e0 references,pr,188825,pr,187744,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 -> #188825 #188834 #188824 #157149 #188639 #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188825,3227595726678a94e5cba55c90caa4180c1fbcd9be564575ff712abbb278ba25 references,pr,188825,pr,188004,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 -> #188825 #188834 #188824 #157149 #188639 #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188825,9a452aaeb5086b979b311a6d0045b8ba9760bb410f5bd25d61dc3c27a3820abe references,pr,188825,pr,188638,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 -> #188825 #188834 #188824 #157149 #188639 #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188825,dcd00f69c7ee8f2aafc9291e8336ad5995e40393f4e0c427839f4744475b35c7 references,pr,188825,pr,188639,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 -> #188825 #188834 #188824 #157149 #188639 #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188825,88fd5df707006a614dfdbcfae312eb1a71a512935938bc62d912536b844d17d2 references,pr,188825,pr,188824,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 -> #188825 #188834 #188824 #157149 #188639 #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188825,38e2d017670a1d33e42ab90df73107e2a897387c68df4efde724f39999cd1744 references,pr,188825,pr,188834,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 -> #188825 #188834 #188824 #157149 #188639 #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188825,315447884db9559477b2b38a3cd17078c5b1d83cf9e46f0564a6c7b0ebda7d86 references,pr,188825,pr,189024,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 -> #188825 #188834 #188824 #157149 #188639 #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188825,4a005108ace968f722ae50210cc4379c12ed032424ff396131bf4fba0aa29af6 references,pr,188834,pr,157149,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 -> #188834 #188824 #157149 #188639 #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188834,c42e541f05bbb5ed43ff5e167e8a507391073c4df30054710abdd635d7d5dc81 references,pr,188834,pr,187690,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 -> #188834 #188824 #157149 #188639 #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188834,a954331c60e9a32c5edb5c193250cf349431b4b1ad167560c5a916206337623d references,pr,188834,pr,187744,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 -> #188834 #188824 #157149 #188639 #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188834,c40f66c759c4c1926a8bc66a7d8bb141dff4d0d737ce6e45b72d38ef292e4718 references,pr,188834,pr,188004,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 -> #188834 #188824 #157149 #188639 #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188834,12b928b36b15c072553f5715b64a8a23df683ee6360e3e9add7cce56ad572f94 references,pr,188834,pr,188638,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 -> #188834 #188824 #157149 #188639 #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188834,178ae165cb08a5167651892896f304977ec2cbac659cd3a030344b7e16701940 references,pr,188834,pr,188639,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 -> #188834 #188824 #157149 #188639 #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188834,612621241e8b492aa855cfe506509819e24165df0f338f5dcc0a3f15fda573d8 references,pr,188834,pr,188824,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 -> #188834 #188824 #157149 #188639 #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188834,4684d341b4e0f4cffd421b81cbde8cc6709387d5ec6bd72eac99b59b6a78042d references,pr,188834,pr,188825,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 -> #188834 #188824 #157149 #188639 #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188834,7afbe2997807bb22089ec73f3da47148e440c9803867699f51738183930af242 references,pr,188834,pr,189024,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 -> #188834 #188824 #157149 #188639 #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188834,12e9883142d7e8d2e49b54e4034cda3959033588bd8bcf58e1a87ff15a359fa8 references,pr,188834,pr,189024,medium,pr.comments[1].body,Starting merge as part of PR stack under #189024,https://github.com/pytorch/pytorch/pull/188834,b13363bce8a66e009c789b419ef240b6eed57ea70c801b25daa5ef4781f3b1e7 review guidance,pr,188834,pr,157149,high,pr.reviews[0].body,test coverage?,https://github.com/pytorch/pytorch/pull/188834,3b21d1eee9b8e58a620a455ee39ac48e78322694cc84843af683fb3075abaf61 review guidance,pr,188834,pr,187690,high,pr.reviews[0].body,test coverage?,https://github.com/pytorch/pytorch/pull/188834,24436352914425aab5d9954af428bc9b5da841408fe19c84c567de370f464c68 review guidance,pr,188834,pr,187744,high,pr.reviews[0].body,test coverage?,https://github.com/pytorch/pytorch/pull/188834,f3b54e758c80dff76e34e33b02b9040db401d23fcc6eba6b467b72c3a7f12575 review guidance,pr,188834,pr,188004,high,pr.reviews[0].body,test coverage?,https://github.com/pytorch/pytorch/pull/188834,eae65968a2245d7fa88c93458519f84003eba6024dff801480999c0f4002d394 review guidance,pr,188834,pr,188638,high,pr.reviews[0].body,test coverage?,https://github.com/pytorch/pytorch/pull/188834,c80bb82c0af2d8506e5f71a59bfa10969151e7c00b809a30c5cddec9d3a5b015 review guidance,pr,188834,pr,188639,high,pr.reviews[0].body,test coverage?,https://github.com/pytorch/pytorch/pull/188834,2442aeac9745d106d4b67d439080302660dea4a93e1cc0983daee7d7a7c18ed9 review guidance,pr,188834,pr,188824,high,pr.reviews[0].body,test coverage?,https://github.com/pytorch/pytorch/pull/188834,1a56d311968fbaeb6118561435c61c91eee9267338c11201d80cddf0f827e0c9 review guidance,pr,188834,pr,188825,high,pr.reviews[0].body,test coverage?,https://github.com/pytorch/pytorch/pull/188834,c3fccec677efac375e34aee9ea553aa69a10c7a925205f3e8d6ca43ff964a424 review guidance,pr,188834,pr,188834,high,pr.reviews[0].body,test coverage?,https://github.com/pytorch/pytorch/pull/188834,37ec2dbbd8ace0626d4018ebcfb50df9be59fdd24671a31e72f250d797078e94 review guidance,pr,188834,pr,189024,high,pr.reviews[0].body,test coverage?,https://github.com/pytorch/pytorch/pull/188834,4d9e54a1ede115bb2dff892a88d2c8866209ac0a9ab0cc21c4c3465104b0e9ba references,pr,188824,pr,157149,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 -> #188824 #157149 #188639 #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188824,6e58d21bd3aee201b04d997225c2e850a581f1da166754e5ddf0aaedcd2c98fb references,pr,188824,pr,187690,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 -> #188824 #157149 #188639 #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188824,6bab0eaae77c57d197c224cef6d6a333dddb5389c79e9f240280ba9e6f4b668a references,pr,188824,pr,187744,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 -> #188824 #157149 #188639 #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188824,d564c7f5a7d4e2a32e58a3aac0d11769acc1da048eb8967d06c7142c5d0b7d00 references,pr,188824,pr,188004,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 -> #188824 #157149 #188639 #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188824,26f20df89a19cb4c9ee71756793b596d3569aea2550f904e7206e64b4b443aa7 references,pr,188824,pr,188638,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 -> #188824 #157149 #188639 #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188824,683c7887232b186a7ee27b2a30d2d2135c59e80996a3240a5e7d2d8a21c6bf38 references,pr,188824,pr,188639,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 -> #188824 #157149 #188639 #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188824,3d6690ac1b0b28f6034982a6b85cab480fa5cc4273681262e6fd96df4f4bdf63 references,pr,188824,pr,188825,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 -> #188824 #157149 #188639 #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188824,4f84314771a240596564651fcc403456441e33224478f086403c9df691d3b47d references,pr,188824,pr,188834,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 -> #188824 #157149 #188639 #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188824,83450f82ec885ffe5ebd159aa86e535fafdc43e80496d0da52d0c6933a7600ad references,pr,188824,pr,189024,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 -> #188824 #157149 #188639 #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188824,ffb2a21312013febca78c25ffd9371b6780612b00b1dc982d9ce4460f37b42f6 review guidance,pr,188824,pr,157149,high,pr.reviews[0].body,LGTM; I verified the tests themselves and against the doctests they were taken from.,https://github.com/pytorch/pytorch/pull/188824,853827d2b22a0faf16b9ac73d931ecb27cf7cc52fda6cb099f6e5738a4d11636 review guidance,pr,188824,pr,187690,high,pr.reviews[0].body,LGTM; I verified the tests themselves and against the doctests they were taken from.,https://github.com/pytorch/pytorch/pull/188824,d0d550cd5ec01b2615295c0fdf6e55920c8969529fb02c78c9ce9bcaa23438c7 review guidance,pr,188824,pr,187744,high,pr.reviews[0].body,LGTM; I verified the tests themselves and against the doctests they were taken from.,https://github.com/pytorch/pytorch/pull/188824,83a798864c89f78f1fa73165b7022eb81da131515f7fa8e8f662f03ae1f2c8bc review guidance,pr,188824,pr,188004,high,pr.reviews[0].body,LGTM; I verified the tests themselves and against the doctests they were taken from.,https://github.com/pytorch/pytorch/pull/188824,4b6b40e214a95dcf156f5a5941183e8551c7f97a47fd3bc17c26bbb7bb21169e review guidance,pr,188824,pr,188638,high,pr.reviews[0].body,LGTM; I verified the tests themselves and against the doctests they were taken from.,https://github.com/pytorch/pytorch/pull/188824,6bc38b857be10f7ba39cb75c1887fcb22373c358423347a031ad027d35f9706c review guidance,pr,188824,pr,188639,high,pr.reviews[0].body,LGTM; I verified the tests themselves and against the doctests they were taken from.,https://github.com/pytorch/pytorch/pull/188824,44d1f1c462bccfc93d93a5cd2826a4d936eeef55e63678b3815b7ac639b5bf13 review guidance,pr,188824,pr,188824,high,pr.reviews[0].body,LGTM; I verified the tests themselves and against the doctests they were taken from.,https://github.com/pytorch/pytorch/pull/188824,685e16f222c488beede22d7b313f8f2e6a2b8155917f8e1f21dfd8c816a669c7 review guidance,pr,188824,pr,188825,high,pr.reviews[0].body,LGTM; I verified the tests themselves and against the doctests they were taken from.,https://github.com/pytorch/pytorch/pull/188824,0b474d5dca90a17f2960a88f0291755be36b444d9d08089f7b721cae2d619980 review guidance,pr,188824,pr,188834,high,pr.reviews[0].body,LGTM; I verified the tests themselves and against the doctests they were taken from.,https://github.com/pytorch/pytorch/pull/188824,ac4ab35131a963205133575ae05956805f55149a8c077b641834f732ad8c5127 review guidance,pr,188824,pr,189024,high,pr.reviews[0].body,LGTM; I verified the tests themselves and against the doctests they were taken from.,https://github.com/pytorch/pytorch/pull/188824,cf45bc85146044f2a50230af54ed80aee8d87a24d8063c43e29e99788782ea57 references,pr,157149,pr,187690,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 -> #157149 #188639 #188638 #188004 #187744 #187690 Motivation and Example Explictly close all open generators in compile_subgraph to ensure that all remaining finally blocks are executed. In,https://github.com/pytorch/pytorch/pull/157149,c451f94957cb53d97f8fd80ec56584df7e3eff1e98c6492cb8e689738bad8194 references,pr,157149,pr,187744,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 -> #157149 #188639 #188638 #188004 #187744 #187690 Motivation and Example Explictly close all open generators in compile_subgraph to ensure that all remaining finally blocks are exec,https://github.com/pytorch/pytorch/pull/157149,8d04a3d1daf192d0e10a032fbc914fbd940d11793a30e64dff5e59f2348a1094 references,pr,157149,pr,188004,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 -> #157149 #188639 #188638 #188004 #187744 #187690 Motivation and Example Explictly close all open generators in compile_subgraph to ensure that all remaining finally blocks,https://github.com/pytorch/pytorch/pull/157149,aaaf34c51eaeaac8532ce0fba292f6095512279676e70ca390d25c19f136a6bf references,pr,157149,pr,188638,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 -> #157149 #188639 #188638 #188004 #187744 #187690 Motivation and Example Explictly close all open generators in compile_subgraph to ensure that all remaining finally,https://github.com/pytorch/pytorch/pull/157149,2801738c053c43c46c1e4de3c1fc2ec3cad12403bbc353081e63645458fd3c86 references,pr,157149,pr,188639,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 -> #157149 #188639 #188638 #188004 #187744 #187690 Motivation and Example Explictly close all open generators in compile_subgraph to ensure that all remaining,https://github.com/pytorch/pytorch/pull/157149,1ba284fdf6ae7ebd616cb992f3797f14ecf56a35e44040136ee4898e66d4f7de references,pr,157149,pr,188824,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 -> #157149 #188639 #188638 #188004 #187744 #187690 Motivation and Example Explictly close all open generators in compile_subgraph to ensure,https://github.com/pytorch/pytorch/pull/157149,af7bc17ca4fde886d8342861678a3a192045cf0927894948b012f046c6cd97af references,pr,157149,pr,188825,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 -> #157149 #188639 #188638 #188004 #187744 #187690 Motivation and Example Explictly close all open generators in compile_su,https://github.com/pytorch/pytorch/pull/157149,6b383605215e080889b7c0a53d70404bbe364f221b92ac8879b4e010f424f179 references,pr,157149,pr,188834,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 -> #157149 #188639 #188638 #188004 #187744 #187690 Motivation and Example Explictly close all open generators in compile_subgraph t,https://github.com/pytorch/pytorch/pull/157149,6a50c54a8537f8206007f2439f021ddc41424a8177dc3e73379b7b04bca6270f references,pr,157149,pr,189024,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 -> #157149 #188639 #188638 #188004 #187744 #187690 Motivation and Example Explictly close all open generators in co,https://github.com/pytorch/pytorch/pull/157149,ad406b163f06924aeb06239e5c47537ae2c79809b16a0c10d7230c452ff03743 references,pr,188639,pr,157149,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 -> #188639 #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188639,3ed4e2e245963c78d9b3a5f4eabf4104f6614682fd6b53c394d7cc196f0b6cb9 references,pr,188639,pr,187690,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 -> #188639 #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188639,a84c5aa47506d18a4cee6f9b6c2d4a0ce6a8a1cfeb8d8e11536a6c82e2d7e214 references,pr,188639,pr,187744,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 -> #188639 #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188639,868c13bb0a7f120f42a5aa035b5376bcc6e59e47fe789e662769737ecdecd8bb references,pr,188639,pr,188004,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 -> #188639 #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188639,7a64d75560b4ae252a294b038bd32bc1397ebb6c745a3ef7c7c1fd6bfa79cac3 references,pr,188639,pr,188638,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 -> #188639 #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188639,ff1e0247b829ec294a878e22ffd9b9ae4d4e79ee05c60f9c3165e27278909b10 references,pr,188639,pr,188824,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 -> #188639 #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188639,1b866880fa6c0f1743c7bf298d2096898bebe208e53506d9579cb933b1489ea5 references,pr,188639,pr,188825,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 -> #188639 #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188639,2e5e0f37e4faead9eb3be56f249712f7f294b47459dd57f45a85c487d0e1f1cd references,pr,188639,pr,188834,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 -> #188639 #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188639,093275fc74b01045e3425ac0bf85da830d405bdf790124574ab29168ccc8cc13 references,pr,188639,pr,189024,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 -> #188639 #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188639,bc7028228e3f5ababde8e0d7facb4b4fd7b43187b372d8c3a848cec11442c038 references,pr,188638,pr,157149,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 #188639 -> #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188638,418ac2e0887aa9cbcf8e8ed800e41f3806b019086984180cfaa1f6aa9fb17979 references,pr,188638,pr,187690,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 #188639 -> #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188638,47dca81196c6741db20dce09a67082208bdd3364eb33bf12e57a2cd72d48cee5 references,pr,188638,pr,187744,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 #188639 -> #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188638,b25cc51e465d6a9a22f9fd32f5202855897f4943800c20b7fae6e55775df1b40 references,pr,188638,pr,188004,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 #188639 -> #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188638,be7b1141499932c313dc886291367923580bc41a3519459dd2b58c32d4cbfdb6 references,pr,188638,pr,188639,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 #188639 -> #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188638,7d3ed13f9f7520eb26614698b090dd4ff7846b63fbbd62e23d43548188a359ed references,pr,188638,pr,188824,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 #188639 -> #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188638,5cac339ad2064b7f1cb43569d2f7937ce45a8730d6249088b8487525ef6bea13 references,pr,188638,pr,188825,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 #188639 -> #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188638,143b2f70eb71d10d176ffa97c3ed1a9ef7954e3a5fbf149c4d7450cadf3abb52 references,pr,188638,pr,188834,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 #188639 -> #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188638,86f77aa94e799ad67b3093d3f892202af41218dfaa5014666a61ae94f695dd92 references,pr,188638,pr,189024,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 #188639 -> #188638 #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188638,131b954f55a6dfb526222446bd143205e81df255a1367ff8145edcfe6c4a8f59 review guidance,pr,188638,pr,157149,high,pr.reviews[1].body,One minor question.,https://github.com/pytorch/pytorch/pull/188638,b39c3004c5edb75908afc282b5aec14272098fcc6635705567c5bc98b59237c1 review guidance,pr,188638,pr,187690,high,pr.reviews[1].body,One minor question.,https://github.com/pytorch/pytorch/pull/188638,810331ec98edae7a37773a3ff6cd96ab42cc757c251a8a01144a2509461ff934 review guidance,pr,188638,pr,187744,high,pr.reviews[1].body,One minor question.,https://github.com/pytorch/pytorch/pull/188638,35dab513829c669d2939f82803a4c17d726c283a991b7e275d3b6a4ad0830d69 review guidance,pr,188638,pr,188004,high,pr.reviews[1].body,One minor question.,https://github.com/pytorch/pytorch/pull/188638,b4ae8e30aecf36310d71e473e3fa770c7d2039d67c8f36d7cf6a2c1b970e7236 review guidance,pr,188638,pr,188638,high,pr.reviews[1].body,One minor question.,https://github.com/pytorch/pytorch/pull/188638,cdaefa7bb209db0a5e7dbd6f44e11890de608bbc5028216b668da4321dff7379 review guidance,pr,188638,pr,188639,high,pr.reviews[1].body,One minor question.,https://github.com/pytorch/pytorch/pull/188638,a2dc5ed18062c061d88af7163e04bda5490daf57dea0ac2f110efac6018a07dd review guidance,pr,188638,pr,188824,high,pr.reviews[1].body,One minor question.,https://github.com/pytorch/pytorch/pull/188638,374cc01388a2fbf2df00c84bf4cac8a1b1faeaf580439b65326ba70bf2f74540 review guidance,pr,188638,pr,188825,high,pr.reviews[1].body,One minor question.,https://github.com/pytorch/pytorch/pull/188638,387996a18616d582a5860404c4c40e8d8b46176ac6519d857c4161fc7b5ff872 review guidance,pr,188638,pr,188834,high,pr.reviews[1].body,One minor question.,https://github.com/pytorch/pytorch/pull/188638,ed62b64cf30576053bec5a2d01b9914f6931c67723bb23947719ce81b8090806 review guidance,pr,188638,pr,189024,high,pr.reviews[1].body,One minor question.,https://github.com/pytorch/pytorch/pull/188638,b0f39001b4dd41832d355d55485d9e05d268e7a286b6201c614731ca18f11529 references,pr,187744,pr,157149,medium,pr.body,"Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 #188639 #188638 #188004 -> #187744 #187690 This mirrors CPython behavior where each generator has its own exception stack, which is linked",https://github.com/pytorch/pytorch/pull/187744,5c1351489313c9a12e347283c9afb8576da2cf69f66d37de51469344a8bd1d6b references,pr,187744,pr,187690,medium,pr.body,"Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 #188639 #188638 #188004 -> #187744 #187690 This mirrors CPython behavior where each generator has its own exception stack, which is linked for the duration the generator is active Re",https://github.com/pytorch/pytorch/pull/187744,e229c5dfefe903889999fcc75a9b79e146a5fb800b9a9289f6431ed0f4e91c2c references,pr,187744,pr,188004,medium,pr.body,"Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 #188639 #188638 #188004 -> #187744 #187690 This mirrors CPython behavior where each generator has its own exception stack, which is linked for the duration the gen",https://github.com/pytorch/pytorch/pull/187744,60090f387360706dc11431526ee8b7921b236d4f9c395446d040ea81a0de06b9 references,pr,187744,pr,188638,medium,pr.body,"Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 #188639 #188638 #188004 -> #187744 #187690 This mirrors CPython behavior where each generator has its own exception stack, which is linked for the duration",https://github.com/pytorch/pytorch/pull/187744,2f5c8676c47e6f5598d5ee79f54798738d1108856d7db731096764d973eecadb references,pr,187744,pr,188639,medium,pr.body,"Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 #188639 #188638 #188004 -> #187744 #187690 This mirrors CPython behavior where each generator has its own exception stack, which is linked for the",https://github.com/pytorch/pytorch/pull/187744,ee2e9539b7f87db096e13513093ebf7117fd7e05b6a757a26f1a4b62710d221a references,pr,187744,pr,188824,medium,pr.body,"Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 #188639 #188638 #188004 -> #187744 #187690 This mirrors CPython behavior where each generator has its own exception stack, which is",https://github.com/pytorch/pytorch/pull/187744,22a58dd243b90e67e9967bcb602c9ed37941a6597278984336ad94ceefdce4ea references,pr,187744,pr,188825,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 #188639 #188638 #188004 -> #187744 #187690 This mirrors CPython behavior where each generator has its own exception,https://github.com/pytorch/pytorch/pull/187744,f4954c950119708d06262f2be67afce7705649fb792ba136ce9d1fc3def4e2a1 references,pr,187744,pr,188834,medium,pr.body,"Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 #188639 #188638 #188004 -> #187744 #187690 This mirrors CPython behavior where each generator has its own exception stack,",https://github.com/pytorch/pytorch/pull/187744,ed951d27259e5c842bdef85e7b3803a4db8b9a67d3d5c6fd56bac9de7d615e73 references,pr,187744,pr,189024,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 #188639 #188638 #188004 -> #187744 #187690 This mirrors CPython behavior where each generator has its own e,https://github.com/pytorch/pytorch/pull/187744,e75cbfcf0de38bee40b8a844d7f7d6b232cc54e2cee6e699a0cc6b6780a1095a references,pr,187744,pr,188639,medium,pr.comments[1].body,Starting merge as part of PR stack under #188639,https://github.com/pytorch/pytorch/pull/187744,69f200072d8a30ed888bbf18e729b4a5c1c516c1b22b5eb352006d7467249f99 references,pr,187690,pr,157149,medium,pr.body,"Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 #188639 #188638 #188004 #187744 -> #187690 Refactor .close, .send and .throw to follow CPython semantics.",https://github.com/pytorch/pytorch/pull/187690,126d07b7e5e5523b8afea874e8e54d61c5082c952e42762d32ce317ed07ebf9a references,pr,187690,pr,187744,medium,pr.body,"Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 #188639 #188638 #188004 #187744 -> #187690 Refactor .close, .send and .throw to follow CPython semantics.",https://github.com/pytorch/pytorch/pull/187690,14ba031bf2d5ce05404e1301bd088c18483f4d5d1f421b80ebb92bffa8e62caf references,pr,187690,pr,188004,medium,pr.body,"Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 #188639 #188638 #188004 #187744 -> #187690 Refactor .close, .send and .throw to follow CPython semantics.",https://github.com/pytorch/pytorch/pull/187690,78a23976e3ed7586ace73434e04bf9689a15c7198d6ae5abe823f425a20b583e references,pr,187690,pr,188638,medium,pr.body,"Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 #188639 #188638 #188004 #187744 -> #187690 Refactor .close, .send and .throw to follow CPython semantics.",https://github.com/pytorch/pytorch/pull/187690,17912f4da340aaa4ad19544f2a4c55a969914b040280c83801b0a925ab1b7798 references,pr,187690,pr,188639,medium,pr.body,"Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 #188639 #188638 #188004 #187744 -> #187690 Refactor .close, .send and .throw to follow CPython semantics.",https://github.com/pytorch/pytorch/pull/187690,d54dd4de95674d7ca06badd0172a85a1088c4aa0808fda592045d7be5d067d67 references,pr,187690,pr,188824,medium,pr.body,"Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 #188639 #188638 #188004 #187744 -> #187690 Refactor .close, .send and .throw to follow CPython semantics.",https://github.com/pytorch/pytorch/pull/187690,9660d74931d413c498c038f62eacacca4a080a26a83c231d6ac87e83e56d7ec3 references,pr,187690,pr,188825,medium,pr.body,"Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 #188639 #188638 #188004 #187744 -> #187690 Refactor .close, .send and .throw to follow CPython semantics.",https://github.com/pytorch/pytorch/pull/187690,576407ed08a57c48a742fd728deba215a0de4f7516f3f128da423b9e3e04b19e references,pr,187690,pr,188834,medium,pr.body,"Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 #188639 #188638 #188004 #187744 -> #187690 Refactor .close, .send and .throw to follow CPython semantics.",https://github.com/pytorch/pytorch/pull/187690,026c0de78e43780282ff77033033e581adcf50d6313e75ded919a6cf5d136da2 references,pr,187690,pr,189024,medium,pr.body,"Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 #188639 #188638 #188004 #187744 -> #187690 Refactor .close, .send and .throw to follow CPython semantics.",https://github.com/pytorch/pytorch/pull/187690,d714642554d3d8b054a030f94fc7779d9705cd057c4cdbbf7a8d9b9fd6ec3dd9 closes,pr,189141,issue,188680,high,pr.closingIssuesReferences,pr #189141 declares a closing reference to issue #188680.,https://github.com/pytorch/pytorch/pull/189141,cf2778cce6ca0963cc17d14239d5dde69a57ae892f491dc9216afbc0ac298950 closes,pr,189141,issue,188680,high,pr.body,"Fixes #188680 torch.allclose treats +0.0 and -0.0 as equal (IEEE 754 defines them as equal under ==), so the default compile-vs-eager correctness check i",https://github.com/pytorch/pytorch/pull/189141,e65a9dd1e31e323f3ddf4db6de366d20c0161910cbe265316c97140fc8782020 references,pr,185424,pr,189021,medium,pr.body,Stack from ghstack (oldest at bottom): -> #185424 #189053 #189052 #189022 #189051 #189021 Adds the CPython Dynamo expected-failure relevance rankings together with the agentic-loop docs used to work the highest-value CPython cove,https://github.com/pytorch/pytorch/pull/185424,0456d9fd15ec1874d33077891d896c0ff50664183b815001176164aafc78caf8 references,pr,185424,pr,189022,medium,pr.body,Stack from ghstack (oldest at bottom): -> #185424 #189053 #189052 #189022 #189051 #189021 Adds the CPython Dynamo expected-failure relevance rankings together with the agentic-loop docs used to work the highest-va,https://github.com/pytorch/pytorch/pull/185424,df7ff46ede1148878b848902840fc8cc5f5f01e2f31de4ad31c869454a3136bf references,pr,185424,pr,189051,medium,pr.body,Stack from ghstack (oldest at bottom): -> #185424 #189053 #189052 #189022 #189051 #189021 Adds the CPython Dynamo expected-failure relevance rankings together with the agentic-loop docs used to work the highest-value CPyt,https://github.com/pytorch/pytorch/pull/185424,fa99e28a98db5e98447814dfa60cbf69ab910281cd1d5e85d22ec966dbde4dcd references,pr,185424,pr,189052,medium,pr.body,Stack from ghstack (oldest at bottom): -> #185424 #189053 #189052 #189022 #189051 #189021 Adds the CPython Dynamo expected-failure relevance rankings together with the agentic-loop docs used to work the hi,https://github.com/pytorch/pytorch/pull/185424,ef07dd34a72c87ccbedfd3c60d3c1e098fe8797dfe50fde3d66055284db13855 references,pr,185424,pr,189053,medium,pr.body,Stack from ghstack (oldest at bottom): -> #185424 #189053 #189052 #189022 #189051 #189021 Adds the CPython Dynamo expected-failure relevance rankings together with the agentic-loop docs used to wor,https://github.com/pytorch/pytorch/pull/185424,35daaac4c49cbf86af3f1db0a5ac2ec13e1ca182c370b1eb1c14bdd4b56511d4 references,pr,189053,pr,185424,medium,pr.body,Stack from ghstack (oldest at bottom): #185424 -> #189053 #189052 #189022 #189051 #189021 A custom hash that does integer arithmetic on id(self) -- e.g. class H(set): __hash__ = lambda s,https://github.com/pytorch/pytorch/pull/189053,e548ccd10f7cd028922958a23b96e57025a20d99547910b434268a5cac5ef2a2 references,pr,189053,pr,189021,medium,pr.body,Stack from ghstack (oldest at bottom): #185424 -> #189053 #189052 #189022 #189051 #189021 A custom hash that does integer arithmetic on id(self) -- e.g. class H(set): __hash__ = lambda self: int(id(self) & 0x7fffffff) -- crashed,https://github.com/pytorch/pytorch/pull/189053,f64560a98be3a6b45a6b594ebad95e794337f6d7d1f9a8365417484005c7b6ed references,pr,189053,pr,189022,medium,pr.body,Stack from ghstack (oldest at bottom): #185424 -> #189053 #189052 #189022 #189051 #189021 A custom hash that does integer arithmetic on id(self) -- e.g. class H(set): __hash__ = lambda self: int(id(self) & 0x7ffff,https://github.com/pytorch/pytorch/pull/189053,072aa56d249097b7ff45d61e858d9cf796d2e6fc55d485e472fac549ef67b14c references,pr,189053,pr,189051,medium,pr.body,Stack from ghstack (oldest at bottom): #185424 -> #189053 #189052 #189022 #189051 #189021 A custom hash that does integer arithmetic on id(self) -- e.g. class H(set): __hash__ = lambda self: int(id(self) & 0x7fffffff) --,https://github.com/pytorch/pytorch/pull/189053,b8410c931764250028a9b3520757ff9d07c1933d46cb1ecca6078376673f5ac3 references,pr,189053,pr,189052,medium,pr.body,Stack from ghstack (oldest at bottom): #185424 -> #189053 #189052 #189022 #189051 #189021 A custom hash that does integer arithmetic on id(self) -- e.g. class H(set): __hash__ = lambda self: int(id(self) &,https://github.com/pytorch/pytorch/pull/189053,97ceb02054af08e5ffcaf74786a4abb5e4c26b2d7c98b3ca4f3f250caa18b5d8 references,pr,189052,pr,185424,medium,pr.body,"Stack from ghstack (oldest at bottom): #185424 #189053 -> #189052 #189022 #189051 #189021 Iterating a deque under Dynamo did not detect mutation during iteration, and reversed(deque) pro",https://github.com/pytorch/pytorch/pull/189052,1852e71cc8d82c9d3f79429d7d313ba2871d09707a64c271d33e226907adbb63 references,pr,189052,pr,189021,medium,pr.body,"Stack from ghstack (oldest at bottom): #185424 #189053 -> #189052 #189022 #189051 #189021 Iterating a deque under Dynamo did not detect mutation during iteration, and reversed(deque) produced the wrong iterator type. DequeVariabl",https://github.com/pytorch/pytorch/pull/189052,8fb4fb812329a572667c6c5c36bc341939f4e170eeced46310585259b5cff9b8 references,pr,189052,pr,189022,medium,pr.body,"Stack from ghstack (oldest at bottom): #185424 #189053 -> #189052 #189022 #189051 #189021 Iterating a deque under Dynamo did not detect mutation during iteration, and reversed(deque) produced the wrong iterator ty",https://github.com/pytorch/pytorch/pull/189052,db234f72a4887cd299c9d409d423feb6dcd344d68a4389eb2832edc1181bd33e references,pr,189052,pr,189051,medium,pr.body,"Stack from ghstack (oldest at bottom): #185424 #189053 -> #189052 #189022 #189051 #189021 Iterating a deque under Dynamo did not detect mutation during iteration, and reversed(deque) produced the wrong iterator type. Dequ",https://github.com/pytorch/pytorch/pull/189052,c4ebfdc668d02d48583eea4e78ef66f9c6737360888d8a8ebeb38d4fe424ff4c references,pr,189052,pr,189053,medium,pr.body,"Stack from ghstack (oldest at bottom): #185424 #189053 -> #189052 #189022 #189051 #189021 Iterating a deque under Dynamo did not detect mutation during iteration, and reversed(deque) produced th",https://github.com/pytorch/pytorch/pull/189052,220acfb8bfa5b63b97176ac802c2ef6649e6a8843c813a01b441b2fd81ae3cc2 references,pr,189131,pr,188867,medium,pr.body,"Summary Split out of #188867 per review feedback (the KleidiAI SME feature and this change are logically independent, and bundling them made #188867 harder to review/re",https://github.com/pytorch/pytorch/pull/189131,6e1cc90fe6f64fe8adb50fac4151fd5f37b654e261a047f853b06d7dc65d1fd9 review guidance,pr,189131,pr,188867,high,pr.reviews[0].body,Pull request overview This PR updates PyTorch’s XNNPACK CMake integration to avoid building XNNPACK’s non-production microkernels-all target (unused by PyTorch) and to keep the GCC 14 warning workaround safe when that target is absent. Changes: Disable XNNPACK_BUILD_ALL_MICROKERNELS to skip compi...,https://github.com/pytorch/pytorch/pull/189131,6c760091f8b58703c3615a075d735c3a490baa2150a3f5b39cb464307afedd1e closes,pr,187989,issue,187988,high,pr.closingIssuesReferences,pr #187989 declares a closing reference to issue #187988.,https://github.com/pytorch/pytorch/pull/187989,fe057327037f32490fe1da1b029a196b7a12a104c13c14fe68196b78724a7838 closes,pr,187989,issue,187988,high,pr.body,Fixes #187988 What linear() accepts a 1-D weight (shape [in_features]): it contracts and drops the last input dimension. mps_linear (forward) already han,https://github.com/pytorch/pytorch/pull/187989,742070ab132d4e67745764059151997c8d75e595f10ad59bf4f3e3de4f57586d references,pr,187989,issue,188721,medium,pr.comments[0].body,"akiness on trunk: B200 Smoke Tests / linux-jammy-cuda13.0-py3.12-gcc11-sm100 / test-osdc (smoke_b200, 1, 1, mt-l-x86iamx-22-225-b200) (gh) (#188721) test/inductor/test_flex_flash.py::TestFlexFlashCUDA::test_cutedsl_captured_alias_views_keep_distinct_layouts_case_offset_cuda vl...",https://github.com/pytorch/pytorch/pull/187989,b780594508f22cbb4946a92876c2b8a32e6711c88c8f0d6f175b327865c25502 review guidance,pr,185309,pr,185309,high,pr.reviews[0].body,"LGTM! Thanks, @cyyever! Nit: could you also update ""Other functions"" in docs/source/sparse.md by adding vector_norm (and replacing obsolete native_norm with norm)?",https://github.com/pytorch/pytorch/pull/185309,223a585b4bb4007b4e1ca19c3e9dc8d7d0ef973a6292e8ed8714bdf1cda92328 review guidance,pr,188954,pr,188954,high,pr.reviews[0].body,"In general looks good, although 2. issue claude raised on catching the unsupported data types with an error message is worth adding. Another thing is on the testing: MPS device is not running the test_linalg.py (I think) where the more rigorous property tests for other backends live. We should pr...",https://github.com/pytorch/pytorch/pull/188954,cc4f9d4df7e86e3e0ed4a672b115e09f1e0fc8d2121876ecdaf4a9abf620899f review guidance,pr,189207,pr,188318,high,pr.reviews[0].body,"The issue is not localized to only SM 8.6, it first appeared on SM 8.6 because that's the primary testing vehicle of those models. I would prefer waiting for the next cuDNN release which should fix this issue (and bumping all dependencies) rather than adding additional churn here",https://github.com/pytorch/pytorch/pull/189207,d9ed2ad7992473bfea5f2eb5780240c8a6d625905daa8a242bc86e5c6d1cefe6 closes,pr,189125,issue,169111,high,pr.closingIssuesReferences,pr #189125 declares a closing reference to issue #169111.,https://github.com/pytorch/pytorch/pull/189125,07de1706810f19165a5ce08733058183a11628aa9f58648e6a4c3fadd150cd79 closes,pr,189125,issue,169111,high,pr.body,"Related Issue closes #169111 Why With dynamic=True, Dynamo passes SymInt scalars as graph inputs, and the tvm backend feeds them straight into torch.jit.trace, which fa",https://github.com/pytorch/pytorch/pull/189125,95ee91b05041ef2874410ac273ad9cda287f71ee7e44ab81e445bab6ad38e3d4 references,pr,189125,pr,189010,medium,pr.body,m in sys.modules so it runs with or without a tvm install The check sits on the trace path so it will not block the relax dispatch added in #189010. cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv @jiayisunx @kadeng @...,https://github.com/pytorch/pytorch/pull/189125,63e3e9a050296351bf074416fde2007058316bcf307b94d24bf289592b38948f review guidance,pr,183663,pr,183663,high,pr.reviews[0].body,Do we have test coverage for the case of an incremental launcher rebuild?,https://github.com/pytorch/pytorch/pull/183663,7afd2bba078d357a1f75572d04fdddc0fa4b625427f0cfd96ca277f54d16cedc review guidance,pr,181826,pr,181826,high,pr.reviews[0].body,Thanks!,https://github.com/pytorch/pytorch/pull/181826,4c3eb64ce9b31662671d0d2122a3ee2c8a2792bbb835fefec0ea551199247a10 closes,pr,188012,issue,187389,high,pr.closingIssuesReferences,pr #188012 declares a closing reference to issue #187389.,https://github.com/pytorch/pytorch/pull/188012,bb7fc90a259110197a44e8812fc7407e979fa3173e16517e6f55f473e9705c9e closes,pr,188012,issue,187389,high,pr.body,"Issue Fixes #187389 Summary Adds a non-MKL CPU reference implementation for CSR sparse triangular_solve. Supports float, double, complex float, complex double,",https://github.com/pytorch/pytorch/pull/188012,287bb375000d35d4ebc464c22b967d3d3ebd011b4fb09e6b3d2bf2de7b9feed2 references,pr,189289,pr,186918,medium,pr.body,"all test classes in test/distributed/test_c10d_pypg.py with hw_classification as part of the ongoing test refactoring initiative. Requires #186918 to be merged first. Changes TestDDPWithWorkSubclass → GENERIC: single-rank fake Python PG on CPU, tests DDP Work subclass dispatch...",https://github.com/pytorch/pytorch/pull/189289,c85465e88cb8f41c388b9330c3eb82b5110b47d53edc9015179e9991c4c3788d closes,pr,188784,issue,188128,high,pr.closingIssuesReferences,pr #188784 declares a closing reference to issue #188128.,https://github.com/pytorch/pytorch/pull/188784,43ffcb35b844f6832d6e091fe310aab4dc33190266b3a1106df6cdda1dc1854c closes,pr,188784,issue,188128,high,pr.body,"Fixes #188128 The fix has two parts: When no dtype is given, read the buffer's declared format (via PyBUF_FORMAT | PyBUF_STRIDES) and map it to the match",https://github.com/pytorch/pytorch/pull/188784,26406626b686eac4b60a632a183893db9924eb15e12cf357d08885fadbf24c00 review guidance,pr,188784,issue,188128,high,pr.reviews[0].body,"Thanks @SwayamInSync, this looks quite good and fixes multiple bugs. The implementation looks solid; I flagged a couple of docs things to polish. The order of protocols/methods to use to attempt to convert the data also matches torch.tensor better with __dlpack__ added, which is what's expected....",https://github.com/pytorch/pytorch/pull/188784,7540f6f3386ae2c064982529eac9f550db4b0ed7f81f04e74a196a0f2dd11a7b references,pr,189263,pr,180247,medium,pr.comments[0].body,".8+. AI verdict: The manywheel-py3_15-rocm7_2-build failure is a pre-existing issue. The same job fails identically on other unrelated PRs (#180247, #180250) run the same day, proving this Cython/Python 3.15 incompatibility predates the suspect commit. Full reasoning on HUD →...",https://github.com/pytorch/pytorch/pull/189263,95e0d885e3f0c86c843e0bf5b1c465f3ea1d2065fa019dda2eb587971d94ca2c references,pr,189263,pr,180250,medium,pr.comments[0].body,"erdict: The manywheel-py3_15-rocm7_2-build failure is a pre-existing issue. The same job fails identically on other unrelated PRs (#180247, #180250) run the same day, proving this Cython/Python 3.15 incompatibility predates the suspect commit. Full reasoning on HUD → linux-bin...",https://github.com/pytorch/pytorch/pull/189263,cb1f1e23ab312125a12a49c6da661fc2b1999e2b9ba26cca6c0e68bdf2f85118 review guidance,pr,181926,pr,3175,high,pr.reviews[0].body,This PR requires intel/torch-xpu-ops#3534 but that would need to land in pytorch repository first.,https://github.com/pytorch/pytorch/pull/181926,f0338bf3a68f40936e2a8a79719618b2b583c6696c8ebaf613b9b0f44a5386b5 review guidance,pr,179425,pr,179425,high,pr.reviews[0].body,Fix lints and add typing. How are we evaluating this?,https://github.com/pytorch/pytorch/pull/179425,1d0d55a3c7d0920978526b6e86d7d00e57c06ced70062d85b1ad4a8879914af2 review guidance,pr,189160,pr,189160,high,pr.reviews[1].body,"Pull request overview This PR refactors autograd’s Python saved-tensor hook and anomaly-metadata implementations to use c10::SafePyObject instead of raw PyObject*, routing lifetime management through PyInterpreter::{incref,decref} and removing bespoke GIL-acquiring destructors. It also changes th...",https://github.com/pytorch/pytorch/pull/189160,5f8059df4ae747074f4dcac3d80412b05ae702961ad981ff5262b4652c3e45b1 closes,pr,189127,issue,187093,high,pr.closingIssuesReferences,pr #189127 declares a closing reference to issue #187093.,https://github.com/pytorch/pytorch/pull/189127,780847842265abccdd68269e066b439011f231d80afd03754aefb246cead60c7 closes,pr,189127,issue,187093,high,pr.body,Fixes #187093. Summary Extend the existing addmm bias-unfuse post-grad pattern to broadcast-bias baddbmm. Only rewrite on GPU when the baddbmm input is b,https://github.com/pytorch/pytorch/pull/189127,498e5c94efe177603078e2a1d3b438d56af13d5a0a68e8891292b13c060604be review guidance,pr,189127,issue,187093,high,pr.reviews[0].body,Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.,https://github.com/pytorch/pytorch/pull/189127,7ddc7025c34d6505184c2fd1702b6b41037fca3e28b1bb5aafc8b2a6a7291c8d closes,pr,189149,issue,186348,high,pr.closingIssuesReferences,pr #189149 declares a closing reference to issue #186348.,https://github.com/pytorch/pytorch/pull/189149,2c9285b3c98631b5be95ae4d0347933120619af424ec633b7b501b320911aa1b closes,pr,189149,issue,186348,high,pr.body,"Fixes #186348 Summary When both inner dimensions of aten.mm are small (K < 8 and N < 8), the cuBLAS/Triton GEMM kernel launch overhead dominates executio",https://github.com/pytorch/pytorch/pull/189149,b9e7da3fc9181b929397020a377afc2ca03a5c59d6526c20b258208c65b86efa references,pr,189149,issue,188721,medium,pr.comments[0].body,"akiness on trunk: B200 Smoke Tests / linux-jammy-cuda13.0-py3.12-gcc11-sm100 / test-osdc (smoke_b200, 1, 1, mt-l-x86iamx-22-225-b200) (gh) (#188721) test/inductor/test_flex_flash.py::TestFlexFlashCUDA::test_captured_table_int64_index_cuda This comment was automatically generat...",https://github.com/pytorch/pytorch/pull/189149,4dc2d72db280b6ae3fdafe9e3575f99c40f03997c81ff561fd06a7e0671406c1 review guidance,pr,182278,pr,182278,high,pr.reviews[0].body,"There are some changes clearly unrelated to RISCV, and some of them create a differentiation across the builds without clear explanation why this is necessary (for example, build OpenBLAS from source for aarch64, but download it from deb for RISCV)",https://github.com/pytorch/pytorch/pull/182278,3975efccac3b46091259f01c16b8299719509f6776f4b9940a57df6f03d9eb1b closes,pr,189124,issue,189121,high,pr.closingIssuesReferences,pr #189124 declares a closing reference to issue #189121.,https://github.com/pytorch/pytorch/pull/189124,d4f9a1c69db12dce5887e4fac8bfd954999ad06507fd88a41930f5b710d1dcfe closes,pr,189124,issue,189121,high,pr.body,Fixes #189121. Summary Extract the existing dynamic-rblock eligibility predicate so the cache gate can see when a single original reduction config can ex,https://github.com/pytorch/pytorch/pull/189124,3b5ef5d2721b8aca455547e05b4ad014d01e73cc38b80ddcda2e5344974c5a4b review guidance,pr,189124,issue,189121,high,pr.reviews[0].body,Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.,https://github.com/pytorch/pytorch/pull/189124,e0ef24998fd5667fe65518e3ff92c48d57c13b144c07014b97992b1ce4c22401 review guidance,pr,185794,pr,185794,high,pr.reviews[0].body,"💡 Codex Review Here are some automated review suggestions for this pull request. Reviewed commit: f3697749e9 ℹ️ About Codex in GitHub Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you Open a pull request for review Mark a draft as ready Comment ""@code...",https://github.com/pytorch/pytorch/pull/185794,ac3cc8721e6b2095f05d5a7eddcedaaa6f45a7425b9025033339718a90c2c68f closes,pr,189145,issue,188405,high,pr.closingIssuesReferences,pr #189145 declares a closing reference to issue #188405.,https://github.com/pytorch/pytorch/pull/189145,d04ec763b5af91db7870f3e8f5bf497fbe71a02de3f4af8bf691a7bc767de742 closes,pr,189145,issue,188325,high,pr.closingIssuesReferences,pr #189145 declares a closing reference to issue #188325.,https://github.com/pytorch/pytorch/pull/189145,1eb69c6adeb2e818033169966a9290b65dadb4bca4095478620312d96eb37e31 closes,pr,189145,issue,188325,high,pr.body,"Related Issue Fixes #188325 Fixes #188405 Why EventVariable.python_type() returns a hardcoded torch.Event, and graph-break reconstruction rebuilds graph-created stream",https://github.com/pytorch/pytorch/pull/189145,d39d2956270488cb86ddea765eb37e3055e132dfac82958454c935ff5b19db3f closes,pr,189145,issue,188405,high,pr.body,"Related Issue Fixes #188325 Fixes #188405 Why EventVariable.python_type() returns a hardcoded torch.Event, and graph-break reconstruction rebuilds graph-created streams/events via b",https://github.com/pytorch/pytorch/pull/189145,0e9b47b817c79f80c368014408c9d381f84dbbfc3ab65821b32c6bd1a0d8119e references,pr,189186,pr,188268,medium,pr.body,"after #188268 is merged, this extends that fix to cutlass -- do not merge before #188268 cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @Xi",https://github.com/pytorch/pytorch/pull/189186,bca4bdeab2c2979dedc1cb52b75be8fb548161537427b9855487b4defbf76b4d references,pr,189186,issue,188721,medium,pr.comments[0].body,"akiness on trunk: B200 Smoke Tests / linux-jammy-cuda13.0-py3.12-gcc11-sm100 / test-osdc (smoke_b200, 1, 1, mt-l-x86iamx-22-225-b200) (gh) (#188721) test/inductor/test_flex_flash.py::TestFlexFlashCUDA::test_cutedsl_captured_alias_views_keep_distinct_layouts_case_offset_cuda Th...",https://github.com/pytorch/pytorch/pull/189186,7dd4372c71c4f9bb06d7bca99a570ee116e75adeab5cabc140b0842ea27ce5bd review guidance,pr,189247,pr,189247,high,pr.reviews[0].body,will let the XPU experts do the final merge,https://github.com/pytorch/pytorch/pull/189247,44faa3e9528e1a3c3de5fc83cf08ea53660b2051b148d552ffbaa41c85546998 references,pr,186056,pr,186055,medium,pr.body,Stack from ghstack (oldest at bottom): -> #186056 #186055 This allows us to conditionally run an optimizer. This enables us to do a gradient nan check before running the optimizer without having to,https://github.com/pytorch/pytorch/pull/186056,10d4d9f750fd9ff9376e638276c049a7126e1077023d3f80672c39cdd9d3cf89 review guidance,pr,189197,pr,189197,high,pr.reviews[0].body,"💡 Codex Review Here are some automated review suggestions for this pull request. Reviewed commit: ee7405910a ℹ️ About Codex in GitHub Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you Open a pull request for review Mark a draft as ready Comment ""@code...",https://github.com/pytorch/pytorch/pull/189197,aa615c6faf6a926c23d85259b09a64f3f3bd523fd8f6aa1150aaa10366bb210f references,pr,189146,issue,187455,medium,pr.body,"/ -44, their dispatch + removing core's last-dim-only gate). What / why Follow-up to the softmax core PR (#187456), stacked on it (Part of #187455). Core migrated last-dim softmax to native Metal and routed non-last-dim axes and very-long last-dim rows to the retained MPSGraph...",https://github.com/pytorch/pytorch/pull/189146,586964e4c83352d281ec7a8e2f282c3e0925b45ad329f25c26480a1870bb8020 references,pr,189146,pr,187456,medium,pr.body,"in — external contributors can't ghstack), so the GitHub diff shows the softmax core PR + this one. To see only what this PR adds on top of #187456, use this fork compare (renders as a normal diff of just this delta): anagnorisis2peripeteia/pytorch@mps-softmax-pr1...mps-softma...",https://github.com/pytorch/pytorch/pull/189146,288aafc334192b64c534d0c0cb6fc22a12906127acc10c70522cdfece57009c0 references,pr,189162,issue,189161,medium,pr.body,"Issue: #189161 What: Add an ""fmax"" reduction type (non-NaN-propagating max) and use it in the softmax two-pass fallback path (prepare_softmax_twopass_fall",https://github.com/pytorch/pytorch/pull/189162,a42d6baf5c893efdecc61db1774458ee1ffa1d0283e24d25a176a02be62c1844 references,pr,188681,issue,188672,medium,pr.body,For #188672 The main motivation is to use the most current compatible CUDA libdevice.10.bc available on the system. Triton bundles its own copy of libd,https://github.com/pytorch/pytorch/pull/188681,a8945a636a8959d7ce2853d2f1364405d6eadc135c0412f2f13339356d12961c closes,pr,189174,issue,165097,high,pr.closingIssuesReferences,pr #189174 declares a closing reference to issue #165097.,https://github.com/pytorch/pytorch/pull/189174,c1ef28523eff95f62816a9f2a0eaa75f5881c79a12c24e9c695d03f0644da8f0 closes,pr,189174,issue,165097,high,pr.body,Fixes: #165097 Summary: Replace hardcoded CUDA/PrivateUse1 availability checks in _write_files_from_queue with generic _get_available_device_type() and a,https://github.com/pytorch/pytorch/pull/189174,54e2421e7a57e13ce87dc12f9dab60d41b5a2e8e68290a2e8a9b75f66a4b7c75 closes,pr,189234,issue,187806,high,pr.closingIssuesReferences,pr #189234 declares a closing reference to issue #187806.,https://github.com/pytorch/pytorch/pull/189234,f11171119df32f49b3d53ef4b6593cf0738e1ecc46d1b24d78bf74ef4c7be21a closes,pr,189234,issue,187806,high,pr.body,"Fixes #187806. Stacked on #189291 (accurate Metal erfc): the shared erfc files sit in both PRs until it lands, then this one rebases and drops them. Root",https://github.com/pytorch/pytorch/pull/189234,dc86398ea7776bea8c3f2f0d64a943f6938aecb5b0f3c4f9012a7f3fe3bd6612 references,pr,189234,pr,189291,medium,pr.body,"Fixes #187806. Stacked on #189291 (accurate Metal erfc): the shared erfc files sit in both PRs until it lands, then this one rebases and drops them. Root cause Exact gelu co",https://github.com/pytorch/pytorch/pull/189234,87556067ad8f8cc32c0192a056e671f7917adebdf88f1f24ccc5d2329bac843d references,pr,186540,issue,141287,medium,pr.body,"_ops.py -k ""segment_reduce and mps"" -v — 10 pass, 0 fail lintrunner — clean Performance benchmark (numbers in description above) References #141287 (MPS operator coverage tracker). References #173292 (closed CPU-fallback attempt). cc @kulinseth @malfet @DenisVieriu97 @jhavukai...",https://github.com/pytorch/pytorch/pull/186540,683424f36c7d0e0e22931c66eda0915f935583f6f6b7ae6a4eb60360fc6ee40a competes with,pr,180247,pr,180248,medium,pr.body,"Stack from ghstack (oldest at bottom): #180250 #180248 -> #180247 Switch pyproject.toml's build backend from setuptools to scikit-build-core (>= 1.0), which invokes CMake directly instead of hav",https://github.com/pytorch/pytorch/pull/180247,aace792942d5e3393396bd500472f5f958af993a17e683fbc735cbad557d9760 competes with,pr,180247,pr,180250,medium,pr.body,"Stack from ghstack (oldest at bottom): #180250 #180248 -> #180247 Switch pyproject.toml's build backend from setuptools to scikit-build-core (>= 1.0), which invokes CMake directly instea",https://github.com/pytorch/pytorch/pull/180247,e7e3ccc937e0ae3610346d7dba82a5368ddbdf662091042868a41e4eb8343353 references,pr,180247,issue,119526,medium,pr.comments[0].body,"c / linux-jammy-cuda13.0-py3-gcc11-slow-gradcheck / test (default, 8, 8, mt-l-x86aavx2-29-113-a10g, module:slowgradcheck) (gh) (disabled by #119526) profiler/test_profiler.py::TestProfiler::test_source_multithreaded_multiple_preexisting_work_in_main_thread_False xpu / linux-no...",https://github.com/pytorch/pytorch/pull/180247,7089545dbee6b37eaa4c8081c9afc351b01b97bc0745914a90d3b54a120c8d93 references,pr,180247,issue,188898,medium,pr.comments[0].body,"c / linux-jammy-cuda13.0-py3-gcc11-slow-gradcheck / test (default, 6, 8, mt-l-x86aavx2-29-113-a10g, module:slowgradcheck) (gh) (disabled by #188898) inductor/test_compiled_autograd.py::FuncTorchHigherOrderOpTestsWithCompiledAutograd::test_vmap_over_vmap_two_inputs periodic / l...",https://github.com/pytorch/pytorch/pull/180247,38c5ed3a3a5ea4a1dc832f6b4c4388f903270bdadea970b53b14ce6b9d480066 review guidance,pr,185251,pr,185251,high,pr.reviews[0].body,"Thanks for the updates! I see a bunch of usages of @onlyCUDA still. For cublas tests, it makes sense to keep these, but for non-cublas tests, I'd expect @onlyAccelerator instead. Wdyt? Oh also I realize it's a more invasive change, but the filename itself still references cuda. I wonder if it mak...",https://github.com/pytorch/pytorch/pull/185251,b82a610e72503d123058711391818323e6c3221b0c6f6009b93b5205c5754fbc review guidance,pr,184279,pr,136012,high,pr.reviews[0].body,"Looks good but, can we verify things like the assert_size_stride, assert alignment, etc that normally get generated remain ? in the test. Can you also compare with the other attempt at fixing this, just to see if theres anything we missed ?",https://github.com/pytorch/pytorch/pull/184279,1aceefc2f660d9fcd6e5d9c0a06511558165cb5ea02d6b4eaa08b349da79e4d3 review guidance,pr,184279,pr,184279,high,pr.reviews[0].body,"Looks good but, can we verify things like the assert_size_stride, assert alignment, etc that normally get generated remain ? in the test. Can you also compare with the other attempt at fixing this, just to see if theres anything we missed ?",https://github.com/pytorch/pytorch/pull/184279,83bc1b05798e61ae3871e74a7bfbb6f3ffd347b726594cde98b7e60475ebac79 closes,pr,186349,issue,132635,high,pr.body,"ects outside compile. Add coverage for custom logger methods and the negative wrapper-method case, and sync the graph-break registry. Fixes #132635 Generated by my agent Test Plan: python test/dynamo/test_reorder_logs.py -k test_ignore_logger python test/dynamo/test_error_mess...",https://github.com/pytorch/pytorch/pull/186349,042a0a662eefd12c8dd0e2de6cee0ae799bbbbdcf18d4852cd6ea3cb43529fd6 closes,pr,186355,issue,131794,high,pr.body,"; pruning the already traced backward graph keeps the change local to lazy backward lowering while preserving a correctness fallback. Fixes #131794 Generated by my agent Benchmark Results Microbenchmark: @torch.compile(backend=""aot_eager"") function returning x.sin(), y.sin(),...",https://github.com/pytorch/pytorch/pull/186355,009268e3f115d18ece86dc9c20ed44e841d65892dd496b5f90e51a3d38dbb04a closes,pr,184121,issue,149982,high,pr.body,192 0.12067 0.06336 0.05747 4096 16384 0.12160 0.06246 0.05725 16384 4096 0.12067 0.06352 0.06349 16384 16384 0.46586 0.24560 0.22838 Fixes #149982 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/184121,922610669dba37ca5786f46db07de6cc41de875da59dcce9997db4273d662556 review guidance,pr,184121,issue,149982,high,pr.reviews[0].body,This is changing the defaults to improve performance. I think this change needs benchmark results of some kind. It looks correct to me but I have no way of telling if it actually improves performance.,https://github.com/pytorch/pytorch/pull/184121,b94086ffaa9f8d52d8634f840d027f5a817910ae934164271c8a8aa9c59c42fd review guidance,pr,184121,pr,184121,high,pr.reviews[0].body,This is changing the defaults to improve performance. I think this change needs benchmark results of some kind. It looks correct to me but I have no way of telling if it actually improves performance.,https://github.com/pytorch/pytorch/pull/184121,dea0c26eb18f2912ac1f646363d54fc3a9629bb8e11a83fbca192e84d8d81ab3 review guidance,pr,184121,issue,149982,high,pr.reviews[1].body,needs better perf validation before landing outside of local testing,https://github.com/pytorch/pytorch/pull/184121,80050b3d51c024590b993b1dabcf014c018ad2eaaadfbe179d8f313311336f8e review guidance,pr,184121,pr,184121,high,pr.reviews[1].body,needs better perf validation before landing outside of local testing,https://github.com/pytorch/pytorch/pull/184121,b4e50c9184a0462c1ab95fde6408e88f75fb87655292219820dd9caf8747ca5d review guidance,pr,184121,issue,149982,high,pr.reviews[2].body,This is changing the defaults to improve performance. I think this change needs benchmark results of some kind. It looks correct to me but I have no way of telling if it actually improves performance.,https://github.com/pytorch/pytorch/pull/184121#pullrequestreview-4319731952,098dd0c4fecdaa69e8a8a00ad90957c101070d3da1706c960e6144212a349039 review guidance,pr,184121,pr,184121,high,pr.reviews[2].body,This is changing the defaults to improve performance. I think this change needs benchmark results of some kind. It looks correct to me but I have no way of telling if it actually improves performance.,https://github.com/pytorch/pytorch/pull/184121#pullrequestreview-4319731952,7c1bac395fc2f3b6ea8d57ca4dbac10d9aa39199bcbf4a2bd86f0c2853458ea1 review guidance,pr,184121,issue,149982,high,pr.reviews[3].body,needs better perf validation before landing outside of local testing,https://github.com/pytorch/pytorch/pull/184121#pullrequestreview-4377383396,c79c2d54529a33cf770111e86fbad908c37787cb95fbc56f388d3d43d6d7fbfa review guidance,pr,184121,pr,184121,high,pr.reviews[3].body,needs better perf validation before landing outside of local testing,https://github.com/pytorch/pytorch/pull/184121#pullrequestreview-4377383396,fd2d0c07b9a988c10d9631776fcc542af477b4df6471497f051c4c02c598f505 review guidance,pr,184121,issue,149982,high,pr.reviews[4].body,"defer on heuristic change until validation. also, can you please, look at NCU ? i believe even with coordinate descent tuning there were terrible bank conflicts interfering with perf.",https://github.com/pytorch/pytorch/pull/184121#pullrequestreview-4479173908,6cb05c1490be09d6283d6befe0ce7da5b0ff595f0d9b362d675a498cecedd39c review guidance,pr,184121,pr,184121,high,pr.reviews[4].body,"defer on heuristic change until validation. also, can you please, look at NCU ? i believe even with coordinate descent tuning there were terrible bank conflicts interfering with perf.",https://github.com/pytorch/pytorch/pull/184121#pullrequestreview-4479173908,a0cd382b07ede235da17b946f0b1ae148aff4e02bab1084f7ca9ff5917d2caa7 closes,pr,183884,issue,186136,high,pr.body,"and job messages, preventing old executor threads from lingering across repeated quiesce cycles and blocking shutdown. Fixes #176968 Fixes #186136 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...",https://github.com/pytorch/pytorch/pull/183884,099b4dad42be3c8d84cfd013bf3818aa20ab9e5246b0f05c8a3ee2132c0e4039 references,pr,183884,issue,186031,medium,pr.comments[10].body,"I rebased and resubmitted onto current main to refresh CI after Dr. CI reported a Torchtitan torchcomms failure that matches existing issue #186031, while the inductor failures were classified as broken trunk. This PR only touches compile-worker code, so I removed the unrelate...",https://github.com/pytorch/pytorch/pull/183884,74476b9abaed0e19bf6315f6cbbf135c5764dcc393c797327d6a8e36dd5cb906 review guidance,pr,183884,pr,176968,high,pr.reviews[0].body,asking @c00w to review,https://github.com/pytorch/pytorch/pull/183884,84a39e043e1422257126fa28f2452ad5aa333d4e9a589bba9e9fe3a4e71143f6 review guidance,pr,183884,pr,183884,high,pr.reviews[0].body,asking @c00w to review,https://github.com/pytorch/pytorch/pull/183884,d07d82a3c86965b60600c136deb2d5de39d56b761920828dc0bb21e2079f695f review guidance,pr,183884,issue,186136,high,pr.reviews[0].body,asking @c00w to review,https://github.com/pytorch/pytorch/pull/183884,eae0afdccb15dce1ce45dd60596229ab6229b29d0f0ae6f62118c5dc41aa019c review guidance,pr,183884,pr,176968,high,pr.reviews[2].body,asking @c00w to review,https://github.com/pytorch/pytorch/pull/183884#pullrequestreview-4478966524,cf51aae369ce43a08c0cf4fb586747f97dc742188f1508c2cc4ed045d3ea7647 review guidance,pr,183884,pr,183884,high,pr.reviews[2].body,asking @c00w to review,https://github.com/pytorch/pytorch/pull/183884#pullrequestreview-4478966524,b6b4c23ff75ab52b7c472568f3e861e0d81ce73d8582cff3919754e6b224f142 review guidance,pr,183884,issue,186136,high,pr.reviews[2].body,asking @c00w to review,https://github.com/pytorch/pytorch/pull/183884#pullrequestreview-4478966524,e0495f902ed0eef98a52fe144b16904ec5bdb60435de8d679def50cfdc87e743 review guidance,pr,183884,pr,176968,high,pr.reviews[4].body,"Core change seems fine (I would probably just kill the whole process tree instead of trying to do it cleanly within the process on shutdown timeout). However, it's interesting to me that we're getting compile jobs queued that are running long enough to block shutdown, do we know why they're takin...",https://github.com/pytorch/pytorch/pull/183884#pullrequestreview-4479508927,66bff33425a1e1b769b9d44a5c64f6c63c7f52f8f753d4be20a77d882f66baae review guidance,pr,183884,pr,183884,high,pr.reviews[4].body,"Core change seems fine (I would probably just kill the whole process tree instead of trying to do it cleanly within the process on shutdown timeout). However, it's interesting to me that we're getting compile jobs queued that are running long enough to block shutdown, do we know why they're takin...",https://github.com/pytorch/pytorch/pull/183884#pullrequestreview-4479508927,e264edf7f4cd92eb1b581990dc0019d308e6e6b610e8855e01361a56339992f6 review guidance,pr,183884,issue,186136,high,pr.reviews[4].body,"Core change seems fine (I would probably just kill the whole process tree instead of trying to do it cleanly within the process on shutdown timeout). However, it's interesting to me that we're getting compile jobs queued that are running long enough to block shutdown, do we know why they're takin...",https://github.com/pytorch/pytorch/pull/183884#pullrequestreview-4479508927,302c9a27c192c34a733e6329ad796f9c732daed71fbf1582a1399bcffdb307d6 closes,pr,188795,issue,150199,high,pr.closingIssuesReferences,pr #188795 declares a closing reference to issue #150199.,https://github.com/pytorch/pytorch/pull/188795,0f3b058b423b3ebef15736f599508a94b5d22198672b918b5c23c6f03bcb983f closes,pr,188795,issue,150199,high,pr.body,Fixes #150199 torch.asarray preserves input's device when a default device is set except for Python scalars and sequences. _asarray_input_has_device help,https://github.com/pytorch/pytorch/pull/188795,c278e48a5bf5ba702db3d00c43c253e23aa228ea680d23c375f4f2034aadaac2 closes,pr,186364,issue,129936,high,pr.body,"lizes only the specific first-dimension symbol it intentionally guards on, so guards on unrelated symbolic dimensions remain visible. Fixes #129936 Generated by my agent Test Plan: python -m pytest test/dynamo/test_subclasses.py::SubclassTests::test_recompile_with_symbool_inpu...",https://github.com/pytorch/pytorch/pull/186364,956d768279b3536e35f45b2a3aaa57ee41531c1f0b262ab70659412521586dc7 closes,pr,186372,issue,129648,high,pr.body,"ices with eager-compatible errors, and graph-break for non-CPU default-device mode rather than cloning NumPy producer nodes unsafely. Fixes #129648 Generated by my agent Test Plan: python test/dynamo/test_repros.py ReproTests.test_tensor_ctor_nested_numpy_ndarrays ReproTests.t...",https://github.com/pytorch/pytorch/pull/186372,57e1002e77746115a49a447a52c6f105c42d0a326bdfa5bb93e62fea87183add closes,pr,186406,issue,128830,high,pr.body,"ave_config_ignore, and non-importable custom callables cannot be faithfully reconstructed without an explicit serialization contract. Fixes #128830 Generated by my agent Test Plan: python test/test_utils_config_module.py TestConfigModule.test_codegen_config_function TestConfig...",https://github.com/pytorch/pytorch/pull/186406,b99e128e3f47ad1fc910f69150cadc4b9e200c912ad35e9faa9bb5ffd1e04432 references,pr,186854,issue,176662,medium,pr.body,std::endian is available since C++20 #176662,https://github.com/pytorch/pytorch/pull/186854,3d8f0339e6df1567219c6a1ce1a9851b2a657e93087c0f02081ad0d4822e66c3 references,pr,188440,issue,176662,medium,pr.body,#176662 cc @jgong5 @mingfeima @XiaobingSuper @sanchitintel @ashokei @jingxu10 @aditew01,https://github.com/pytorch/pytorch/pull/188440,37b9f47372d6066a089a6eeb96964c5008a60dc1428b633d5984d707b5292d42 review guidance,pr,188440,issue,176662,high,pr.reviews[1].body,Looks good to me. Can we validate on Android with old toolchains too?,https://github.com/pytorch/pytorch/pull/188440,9bb7a03634c8b9aaf1b7b9e2c7968fcc04dcce14bbdde81ac52d75eba1461262 closes,pr,186415,issue,128556,high,pr.body,"e inside a fullgraph compiled function with a tensor constant stored as a plain attribute, both with and without a registered buffer. Fixes #128556 Generated by my agent Test Plan: Manual issue repro before fix: reproduced both with-buffer and without-buffer failures. Manual i...",https://github.com/pytorch/pytorch/pull/186415,d919a99f00fce1e44ab3102b69b54d66843021fd31f2e60c7b62a5f11248768c references,pr,186415,issue,128556,medium,pr.comments[2].body,e constructors legitimately create Parameters and should be exempted from the sourceless-Parameter graph break — is well-motivated by issue #128556. The scoping (only Module subclasses get the allowance) avoids over-permitting non-module classes. LGTM with the optional suggest...,https://github.com/pytorch/pytorch/pull/186415,3386c2a7f0601801a31edd1599b621ec429f0333c892f34b0e06d2ce5ef58df7 references,pr,186415,issue,128556,medium,pr.comments[6].body,"_buffer` / `_without_buffer`) in `test_modules.py` properly test the fullgraph path with `CompileCounter`, covering the exact scenario from #128556. The `test_nn_parameter_ctor_graph_breaks` test in `test_repros.py` is unchanged and continues to validate that direct sourceless...",https://github.com/pytorch/pytorch/pull/186415,cc4f9fd25e2a698fa05eb0671295d852cc18e276f859026425779ae495225c06 closes,pr,186418,issue,128554,high,pr.body,focused regression test that saves and loads an aot_export_module GraphModule and checks the loaded module computes the same result. Fixes #128554 Generated by my agent Test Plan: python test/functorch/test_aotdispatch.py TestAOTExport.test_aot_export_module_graph_module_seria...,https://github.com/pytorch/pytorch/pull/186418,5da41d1a7ad9d26e5b75786a82aae0d5dc6fb0ebdb6779ec22ed5113cf2ce3dd closes,pr,186403,issue,129131,high,pr.body,the same process and the compiler should not mutate accelerator state unless it is compiling for an already-initialized accelerator. Fixes #129131 Generated by my agent Test Plan: python test/dynamo/test_repros.py CUDAReproTests.test_cpu_compile_does_not_initialize_cuda python...,https://github.com/pytorch/pytorch/pull/186403,d092f175791b4183be87196864e65986651547a51eb4598e6dbfe2ca117b5298 references,pr,189076,pr,184114,medium,pr.body,codegen paths can read it directly. Generated with assistance from Claude. The cpp-wrapper side assert_alignment fix needs to be done after #184114 lands. Differential Revision: D110885058,https://github.com/pytorch/pytorch/pull/189076,f8c1f23429c7671a379441cbf84e2d6ffc14b92d1e55dae8034a925de1ed45cb closes,pr,189273,issue,189267,high,pr.closingIssuesReferences,pr #189273 declares a closing reference to issue #189267.,https://github.com/pytorch/pytorch/pull/189273,fed9694831af3109ce7b48cd4d6a5f70baa7285d3a4e359e90c96ff5df48f3cd closes,pr,189273,issue,189267,high,pr.body,Fixes #189267,https://github.com/pytorch/pytorch/pull/189273,c0cfe81d4a4f0e6d027bd83db0984065623d00428422563d5c347c57b40c95b4 closes,pr,189273,issue,189267,high,pr.comments[1].body,See #189267 (comment),https://github.com/pytorch/pytorch/pull/189273,afb013de2d9bd407de02ca4ab54ad963ae083d8da723b852fc0934e4d928252f references,pr,180250,pr,180247,medium,pr.body,"Stack from ghstack (oldest at bottom): -> #180250 #180248 #180247 With setup.py removed (#180248), workflows and habits that still invoke python setup.py {install,develop,bdist_wheel,clean} would simply fa",https://github.com/pytorch/pytorch/pull/180250,d2b6c6fb4c2dcc6fc9253e2d471b0770aec65cc17c12b177879992390793b1e0 references,pr,180250,pr,180248,medium,pr.body,"Stack from ghstack (oldest at bottom): -> #180250 #180248 #180247 With setup.py removed (#180248), workflows and habits that still invoke python setup.py {install,develop,bdist_wheel,clean} would s",https://github.com/pytorch/pytorch/pull/180250,61b41deb9f7fcc726402f34948d9f93a65afe0938377db30385df9b28a2f20bd references,pr,180248,pr,180247,medium,pr.body,"Stack from ghstack (oldest at bottom): #180250 -> #180248 #180247 With the build fully driven by scikit-build-core (#180247), the setuptools path is dead code. Remove setup.py and its helpers (build_pytorc",https://github.com/pytorch/pytorch/pull/180248,801e099f9c6745b314a688fca714ff52d3b5daf9e98218b0639cf5bf1b0158f7 references,pr,180248,pr,180250,medium,pr.body,"Stack from ghstack (oldest at bottom): #180250 -> #180248 #180247 With the build fully driven by scikit-build-core (#180247), the setuptools path is dead code. Remove setup.py and its he",https://github.com/pytorch/pytorch/pull/180248,fe401e6006100501e2c9f3c3ec11fc957a22f444ebe0e233d025751d7ee436a0 closes,pr,186849,issue,176178,high,pr.body,e key. Preserve exact initial tensor strides for non-compact fake tensors so padded layouts compile against the same runtime strides. Fixes #176178 Generated by my agent Test Plan: timeout 900s env CUDA_VISIBLE_DEVICES=0 TORCHINDUCTOR_CACHE_DIR=/data/users/jansel/pytorch-issue...,https://github.com/pytorch/pytorch/pull/186849,48efa2ef10cc320143cc583a93c7c1e9862a03e8f69f95132a0dc1a0ec67544c closes,pr,185150,issue,161807,high,pr.body,"ules, but that would only hide the source bug: the GraphModule produced by FX tracing was already returning the wrong container type. Fixes #161807 Generated by my agent Test Plan: python test/test_fx.py TestFX.test_symbolic_trace_preserves_ordered_dict_output (blocked before...",https://github.com/pytorch/pytorch/pull/185150,ddf6e13d96f5d953a1bad546db255bdccda0bb803e32633405d39f5e4a8f67a3 review guidance,pr,185150,issue,161807,high,pr.reviews[0].body,"Agent: fx graph should not have intermediate nodes that return OrderedDict. I don't know if this PR is salvageable, if it isn't then please close it",https://github.com/pytorch/pytorch/pull/185150,34c25d4907f9905d5d85ab6f11732d863333fa15613c092f225ea39e42948b1a review guidance,pr,185150,pr,185150,high,pr.reviews[0].body,"Agent: fx graph should not have intermediate nodes that return OrderedDict. I don't know if this PR is salvageable, if it isn't then please close it",https://github.com/pytorch/pytorch/pull/185150,991ba00c3e20c6881518bf5a2a1267df382630b790cbffde02e708e926c9fd6f review guidance,pr,185150,issue,161807,high,pr.reviews[1].body,"My agent says I addressed the review feedback by removing the intermediate call_function(collections.OrderedDict) reconstruction node. FX now carries OrderedDict outputs as an internal immutable ordered aggregate and reconstructs collections.OrderedDict only at the output boundary, including nest...",https://github.com/pytorch/pytorch/pull/185150,528ffdbb09ed3762de148ffb328b065c4271c7fefea401598349b9990931468c review guidance,pr,185150,pr,185150,high,pr.reviews[1].body,"My agent says I addressed the review feedback by removing the intermediate call_function(collections.OrderedDict) reconstruction node. FX now carries OrderedDict outputs as an internal immutable ordered aggregate and reconstructs collections.OrderedDict only at the output boundary, including nest...",https://github.com/pytorch/pytorch/pull/185150,74e5b2fca648d827fc408850631d85c2d2308da1e67d878c4bd22dd47639150a closes,pr,186425,issue,128160,high,pr.body,"oo. Tests cover direct hooks, assigned hook functions, nested helper calls, and hooks whose code object comes from a skiplisted file. Fixes #128160 Generated by my agent Benchmark Results: Command: repeated 40-iteration torch.compile(..., backend=""eager"") loop, 7 reps. Main ba...",https://github.com/pytorch/pytorch/pull/186425,2d1decf6a2a63f3cf9470b4e4bdf030e3e07f4d56373a337b8fdb7aed508fcd7 closes,pr,184319,issue,128063,high,pr.body,on. This lets the first split-reduction level share the pointwise tiling instead of blocking fusion on incompatible flat split sizes. Fixes #128063 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/184319,37f5cd6736fc33c5e326c07afe75cd066dd9d4990d5e878245d7a6b3b896a5ba references,pr,184319,issue,128063,medium,pr.comments[8].body,"is a nice complement. --- ### On benchmarks The trigger was `review these changes`, so I focused on the code. I'm not able to run the issue #128063 GPU benchmark from this environment, so I can't post measured numbers here — that's better produced from a GPU CI run or local `d...",https://github.com/pytorch/pytorch/pull/184319,18a34639ca43e02bef41a451dff90d2a28585cc24d3037d609d4aa43ee0cc047 review guidance,pr,184319,issue,128063,high,pr.reviews[1].body,"I'd rather do this in scheduler, as was the original attempt of this. we can also use the cancel split reduction facility for this.",https://github.com/pytorch/pytorch/pull/184319,7edf5645e6e83563d006a47cda5e5cbc3e29fdf454cf9d4eb21a54aa28e46f4c review guidance,pr,184319,pr,184319,high,pr.reviews[1].body,"I'd rather do this in scheduler, as was the original attempt of this. we can also use the cancel split reduction facility for this.",https://github.com/pytorch/pytorch/pull/184319,58d3f57887ba4702f6343b9f4a7fcb81845b8a7f89f49c5680cd61b9383e89b2 review guidance,pr,184319,issue,128063,high,pr.reviews[3].body,"I'd rather do this in scheduler, as was the original attempt of this. we can also use the cancel split reduction facility for this.",https://github.com/pytorch/pytorch/pull/184319#pullrequestreview-4461362270,9341d026f8edeb7bea8a130e9eb132e94fc13a2218a5137310c797a8405749f0 review guidance,pr,184319,pr,184319,high,pr.reviews[3].body,"I'd rather do this in scheduler, as was the original attempt of this. we can also use the cancel split reduction facility for this.",https://github.com/pytorch/pytorch/pull/184319#pullrequestreview-4461362270,688fa08cc23cfa664969854367f35a74caab3a5372493f3147eec3a09cda94b3 review guidance,pr,184319,issue,128063,high,pr.reviews[4].body,see https://github.com/pytorch/pytorch/pull/141082,https://github.com/pytorch/pytorch/pull/184319#pullrequestreview-4461364379,2b9569d0dd732a489ffdf3fc532dc5bc783c356898b307b5577c08d1c5e96f68 review guidance,pr,184319,pr,184319,high,pr.reviews[4].body,see https://github.com/pytorch/pytorch/pull/141082,https://github.com/pytorch/pytorch/pull/184319#pullrequestreview-4461364379,f4a28a891f676ea327be8371fade1dd17eb180a1eede50fe1f4fc9a8a58b2bb8 closes,pr,186429,issue,127153,high,pr.body,"the exported graph. Add a regression test covering ONNX dynamo export for the issue's pad_sequence(torch.tensor_split(...)) pattern. Fixes #127153 Generated by my agent Benchmark Results: On a module returning torch._C._nn.pad_sequence([x, y, z], batch_first=True) with input s...",https://github.com/pytorch/pytorch/pull/186429,92e5b7c8ba49b1a2a0a24562750931afe7efa47a5334d44f7fbae0ec74892d75 closes,pr,186431,issue,127112,high,pr.body,hat contain FX GraphModule submodules. The narrower marker-based filter fixes the reported failure without hiding those user modules. Fixes #127112 Generated by my agent Test Plan: PYTORCH_TEST_WITH_DYNAMO=1 pytest -q test/test_module_tracker.py::TestModuleTracker::test_user_g...,https://github.com/pytorch/pytorch/pull/186431,86adb66588d37ee38793560fad2d50eafdeb6a71c59da2befaf5134ec61a3be5 closes,pr,186443,issue,126834,high,pr.body,"checks in the meta wrapper, but reusing random_from_to_impl keeps the eager and meta validation rules coupled to one implementation. Fixes #126834 Generated by my agent Test Plan: ninja -C build torch_python ninja -C build install python -m pytest test/test_fake_tensor.py::Fak...",https://github.com/pytorch/pytorch/pull/186443,1ea879bd6c1740afd9f887bc9911004240500c71b54edaf9ac6feefc4cc1cb93 review guidance,pr,188301,pr,188301,high,pr.reviews[0].body,LGTM.,https://github.com/pytorch/pytorch/pull/188301,448f744730d2af5d53d71f39baad736eddfb703e56ebfe3ca83b3a128d63a88f review guidance,pr,188301,pr,188301,high,pr.reviews[1].body,"@pkourdis , the failures should be related. Please check all these failures.",https://github.com/pytorch/pytorch/pull/188301,8f73281bd1a566c5a00ae5f37f9a3d63437d0e11c2b6b589b87a34906db975b6 closes,pr,184124,issue,148751,high,pr.body,ts at the earliest valid location after all replacement inputs and skips matches that would require reordering across external users. Fixes #148751 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/184124,4da0d513ae5b3458ed88f2d973eea18e6cab35e659bb3011c6bb77aa2093b701 review guidance,pr,184124,issue,148751,high,pr.reviews[0].body,"Two things: re-sorting the graph after every pattern is needless compile time overhead, and should only be opt in. I dont see why we dont just do a local, topo sort, of the affected region.",https://github.com/pytorch/pytorch/pull/184124,51a4f12dd1861bb9e5eb354648d7cf94796aa6774cb8c6016b7b8e310d884f55 review guidance,pr,184124,pr,184124,high,pr.reviews[0].body,"Two things: re-sorting the graph after every pattern is needless compile time overhead, and should only be opt in. I dont see why we dont just do a local, topo sort, of the affected region.",https://github.com/pytorch/pytorch/pull/184124,ace69f8a36eee582b66be1e68d19008d09402305c65d50db2f2d308548469e85 review guidance,pr,184124,issue,148751,high,pr.reviews[1].body,"Two things: - re-sorting the graph after every pattern is needless compile time overhead, and should only be opt in. - I dont see why we dont just do a local, topo sort, of the affected region.",https://github.com/pytorch/pytorch/pull/184124#pullrequestreview-4479160474,5ce51719a3dfd94410f98ba802042728d160d0c1245c28d14123df8ae696ce86 review guidance,pr,184124,pr,184124,high,pr.reviews[1].body,"Two things: - re-sorting the graph after every pattern is needless compile time overhead, and should only be opt in. - I dont see why we dont just do a local, topo sort, of the affected region.",https://github.com/pytorch/pytorch/pull/184124#pullrequestreview-4479160474,f790fa96b0203edd0873f6fb91253234513b4d9504d3fcce32a26d880d952c03 closes,pr,184343,issue,125277,high,pr.body,"mes during max-autotune, while preserving non-Hopper behavior and exhaustive search. Add tests for target-device capability handling. Fixes #125277 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...",https://github.com/pytorch/pytorch/pull/184343,285b88f12fd7bb088ce3cbe9eb66048aa117df31781dd50d022fc33e2cbf9517 closes,pr,186969,issue,185955,high,pr.closingIssuesReferences,pr #186969 declares a closing reference to issue #185955.,https://github.com/pytorch/pytorch/pull/186969,f735320ac9a65b0147bd2e3234a34fb81c5367fa37becd08091ec706fb91a750 supersedes,pr,186969,pr,185989,medium,pr.body,"-4.15.0 Notes This touches .ci/docker/, so it triggers a full CI Docker image rebuild (expected). Context: supersedes/continues the work in #185989. Authored with assistance from an AI agent. cc @justinchuby @malfet @pytorch/pytorch-dev-infra",https://github.com/pytorch/pytorch/pull/186969,d7d84043ad7808464791f7405eee3540cd1bf11db4fe4ec2fbb0d014f8314bf0 closes,pr,187171,issue,187170,high,pr.closingIssuesReferences,pr #187171 declares a closing reference to issue #187170.,https://github.com/pytorch/pytorch/pull/187171,402585a40db3bd336511039f0696dcf41cd3a276133728645b64b78da45e6510 closes,pr,187171,issue,187170,high,pr.body,"est_comprehensive_addbmm_cpu_float16 by comparing compiled f16 output against eager f16 output instead of an idealized f32 reference. Fixes #187170 Problem addbmm uses make_fallback in Inductor, meaning both compiled and eager paths call the same ATen kernel. The test was fail...",https://github.com/pytorch/pytorch/pull/187171,798e5c213455577388fe19463106c17467e4cb2561a1fe97bd6a6766496a4c64 review guidance,pr,187171,issue,187170,high,pr.reviews[1].body,"SGTM, thank you.",https://github.com/pytorch/pytorch/pull/187171,9475ea1ce8d0341528bd1cc2d1c2903b60af7b85d96a867161a00518ade0e1e0 references,pr,188105,pr,188181,medium,pr.body,Stack from ghstack (oldest at bottom): -> #188105 #188181 Authored with assistance from Claude (an AI assistant). Co-Authored-By: Claude Opus 4.8 (1M context) noreply@anthropic.com cc @voznesenskym,https://github.com/pytorch/pytorch/pull/188105,903955fb865510a2fa1ea396f9744d465d730b207807f74ad00760642fe771c5 review guidance,pr,188105,pr,188105,high,pr.reviews[1].body,Can we also add tests to test/dynamo?,https://github.com/pytorch/pytorch/pull/188105,ad1862e95ab4c565e7d29702e0c50452b869b82efac40725610c35bc78212887 review guidance,pr,188105,pr,188181,high,pr.reviews[1].body,Can we also add tests to test/dynamo?,https://github.com/pytorch/pytorch/pull/188105,ae9f5027de953a9febc05aa6c35808cac52f3324c408c23cd8b8608b6ecd02b2 closes,pr,184347,issue,125077,high,pr.body,D slices. Extend lazy compile metadata to carry R1/R2 block sizes and add block-pointer regressions for Python and cpp_wrapper paths. Fixes #125077 Generated by my agent cc @jgong5 @mingfeima @XiaobingSuper @sanchitintel @ashokei @jingxu10 @jerryzh168 @aditew01 @voznesenskym @...,https://github.com/pytorch/pytorch/pull/184347,92d77909d7a6285d953c70c16c0c8dbc10cf395e88b301602a047b5a278696e1 review guidance,pr,184347,issue,125077,high,pr.reviews[0].body,"For things like this that might be variable profitability, lets wait till we have better signal on microbenchmarks (which I am working on)",https://github.com/pytorch/pytorch/pull/184347,f41d768c953cb2b2a160604fa1323585e3f44a3a1adec829378495ae310c4f69 review guidance,pr,184347,pr,184347,high,pr.reviews[0].body,"For things like this that might be variable profitability, lets wait till we have better signal on microbenchmarks (which I am working on)",https://github.com/pytorch/pytorch/pull/184347,1b76f9ac8311329b140826cd131da08329729c70d8b527006ee50198136eb8e0 closes,pr,186454,issue,124863,high,pr.body,"ore removal, and warning during registration gives users a direct migration signal without changing dispatcher semantics immediately. Fixes #124863 Generated by my agent Test Plan: python test/test_custom_ops.py TestCustomOp.test_incorrect_schema_types TestCustomOp.test_is_ten...",https://github.com/pytorch/pytorch/pull/186454,54fe8042694a9f26a119c95171e1753d79beb7775ad40689fe7bf4ac80907fbf closes,pr,186455,issue,124747,high,pr.body,"Stack from ghstack (oldest at bottom): -> #186455 Issue #124747 failed when running the AOTAutograd memory leak test with PYTORCH_TEST_WITH_DYNAMO=1. There were two root causes. First, direct functorch a",https://github.com/pytorch/pytorch/pull/186455,be4a6a02d1b3ca8535d00d99246ea8aa5d5a414fb211f7021f9d0be86e8c54d4 closes,pr,186457,issue,124344,high,pr.body,"h.ops binding parser then rejects the FakeTensor before Dynamo can model eager's scalar conversion, producing the user-visible failure from #124344. Teach Dynamo's in-graph torch function path to inspect OpOverload and OpOverloadPacket schemas under capture_scalar_outputs. Whe...",https://github.com/pytorch/pytorch/pull/186457,015c78831bb3dacfe5571783121f4eed593e8afcf91e4ae24a23c14c0692df4c closes,pr,186459,issue,124181,high,pr.body,nly hide the symptom for this operator. Fixing fake view stride metadata addresses the traced alias/layout information at the source. Fixes #124181 Generated by my agent Test Plan: python test/dynamo/test_aot_autograd.py AotAutogradFallbackTests.test_aot_eager_group_norm_prese...,https://github.com/pytorch/pytorch/pull/186459,3c724dbba27d33d1612f264acb76e579275f6074115a640cde6bd820d8d847b5 closes,pr,186460,issue,124163,high,pr.body,all whitelist rather than exposing every optimize option so explain only accepts options that are meaningful for this debugging path. Fixes #124163 Generated by my agent Test Plan: python test/dynamo/test_explain.py -v python test/dynamo/test_backends.py TestExplainWithBackend...,https://github.com/pytorch/pytorch/pull/186460,21de03ce8e335196e62f5e80a4e17ca6dc1180600f0c758fca2e3df0601cc540 closes,pr,186461,issue,124110,high,pr.body,"o handle boolean ranges without numeric interval ordering. The behavior remains conservative when the boolean value cannot be proven. Fixes #124110 Generated by my agent Benchmark Results: Microbenchmark command recorded in state/124110/notes. Results for 20,000 iterations: ca...",https://github.com/pytorch/pytorch/pull/186461,765fb9993101191c2b33577fce6b56372a04906f4567c7f265805200f550bd8d closes,pr,185433,issue,158088,high,pr.body,dentOutputException pass through that debug backend unchanged while preserving the existing wrapping for other unexpected exceptions. Fixes #158088 Generated by my agent Benchmark Results: A micro-benchmark of the affected generated-code operation compares the old direct .item...,https://github.com/pytorch/pytorch/pull/185433,ec5ba75f37a01b0467e009b9329cdc1ea4665ea865507d992bb287330f567d2a references,pr,185433,issue,186031,medium,pr.comments[9].body,My agent says the Torchtitan failure matches the tracked torchcomms issue #186031 and is unrelated to this FakeTensor/Inductor change. I removed the optional ciflow/torchtitan label and refreshed the ghstack branch so the,https://github.com/pytorch/pytorch/pull/185433,dae220fb6f92a99d862de1d92c2d1165c1a5dc319cf1ede9eb20e0f03fd13cb2 review guidance,pr,185433,issue,158088,high,pr.reviews[0].body,"I appreciated the benchmark results but the entire approach feels wrong headed. First, we have an option to ""hop-ify"" inductor compiled code. When this occurs, fake tensor execution gets intercepted at the level of the entire inductor region. No problem. But even if we didn't have that, eventuall...",https://github.com/pytorch/pytorch/pull/185433,c3b93b4a112360790174c9ba6f0dd3eb82c2dbc3877a94d5e545a0efb4c1d606 review guidance,pr,185433,pr,185433,high,pr.reviews[0].body,"I appreciated the benchmark results but the entire approach feels wrong headed. First, we have an option to ""hop-ify"" inductor compiled code. When this occurs, fake tensor execution gets intercepted at the level of the entire inductor region. No problem. But even if we didn't have that, eventuall...",https://github.com/pytorch/pytorch/pull/185433,c6bceff5483cee7de56913a5668415290055375fa33d20426bc30e20f5603bba review guidance,pr,185433,issue,158088,high,pr.reviews[1].body,"I appreciated the benchmark results but the entire approach feels wrong headed. First, we have an option to ""hop-ify"" inductor compiled code. When this occurs, fake tensor execution gets intercepted at the level of the entire inductor region. No problem. But even if we didn't have that, eventuall...",https://github.com/pytorch/pytorch/pull/185433#pullrequestreview-4414788792,4c859847ce37da5f52deeb12ecb30f3636ca3f396ac5528cc685e22052630d17 review guidance,pr,185433,pr,185433,high,pr.reviews[1].body,"I appreciated the benchmark results but the entire approach feels wrong headed. First, we have an option to ""hop-ify"" inductor compiled code. When this occurs, fake tensor execution gets intercepted at the level of the entire inductor region. No problem. But even if we didn't have that, eventuall...",https://github.com/pytorch/pytorch/pull/185433#pullrequestreview-4414788792,3ad4c7182cba742248e3b1258e2de38b2c453f73362f0d3369d4fb97d22fcc7a closes,pr,185776,issue,148475,high,pr.body,"tensor failures. Limiting the conversion to active exception regions keeps the existing diagnostics outside user exception handling. Fixes #148475 Generated by my agent Benchmark Results: Baseline valid expand_as compile+first-run: mean 18.723 ms, median 18.645 ms over 30 iter...",https://github.com/pytorch/pytorch/pull/185776,c341a9e8c4e9eb83369a1f7ff19d7b1fcc7c903e3be35221e17ec98210d112d3 closes,pr,186463,issue,123663,high,pr.body,"-symbolic inplace cases, and add foreach coverage that asserts the ErrorInputs fail with the same error as the Tensor reference path. Fixes #123663 Generated by my agent Test Plan: python test/test_meta.py TestMetaCPU.test_dispatch_meta_inplace__foreach_add_cpu_bool TestMetaCP...",https://github.com/pytorch/pytorch/pull/186463,77c021a84cc245242cb82cd95d7f30fba62d0d4a4942bb0a1dda4455adbb9ed0 review guidance,pr,184365,pr,168073,high,pr.reviews[1].body,Do you have any performance results you could share ? still need to do review,https://github.com/pytorch/pytorch/pull/184365,97c5bb31691fb38d2bdc47298e8494c0b28e362dd5ee9340d3de58e6ee249b56 closes,pr,186464,issue,123651,high,pr.body,"quotient to be exact. This keeps the original range information and lets the source symbol specialize to the precise divisible value. Fixes #123651 Generated by my agent Benchmark Results: Focused ShapeEnv divisibility replacement microbenchmark, 300 iterations x5: Baseline ma...",https://github.com/pytorch/pytorch/pull/186464,4b66d7b37a63e0f3f32f5d28e49df7e7a3e4b5b9a3d4d55705ea9866ae4ea188 review guidance,pr,178737,pr,178737,high,pr.reviews[0].body,"Summary Enables torch.backends.cusparselt.version() and initCusparseltBindings on ROCm via a new USE_HIPSPARSELT cmake option, plus tightens 5 sparse-semi-structured test skips to be conditional on _IS_HIPSPARSELT_AVAILABLE. Three substantive problems below, plus a BC note worth surfacing in the...",https://github.com/pytorch/pytorch/pull/178737,d1912734a70cd6499ed7a7704d1f0f58e5d6a42e2d20e6b3e315787a98521b07 review guidance,pr,178737,pr,178737,high,pr.reviews[1].body,lgtm; please address Jeff's comments.,https://github.com/pytorch/pytorch/pull/178737,830cc12a6b253855d2969a1e6e66af79bab5d001bde9cd17fcde6845fb928f46 closes,pr,186466,issue,123470,high,pr.body,ions consume the SymFloat keeps value-independent custom ops compiling while preserving correct failures for data-dependent metadata. Fixes #123470 Generated by my agent Test Plan: python test/dynamo/test_misc.py MiscTests.test_custom_op_float_arg_accepts_scalar_tensor MiscTes...,https://github.com/pytorch/pytorch/pull/186466,cf01aa2d981033984e5fb1c2fae7e1f9197a05dd7b70424da0aef1f7e299dcc9 references,pr,186466,pr,185132,medium,pr.body,"akeTensor during compile and a real CPU tensor at runtime. lintrunner -a Duplicate check: inspected related custom-op PRs #186305, #184099, #185132, and #185096; none touches the Dynamo/FakeTensor float-schema scalar Tensor path. Benchmark Results: Command on main and fix bran...",https://github.com/pytorch/pytorch/pull/186466,abeb50cd1d84805c064886b30281b7fc75969ee17d765a7d73f6168a1d3be064 closes,pr,186467,issue,123411,high,pr.body,and named traversal. Changing all buffer naming would be broader and more likely to break existing compiled-module state_dict users. Fixes #123411 Generated by my agent Benchmark Results: Command: in-process microbenchmark comparing the previous no-sync OptimizedModule.call bo...,https://github.com/pytorch/pytorch/pull/186467,6d0d4c7f1699b9ba105904f872c847d4acd5cf6de44318a0bf613cb13e049900 closes,pr,186557,issue,90923,high,pr.body,torchvision stub due to the environment's torchvision::nms collection failure lintrunner -a $(git diff --name-only) git diff --check Fixes #90923 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv @...,https://github.com/pytorch/pytorch/pull/186557,2a3be9ca540f86315e4e121e4020a7ec564f94facae60c82b7f33c1ae95fb076 closes,pr,186470,issue,122860,high,pr.body,nsupported user program. Handling both call paths gives a consistent diagnostic at the boundary where parameter mutation is detected. Fixes #122860 Generated by my agent Test Plan: python test/export/test_experimental.py TestExperiment.test_joint_parameter_mutation_error TestE...,https://github.com/pytorch/pytorch/pull/186470,225372329518d0773719405ad91fd3716e5d4319002ea4cf5b029db746991471 closes,pr,186471,issue,122773,high,pr.body,"y overlap helper available to fail in other symbolic out/in-place meta paths, so the fix is placed at the root overlap-status helper. Fixes #122773 Generated by my agent Test Plan: Reproduced the original issue before the fix. ninja -C build torch_python ninja -C build install...",https://github.com/pytorch/pytorch/pull/186471,b1c52951b17718d1fdc9b3cacfa35af3ab16b03c33c91fce3adf62a793e23914 closes,pr,185926,issue,145785,high,pr.body,would leave the underlying serde invariant broken for other tuple arguments. Encoding tuple-ness in the schema fixes the root cause. Fixes #145785 Generated by my agent Test Plan: python test/export/test_serialize.py TestDeserialize.test_hoo_tuple_input python test/export/test...,https://github.com/pytorch/pytorch/pull/185926,371f589e92bd13558e7215a52546c1bb09ccafa5071ed1839fcb3a2f54c07527 closes,pr,185929,issue,145773,high,pr.body,is small Python reference-clearing overhead. The fix does not add cudagraph step markers or synchronization inside the measured loop. Fixes #145773 Generated by my agent Test Plan: python -m py_compile benchmarks/dynamo/common.py benchmarks/dynamo/huggingface.py benchmarks/dyn...,https://github.com/pytorch/pytorch/pull/185929,4416c6ef5afa996fea59ac454d031a9786dd377cd935a9766ff5b209a8d7eb9c closes,pr,189233,issue,175190,high,pr.closingIssuesReferences,pr #189233 declares a closing reference to issue #175190.,https://github.com/pytorch/pytorch/pull/189233,32f34ef5dbe319744c188933f512d5d0d7fdc2299fb1de868abb7d2c594c2285 closes,pr,189233,issue,175190,high,pr.body,Fixes #175190 Summary AvgPool2d / AdaptiveAvgPool2d backward on MPS either aborts the process (SIGABRT — buffer is not large enough from MPSNDArray) or s,https://github.com/pytorch/pytorch/pull/189233,bb0cdbf68a6a12cfe971ed8fe9b50ea6d33c9d3ffd526e84f290a86436d70e6c closes,pr,189045,issue,41508,high,pr.closingIssuesReferences,pr #189045 declares a closing reference to issue #41508.,https://github.com/pytorch/pytorch/pull/189045,d48f921c4f734cddd07343127bd2a2d94b93fc0ca5d8f16d31cff021218d3eb8 closes,pr,189045,issue,41508,high,pr.body,---Fixes #41508 Root cause nn.MultiheadAttention produces NaN gradients when a key_padding_mask (or attn_mask) fully masks out every key for a given query,https://github.com/pytorch/pytorch/pull/189045,8771659b072098b0222753d86c142f4b7876368c2ca3374950238012bf07baa3 closes,pr,184858,issue,174541,high,pr.body,"aliasing remains sound. This fixes the lookup root cause rather than changing cache accounting, which would only hide the recompiles. Fixes #174541 Generated by my agent Test Plan: python test/dynamo/test_dicts.py -k runtime_lookup python test/dynamo/test_getitem.py -k dict py...",https://github.com/pytorch/pytorch/pull/184858,82d116e499f1b0eeb94e3cefe018ae5d95dfe66e3b8545a1e9050e0d711006f0 references,pr,184858,issue,174541,medium,pr.comments[2].body,"minates unnecessary recompilations when different module instances use different keys into the same dict, directly fixing the root cause of #174541. The design is sound: guard the key's type (not value), emit `dict.__getitem__` at runtime, fall back for missing keys (with `DIC...",https://github.com/pytorch/pytorch/pull/184858,eff85f0cd1dbf9a6ede8cb4d6d004433d227ff89c4efcdb8d5b89303fd1d0c0b review guidance,pr,184858,issue,174541,high,pr.reviews[0].body,This change will collide with the work we're doing to improve LazyConstantVariables,https://github.com/pytorch/pytorch/pull/184858,71850b649872e0e76a2a72a20325ff68aba1c6bbc3da123ebff7bbf55e64fa41 review guidance,pr,184858,pr,184858,high,pr.reviews[0].body,This change will collide with the work we're doing to improve LazyConstantVariables,https://github.com/pytorch/pytorch/pull/184858,5717b67c15afb1db7714a3ed0c6d6e3e6b986d3e26c477fb9c8719ef44ea8416 review guidance,pr,184858,issue,174541,high,pr.reviews[1].body,This change will collide with the work we're doing to improve LazyConstantVariables,https://github.com/pytorch/pytorch/pull/184858#pullrequestreview-4520489678,c1b1b37814a636db2b74c2df8232632c8c6646c528232975f5c5a724863d3719 review guidance,pr,184858,pr,184858,high,pr.reviews[1].body,This change will collide with the work we're doing to improve LazyConstantVariables,https://github.com/pytorch/pytorch/pull/184858#pullrequestreview-4520489678,0a583a8a3c22b682977632af55509e2932d65fa0652c950fe653389ba5bc68ad closes,pr,189215,issue,186605,high,pr.body,"Stack from ghstack (oldest at bottom): -> #189215 Fixes #186867 Fixes #186866 Fixes #186605 Fixes #186604 Follow-up from #188855, which did not fix the failures Updates largeTensorTest to synchronize and collect any unused MPS buff",https://github.com/pytorch/pytorch/pull/189215,3c89d6703fe4938996a546037c529ef9634b1bb489c731106e724595e23657fe closes,pr,189215,issue,186867,high,pr.body,"Stack from ghstack (oldest at bottom): -> #189215 Fixes #186867 Fixes #186866 Fixes #186605 Fixes #186604 Follow-up from #188855, which did not fix the failures Updates largeTensorTest to synchronize and",https://github.com/pytorch/pytorch/pull/189215,c5808267604307cf271496d6487eecf9586c01ca24ea16fcbc479181d4b486f2 closes,pr,185947,issue,145596,high,pr.body,iling for data that should stay in the graph. The tensor rewrite fixes the root scalarization problem for the reported case directly. Fixes #145596 Generated by my agent Test Plan: python test/dynamo/test_functions.py -k 'test_logit_with_tensor_eps' python repro snippets for C...,https://github.com/pytorch/pytorch/pull/185947,4996589af69e54bf0bc8451d0406f4fafbd98f742b33815d91d96d3ef37b4c36 closes,pr,185952,issue,145574,high,pr.body,upports common escaped Work methods conservatively so compiled handles do not lose basic API shape if returned from a compiled frame. Fixes #145574 Generated by my agent Test Plan: python test/dynamo/test_fake_distributed.py -k test_compiled_async_all_reduce -v python test/exp...,https://github.com/pytorch/pytorch/pull/185952,301ec2284dc37b1f42427914076721f921019f058c27ec7a7ec962d7821b6f29 closes,pr,185971,issue,145445,high,pr.body,th. Replaying against the live source object fixes the root cause while retaining fullgraph support for ordinary random.Random calls. Fixes #145445 Generated by my agent Test Plan: python -m pytest test/dynamo/test_higher_order_ops.py::HigherOrderOpTests::test_cond_random_obje...,https://github.com/pytorch/pytorch/pull/185971,5ea240800b2e4cfd81682646123967d021fbfba01521e4ae29933b72179b236a closes,pr,185972,issue,145383,high,pr.body,xistence and compiler-kind check independent from the locale encoding while preserving the same detection behavior for normal output. Fixes #145383 Generated by my agent Test Plan: python test/inductor/test_compile.py TestStandaloneInductor.test_is_msvc_cl_handles_invalid_help...,https://github.com/pytorch/pytorch/pull/185972,9247adb0701d344bb97b8a899ca13206d148e464b591b6cb3e171c0ca5394194 closes,pr,185984,issue,145220,high,pr.body,"produce stale metadata. The final check is intentionally narrow: allow only non-input, non-view, non-distinct-alias internal resizes. Fixes #145220 Generated by my agent Test Plan: python test/dynamo/test_repros.py -k out_variant_resize TORCH_LOGS=graph_breaks python - <<'PY'...",https://github.com/pytorch/pytorch/pull/185984,9f281e307b04cb824163274e0adf815be3496e97e1699135cb3caf8349df28d6 references,pr,185984,issue,145220,medium,pr.comments[8].body,a different code path that doesn't raise `Unsupported`? ### Overall Assessment The change is sound for its narrowly scoped purpose (fixing #145220 for `torch.all`). The node metadata restoration logic correctly preserves the pre-resize shapes on producer nodes while updating l...,https://github.com/pytorch/pytorch/pull/185984,1e4a8c16c575adb044289e73d1feb200e909878bd0e5d5aa2020dce59dc82add review guidance,pr,185984,issue,145220,high,pr.reviews[0].body,"my codex tells me that this is what the output graph looks like: class GraphModule(torch.nn.Module): def forward(self, L_input_tensor_: ""b8[2, 3, 4]""): l_input_tensor_ = L_input_tensor_ empty: ""b8[2, 3]"" = torch.empty((0,), dtype=torch.bool) output_tensor: ""b8[2, 3]"" = empty.to(device(type='cpu')...",https://github.com/pytorch/pytorch/pull/185984,f1dd543285bdbc61d08d6ab5e8b1cd7c8e10f0a6f441e702ff1ef99c6ef6b77a review guidance,pr,185984,pr,185984,high,pr.reviews[0].body,"my codex tells me that this is what the output graph looks like: class GraphModule(torch.nn.Module): def forward(self, L_input_tensor_: ""b8[2, 3, 4]""): l_input_tensor_ = L_input_tensor_ empty: ""b8[2, 3]"" = torch.empty((0,), dtype=torch.bool) output_tensor: ""b8[2, 3]"" = empty.to(device(type='cpu')...",https://github.com/pytorch/pytorch/pull/185984,deaa4da9e87c900481ffd90e3689a6eb3202c341ff5af648afa3140a1d1471a1 review guidance,pr,185984,issue,145220,high,pr.reviews[1].body,the out metadata concern looks fixed to me but I don't know enough about TensorVariable here to review this,https://github.com/pytorch/pytorch/pull/185984,51346e7113394db310a70900969dafe2fcf906e2ee24de50770e5d2fbb3b020c review guidance,pr,185984,pr,185984,high,pr.reviews[1].body,the out metadata concern looks fixed to me but I don't know enough about TensorVariable here to review this,https://github.com/pytorch/pytorch/pull/185984,0ae6a153a4a3d1b587d940f3f133a4b145a427d30ca396cd629966247f1e1f78 review guidance,pr,185984,issue,145220,high,pr.reviews[2].body,"my codex tells me that this is what the output graph looks like: ``` class GraphModule(torch.nn.Module): def forward(self, L_input_tensor_: ""b8[2, 3, 4]""): l_input_tensor_ = L_input_tensor_ empty: ""b8[2, 3]"" = torch.empty((0,), dtype=torch.bool) output_tensor: ""b8[2, 3]"" = empty.to(device(type='c...",https://github.com/pytorch/pytorch/pull/185984#pullrequestreview-4450943302,14d1b3b7881a98adf646b3b33091e3eab21c53a4b180893f3f3ab4951a88d9b7 review guidance,pr,185984,pr,185984,high,pr.reviews[2].body,"my codex tells me that this is what the output graph looks like: ``` class GraphModule(torch.nn.Module): def forward(self, L_input_tensor_: ""b8[2, 3, 4]""): l_input_tensor_ = L_input_tensor_ empty: ""b8[2, 3]"" = torch.empty((0,), dtype=torch.bool) output_tensor: ""b8[2, 3]"" = empty.to(device(type='c...",https://github.com/pytorch/pytorch/pull/185984#pullrequestreview-4450943302,d89082d561701f0a745f452c29733526791caa513c2e011676ecc219231634c5 review guidance,pr,185984,issue,145220,high,pr.reviews[3].body,the out metadata concern looks fixed to me but I don't know enough about TensorVariable here to review this,https://github.com/pytorch/pytorch/pull/185984#pullrequestreview-4480120143,4d146f643ee88523152d1784b1e1b9d8e1e3b23d2e022530f5343d5cb8bdbd47 review guidance,pr,185984,pr,185984,high,pr.reviews[3].body,the out metadata concern looks fixed to me but I don't know enough about TensorVariable here to review this,https://github.com/pytorch/pytorch/pull/185984#pullrequestreview-4480120143,879f529e3cd6323e79bf9309632187bc043d89d7909dd365abceb798e21203d4 references,pr,187509,pr,187508,medium,pr.body,Stack from ghstack (oldest at bottom): -> #187509 #187508 Add a MultiProcessTestCase suite that exercises the core distributed collective and point-to-point API across an explicit backend allowlist,https://github.com/pytorch/pytorch/pull/187509,66015015d655cc2db8725f527dd8e13b67347072cfb7402b418ba684a6eb35f0 review guidance,pr,187509,pr,187508,high,pr.reviews[0].body,Seems fine. Just unclear what the point of referenced nccl2 backend is.,https://github.com/pytorch/pytorch/pull/187509,5932da0df6aaeb2766c7dd99383a56652f1f875ce1ba405d786ad5528870a411 review guidance,pr,187509,pr,187509,high,pr.reviews[0].body,Seems fine. Just unclear what the point of referenced nccl2 backend is.,https://github.com/pytorch/pytorch/pull/187509,ec91c9d89e6390e64f45d6ac8b5c62d7dc1dca7c6405411d571eae81cbde5389 closes,pr,189142,issue,189106,high,pr.closingIssuesReferences,pr #189142 declares a closing reference to issue #189106.,https://github.com/pytorch/pytorch/pull/189142,bdfacf07d66de8b9a019d4934dd225cb2cb3c379d017caceddf35e149f58f557 closes,pr,189142,issue,189106,high,pr.body,"Fixes #189106 Description Currently, get_signature_for_torch_op returns incorrect annotations for operations that return multiple tensors (e.g., aten.var",https://github.com/pytorch/pytorch/pull/189142,735b9b1d7c04fcba668dbeff2ec288a6eff5f4f6d63c4f3de685538e27377428 references,pr,187508,pr,187509,medium,pr.body,"Stack from ghstack (oldest at bottom): #187509 -> #187508 Add ProcessGroupNCCL2 under c10d/nccl2 as a Backend implementation that delegates the core collective, point-to-point, barrier,",https://github.com/pytorch/pytorch/pull/187508,da3fe5c4360d5e0f5ca9394f8eea96d8158faef30231af8f52aa7a9fc6a84547 review guidance,pr,187508,pr,187508,high,pr.reviews[0].body,What problem are we trying to solve with this is new backend?,https://github.com/pytorch/pytorch/pull/187508,00e89f317f361291a6625991be18b9f0ad7c2846a3a1a5166c0fa1a71ebc2803 review guidance,pr,187508,pr,187509,high,pr.reviews[0].body,What problem are we trying to solve with this is new backend?,https://github.com/pytorch/pytorch/pull/187508,49d52b47c9b4e8b37f4655ec21f07baf2026d8b9f5e6479d92da8a0966091364 references,pr,189165,pr,187859,medium,pr.body,"Stack from ghstack (oldest at bottom): -> #189165 #187859 #188376 #188367 #188366 Adds torch.compiler.precompile, an ahead-of-time precompile API that captures a whole computation with make_fx and",https://github.com/pytorch/pytorch/pull/189165,8e428870c178c84d5083be2baa5ab5d400a60238f0ad06c5e458bf3916b0e904 references,pr,189165,pr,188366,medium,pr.body,"Stack from ghstack (oldest at bottom): -> #189165 #187859 #188376 #188367 #188366 Adds torch.compiler.precompile, an ahead-of-time precompile API that captures a whole computation with make_fx and lowers it to a self-cont",https://github.com/pytorch/pytorch/pull/189165,47f51fcb3c75f0cb815169a5adb84990f4911036845ff908241b7a6753f0a734 references,pr,189165,pr,188367,medium,pr.body,"Stack from ghstack (oldest at bottom): -> #189165 #187859 #188376 #188367 #188366 Adds torch.compiler.precompile, an ahead-of-time precompile API that captures a whole computation with make_fx and lowers it to a s",https://github.com/pytorch/pytorch/pull/189165,55355fa39d931251bb880848dbdd128ce520070505670905a54dcbc943d231ee references,pr,189165,pr,188376,medium,pr.body,"Stack from ghstack (oldest at bottom): -> #189165 #187859 #188376 #188367 #188366 Adds torch.compiler.precompile, an ahead-of-time precompile API that captures a whole computation with make_fx and lowers i",https://github.com/pytorch/pytorch/pull/189165,f604066512f1f979aaf11ac45f014dc7e0bd7d388a03794acfd4dcab64e7ecbf references,pr,188366,pr,187859,medium,pr.body,Stack from ghstack (oldest at bottom): #189165 #187859 #188376 #188367 -> #188366 _create_runtime_wrapper codegens the output-alias (_alias_fn) and input-mutation (_apply_mutations) epilogue hel,https://github.com/pytorch/pytorch/pull/188366,d87a4ee15d69681ed5b4c1f1dd7125c3a1c6c7c304677a087a3dbcdbcfc7dbc5 references,pr,188366,pr,188367,medium,pr.body,Stack from ghstack (oldest at bottom): #189165 #187859 #188376 #188367 -> #188366 _create_runtime_wrapper codegens the output-alias (_alias_fn) and input-mutation (_apply_mutations) epilogue helpers as standalo,https://github.com/pytorch/pytorch/pull/188366,842dec309f48d25ee96822bfdbb27c88aa9c73afcb3d6873931b8fdf75626e8e references,pr,188366,pr,188376,medium,pr.body,Stack from ghstack (oldest at bottom): #189165 #187859 #188376 #188367 -> #188366 _create_runtime_wrapper codegens the output-alias (_alias_fn) and input-mutation (_apply_mutations) epilogue helpers as,https://github.com/pytorch/pytorch/pull/188366,e77a551f6cc000fa87006f4e6a20987886d636ac764780233d6ee29b22f16303 references,pr,188366,pr,189165,medium,pr.body,Stack from ghstack (oldest at bottom): #189165 #187859 #188376 #188367 -> #188366 _create_runtime_wrapper codegens the output-alias (_alias_fn) and input-mutation (_apply_mutations) epil,https://github.com/pytorch/pytorch/pull/188366,3b6be4df81470e638f4d19ba5447344094f2d4a05550ab1a9260efa5e1554f07 references,pr,188367,pr,187859,medium,pr.body,Stack from ghstack (oldest at bottom): #189165 #187859 #188376 -> #188367 #188366 _compile_and_exec_source -- the chokepoint that compiles a generated wrapper source string into a live function,https://github.com/pytorch/pytorch/pull/188367,04b4b5afd1ef3b0077fc2f211ef6859a282f4325a8c5df02abb7cbfcd2380590 references,pr,188367,pr,188366,medium,pr.body,Stack from ghstack (oldest at bottom): #189165 #187859 #188376 -> #188367 #188366 _compile_and_exec_source -- the chokepoint that compiles a generated wrapper source string into a live function -- lived in subclass_codege,https://github.com/pytorch/pytorch/pull/188367,b6bf432c6c8d7953e8ca729e0802df4ffc492a6002eaf389aff5379c26f5211f references,pr,188367,pr,188376,medium,pr.body,Stack from ghstack (oldest at bottom): #189165 #187859 #188376 -> #188367 #188366 _compile_and_exec_source -- the chokepoint that compiles a generated wrapper source string into a live function -- lived,https://github.com/pytorch/pytorch/pull/188367,7441ee49fa488f606810f4094e6cb96d3b0d3488d8826718711006a72d296196 references,pr,188367,pr,189165,medium,pr.body,Stack from ghstack (oldest at bottom): #189165 #187859 #188376 -> #188367 #188366 _compile_and_exec_source -- the chokepoint that compiles a generated wrapper source string into a live f,https://github.com/pytorch/pytorch/pull/188367,452ea5ed84cf6a4958c9b96e59a4365f3dc65a3ea1f57a9458f1c3f493adfe03 references,pr,188376,pr,187859,medium,pr.body,"Stack from ghstack (oldest at bottom): #189165 #187859 -> #188376 #188367 #188366 source_emit.py is a small leaf module with one job: given a live Python object, return a Python expression strin",https://github.com/pytorch/pytorch/pull/188376,58f0fcd8e8a74590d10cb4febaf3531b0dc84b0d846c29f2fbd55781cb61bb69 references,pr,188376,pr,188366,medium,pr.body,"Stack from ghstack (oldest at bottom): #189165 #187859 -> #188376 #188367 #188366 source_emit.py is a small leaf module with one job: given a live Python object, return a Python expression string that, when exec'd, recons",https://github.com/pytorch/pytorch/pull/188376,25b31acb2fba615c6aace26488decb78d1e2f036adce0a875684b6cbc261485b references,pr,188376,pr,188367,medium,pr.body,"Stack from ghstack (oldest at bottom): #189165 #187859 -> #188376 #188367 #188366 source_emit.py is a small leaf module with one job: given a live Python object, return a Python expression string that, when exec'd",https://github.com/pytorch/pytorch/pull/188376,8007a26ce318f015c5e03ffde8eb4679b3ec20a8bd0102feeba008e1c20acf76 references,pr,188376,pr,189165,medium,pr.body,"Stack from ghstack (oldest at bottom): #189165 #187859 -> #188376 #188367 #188366 source_emit.py is a small leaf module with one job: given a live Python object, return a Python expressi",https://github.com/pytorch/pytorch/pull/188376,daab47519f1e92ae7e46bd06d396900f4a3990baa56b8ef32de1d2080a99ab64 references,pr,187859,pr,188366,medium,pr.body,"Stack from ghstack (oldest at bottom): #189165 -> #187859 #188376 #188367 #188366 Add torch._functorch.aot_autograd.compile_to_python(gm, example_inputs) -> (python, cache), the outer half of the backend contract behind t",https://github.com/pytorch/pytorch/pull/187859,daf4c0f77cb5993e1f8f85927203fbf920784517bf8b2cba5a319d8a7b16bf2c references,pr,187859,pr,188367,medium,pr.body,"Stack from ghstack (oldest at bottom): #189165 -> #187859 #188376 #188367 #188366 Add torch._functorch.aot_autograd.compile_to_python(gm, example_inputs) -> (python, cache), the outer half of the backend contract",https://github.com/pytorch/pytorch/pull/187859,b46494c3e20ed0a6e21f2dc1cfcd667c1ee2aa7b205b5bde9822df4f365069e1 references,pr,187859,pr,188376,medium,pr.body,"Stack from ghstack (oldest at bottom): #189165 -> #187859 #188376 #188367 #188366 Add torch._functorch.aot_autograd.compile_to_python(gm, example_inputs) -> (python, cache), the outer half of the backend c",https://github.com/pytorch/pytorch/pull/187859,172592f6da2807ec44cbd40bf8b942c25cb1f589cb5b638796230b08e3c8f47e references,pr,187859,pr,189165,medium,pr.body,"Stack from ghstack (oldest at bottom): #189165 -> #187859 #188376 #188367 #188366 Add torch._functorch.aot_autograd.compile_to_python(gm, example_inputs) -> (python, cache), the outer ha",https://github.com/pytorch/pytorch/pull/187859,71fb28f1c07ab1522e7edd316db68c56d1f1f568492aa3f53a6b76fb64c51a39 closes,pr,186007,issue,144376,high,pr.body,"the backward output mapping. Preserve the order through deduped inputs, tensor subclass remapping, and compiled autograd cache keys. Fixes #144376 Test Plan: python -m py_compile torch/_functorch/_aot_autograd/graph_compile.py torch/_functorch/_aot_autograd/input_output_analys...",https://github.com/pytorch/pytorch/pull/186007,cb8b1d5945f1ebfdcbf7cbc7ac4dd2ad3460c8baec0bba123a85fb859f6b332b closes,pr,184094,issue,85852,high,pr.closingIssuesReferences,pr #184094 declares a closing reference to issue #85852.,https://github.com/pytorch/pytorch/pull/184094,44cffcf19d2435b72791b1694ec4d26315ce98e5b3093f13432cb0b99bb85328 closes,pr,184094,issue,85852,high,pr.body,"Issue Fixes #85852 Summary addmm_impl_cpu_ delegates to BLAS gemm, which cannot operate in-place on its input matrices. When a caller passes the same tensor f",https://github.com/pytorch/pytorch/pull/184094,9732300027f1d1c091296f59e947b10816bf97fbb8f92102cc4d5c1521ab0537 closes,pr,186011,issue,144360,high,pr.body,"eserves the existing cached empty-frame skip behavior while using the eval-frame strategy mechanism instead of adding a new sentinel. Fixes #144360 Generated by my agent Benchmark Results: A 30-iteration cold compile microbenchmark for the issue repro using backend=""eager"" and...",https://github.com/pytorch/pytorch/pull/186011,ea8cbcc3d0d1089bcd7efc613a37ebc407f98ec4991f655b4a0cee1e5dc97c52 references,pr,186011,issue,186031,medium,pr.comments[7].body,"tegration failing with `RuntimeError: Could not resolve the process group registered under the name 19/20`, which matches preexisting issue #186031 about torchcomms-managed process groups under torch.compile. I did not change code for this PR because the failing path is outsid...",https://github.com/pytorch/pytorch/pull/186011,9e163735949d89bafb69a7d85f12deeee9f2baa6b7098fabfb883e0dd29ca7c5 review guidance,pr,188472,pr,188472,high,pr.reviews[0].body,LGTM!,https://github.com/pytorch/pytorch/pull/188472,9ea68cbd64550c0a9d34586c5db7f2aac12dc4358cd0b024b83749bc8f6f23c9 closes,pr,186016,issue,144247,high,pr.body,"ve broadened the patch beyond the Inductor bug. Keeping the validation in Inductor lowering directly fixes the failing compiler path. Fixes #144247 Generated by my agent Test Plan: d=$(mktemp -d); printf 'raise ImportError(""stub torchvision unavailable for this targeted test"")...",https://github.com/pytorch/pytorch/pull/186016,ebf25e888e5cbee34c7184d72abf74d5657dc54fb086774a77bd375351f1570f closes,pr,184963,issue,171977,high,pr.body,compiled backends is not currently supported. This change is scoped to shallow copy.copy and does not change copy.deepcopy behavior. Fixes #171977 Generated by my agent Test Plan: python test/export/test_export.py TestExport.test_export_copy_copy_requires_grad_unsupported pyth...,https://github.com/pytorch/pytorch/pull/184963,a5d061066a53779978129f6262f37047bc3f6e8ff1a3d178a3a7f52689296d97 review guidance,pr,184963,issue,171977,high,pr.reviews[0].body,This feels like a foot in the gun from export. they should trace with fake tensor. I would want weigh in from export before adding this amount of complication to fake tensor.,https://github.com/pytorch/pytorch/pull/184963,91d7e40045a8bf18f5f3433f90219b1c28c7b0965afffe030ec3e0e4467f741d review guidance,pr,184963,pr,184963,high,pr.reviews[0].body,This feels like a foot in the gun from export. they should trace with fake tensor. I would want weigh in from export before adding this amount of complication to fake tensor.,https://github.com/pytorch/pytorch/pull/184963,0f32208be2b9adce2990c19d3d4f8a0940be22a5d921bc689d25921af3102791 review guidance,pr,184963,issue,171977,high,pr.reviews[1].body,This feels like a foot in the gun from export. they should trace with fake tensor. I would want weigh in from export before adding this amount of complication to fake tensor.,https://github.com/pytorch/pytorch/pull/184963#pullrequestreview-4367854457,61a64ac1e39bba9b99caba20e09904cc9d731d7017432f3ad884fcdcf08500ff review guidance,pr,184963,pr,184963,high,pr.reviews[1].body,This feels like a foot in the gun from export. they should trace with fake tensor. I would want weigh in from export before adding this amount of complication to fake tensor.,https://github.com/pytorch/pytorch/pull/184963#pullrequestreview-4367854457,2462f2beb27e84758806225cd46f620b6d3f41e2c857d13ff041e07669c44351 closes,pr,186021,issue,143649,high,pr.body,"ocalized to CPU C++ codegen instead of adding Python-side prechecks, which would miss generated-code paths and add extra tensor work. Fixes #143649 Generated by my agent Benchmark Results: A small CPU int64 non-zero divisor microbenchmark was run with torch.set_num_threads(1),...",https://github.com/pytorch/pytorch/pull/186021,0c9b4ce915fbc3f61d0b575fbd6cc051404382ab454cc10a9c52b033e53632ed closes,pr,186833,issue,170539,high,pr.body,so they stay in place. This follows the abandoned #170540 direction but revalidates the previous getitem CI concern on current main. Fixes #170539 Generated by my agent Test Plan: python test/export/test_export_opinfo.py -v TestExportOpInfoCPU.test_fake_export_sparse_sampled_a...,https://github.com/pytorch/pytorch/pull/186833,55449a98ab0e4f6e1e24eb6b4b8388ef24a8650480f9849fe70493683c2aa409 closes,pr,186472,issue,122578,high,pr.body,"Stack from ghstack (oldest at bottom): -> #186472 Fixes #122578 When Dynamo sees mutation on an nn.Module, it restarts analysis and tracks the module as an UnspecializedNNModuleVariable. For ModuleList-l",https://github.com/pytorch/pytorch/pull/186472,883f41b1dc91f02358df3c56709b0811685a419a608ddc3bdac6891191845306 closes,pr,186478,issue,122386,high,pr.body,ze/stride metadata and copying data back. This keeps the fix in the generic mutation replay path instead of special-casing histogram. Fixes #122386 Generated by my agent Benchmark Results: Metadata-resize mutation epilogue microbenchmark command: python - <<'PY' ... PY. main:...,https://github.com/pytorch/pytorch/pull/186478,4615d81655e7780f918aead21d73c733d6159489d40a2f4007c94dd86152db97 closes,pr,186487,issue,122200,high,pr.body,backends isolated from ambient functorch modes without losing the information needed for fake tensor propagation and runtime guards. Fixes #122200 Generated by my agent Benchmark Results: A repeated compile micro-benchmark was run because this touches Dynamo compile entry. Com...,https://github.com/pytorch/pytorch/pull/186487,e20efe30f8be1ba4f18aa0f00fb24b461bf45aa6171db3a66665d2f4978cbff5 closes,pr,186488,issue,122129,high,pr.body,lue. Improving the schema formatter fixes the root cause of the obscure message and benefits the same conversion path more generally. Fixes #122129 Generated by my agent Test Plan: MAX_JOBS=16 python setup.py develop python test/dynamo/test_misc.py MiscTests.test_custom_op_int...,https://github.com/pytorch/pytorch/pull/186488,ba698a25d0d88c5e4887749f02bafa8e187f1faf0ff3bfd5b77a5de5cb6688fd closes,pr,186248,issue,137096,high,pr.body,Min/Max avoids leaving equivalent flattened expressions unsimplified while still requiring ShapeEnv to prove every removed argument. Fixes #137096 Generated by my agent Benchmark Results: Mixed cold simplify workload with 400 legacy Max expressions and 400 newly handled Min ex...,https://github.com/pytorch/pytorch/pull/186248,09929b7f2206a08cab6eeea3742f176d02429630756abafc87713bb0146a1ec9 closes,pr,186189,issue,139080,high,pr.body,marker from the dynamo expected-failure wrapper. This keeps the Dynamo xfail signal authoritative when a stale entry starts passing. Fixes #139080 Generated by my agent Test Plan: PYTORCH_TEST_WITH_DYNAMO=1 python test/dynamo/test_modules.py NNModuleTests.test_lazy_module2 PYT...,https://github.com/pytorch/pytorch/pull/186189,48b818ae5e2c7a1c73c590fd85a68ca2394e1fcfe7a0a1b93cf1c6c632c6cbf7 closes,pr,184048,issue,161132,high,pr.body,"/Triton gate with device-specific skips, and keep stale skipped timing data from preventing pytest sharding of the large OpInfo file. Fixes #161132 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...",https://github.com/pytorch/pytorch/pull/184048,bc14d4b3bcd8bb55d6394ca9bbf168f1101a2348a0569c0fed9b053fb5f0fa22 references,pr,188181,pr,188105,medium,pr.body,"Stack from ghstack (oldest at bottom): #188105 -> #188181 Deduplicates the identical dict_vt field + get_dict_vt accessor that lived on UserDefinedObjectVariable, the user-function varia",https://github.com/pytorch/pytorch/pull/188181,0cffde1c557ac7f76ea911a3235796cf477568b7d60ff8c4f5ba0d0bdefd2431 closes,pr,184223,issue,139629,high,pr.body,"TDispatcher already handled. Covers Python and C++ wrapper fallback paths, fallback out variants, and special fallback codegen paths. Fixes #139629 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...",https://github.com/pytorch/pytorch/pull/184223,6e31612197dea9bf7cb3d6cbecd06602946583a272fe1b04c36701c2463712f7 closes,pr,186493,issue,122029,high,pr.body,"-line Python and preserves existing assertions in generated form, including debug-only mutated-input and aliased-intermediate checks. Fixes #122029 Generated by my agent Benchmark Results Command: python benchmarks/dynamo/microbenchmarks/overheads.py Before: requires_grad=Fals...",https://github.com/pytorch/pytorch/pull/186493,c34d3307f008d129e7970edb9689645e286790d9159c3e4ae7a162a1b0ade2f0 closes,pr,189038,issue,160744,high,pr.closingIssuesReferences,pr #189038 declares a closing reference to issue #160744.,https://github.com/pytorch/pytorch/pull/189038,39b75591d90c2366c9fb2da80c035268ecc7e629eace1761f9a9be3c5d900dc9 closes,pr,189038,issue,160744,high,pr.body,"Fixes #160744 On MPS, copy_ from a negative floating point source into a torch.uint8 destination can produce 0 for broadcasted shapes like [3] and [2, 2]",https://github.com/pytorch/pytorch/pull/189038,cf52e14164c59dc59df56b501a197f0117e56a95dd2050f56729e2421875be98 review guidance,pr,189038,issue,160744,high,pr.reviews[0].body,Casts outside of the target range are UB by default. They don't need to match,https://github.com/pytorch/pytorch/pull/189038,23f0ba35389302c16a112292e12a02927eeb70dd135b85f4d949ae36a559385d closes,pr,186494,issue,121526,high,pr.body,eeps the fix limited to the sound case needed by CUDA user Triton code instead of treating all is_pinned calls as metadata constants. Fixes #121526 Generated by my agent Test Plan: python test/dynamo/test_misc.py -k test_cuda_tensor_is_pinned_constant_false lintrunner -a git d...,https://github.com/pytorch/pytorch/pull/186494,151553873946f04048b1da1c93f514dd6ce08fd64bd7c8f03c9b6804bab9415a closes,pr,186495,issue,121318,high,pr.body,Plan: python test/export/test_export_logging.py lintrunner -a lintrunner -a torch/export/_trace.py test/export/test_export_logging.py Fixes #121318 Generated by my agent,https://github.com/pytorch/pytorch/pull/186495,e7501de00ef4f1b523bb726cd888fd673df5d1898b6d9053d6e826c3626a1ab5 closes,pr,188047,issue,188023,high,pr.closingIssuesReferences,pr #188047 declares a closing reference to issue #188023.,https://github.com/pytorch/pytorch/pull/188047,5aaa7330e4162d293e3fb594d182b5fd1abe0b0fbb263f9735c73cb3ceba0eb0 closes,pr,188047,issue,188023,high,pr.body,Fixes #188023 Summary torch.from_dlpack() aborted the entire Python process (SIGABRT / libc++abi: terminating due to uncaught exception) when given any a,https://github.com/pytorch/pytorch/pull/188047,bb8a2f694d01821b7d8116811118e3d6caaa8ad59ba6a10979175319634998da review guidance,pr,188047,issue,188023,high,pr.reviews[0].body,"Thanks for picking this up. I think this is going in the right direction, especially the C++ guard before from_blob so the capsule path also raises instead of aborting. One area that may be worth tightening is the test coverage around copy semantics, not just the negative-stride crash case. Right...",https://github.com/pytorch/pytorch/pull/188047,76549429b3970f30bdb99b407d296fb5ea4acbc62af6ad327465544a84ce8755 closes,pr,184910,issue,174176,high,pr.body,eps the supported Tensor.set_ path separate from unsupported direct op calls instead of weakening functionalization's lift assertion. Fixes #174176 Generated by my agent Test Plan: python test/dynamo/test_misc.py MiscTests.test_direct_aten_set_unsupported MiscTests.test_inplac...,https://github.com/pytorch/pytorch/pull/184910,caa3827a32a4c9c30b4b5525e296598cace80113c0a7bbd5a99282670f08f312 closes,pr,185069,issue,166093,high,pr.body,stead of disabling AOTI PCH for Windows cross targets so the configured PCH optimization remains available once the command is valid. Fixes #166093 Generated by my agent Test Plan: python test/inductor/test_compile.py TestStandaloneInductor.test_windows_cross_target_link_flags...,https://github.com/pytorch/pytorch/pull/185069,b90b489c1b69ae40f464a1a4711bdeca2195ad73d0dbc062b21d4979049f2d31 review guidance,pr,185069,issue,166093,high,pr.reviews[0].body,"Looks plausible to me. To land this, we should confirm that it actually fixes the Linux-to-Windows AOTI cross-compilation issue. I do not see an end to end test for that. If it is not possible to exercise this through the CI, we should at least confirm this manually before landing. Not sure if yo...",https://github.com/pytorch/pytorch/pull/185069,0296715bdbcb38e6e94f87f109a0242356d5216397c02f4fe9433a0f1c54633a review guidance,pr,185069,pr,185069,high,pr.reviews[0].body,"Looks plausible to me. To land this, we should confirm that it actually fixes the Linux-to-Windows AOTI cross-compilation issue. I do not see an end to end test for that. If it is not possible to exercise this through the CI, we should at least confirm this manually before landing. Not sure if yo...",https://github.com/pytorch/pytorch/pull/185069,3eabebdf1b017af7a301b53b56a6be9bfbd2405d1e59381da5700e7d85cc7fb5 review guidance,pr,185069,issue,166093,high,pr.reviews[2].body,"Looks plausible to me. To land this, we should confirm that it actually fixes the Linux-to-Windows AOTI cross-compilation issue. I do not see an end to end test for that. If it is not possible to exercise this through the CI, we should at least confirm this manually before landing. Not sure if yo...",https://github.com/pytorch/pytorch/pull/185069#pullrequestreview-4471530613,7d800b8c16f3c82135d7ce97c8e5c60df8ca82feb833208c6e9aa3cb5fd88893 review guidance,pr,185069,pr,185069,high,pr.reviews[2].body,"Looks plausible to me. To land this, we should confirm that it actually fixes the Linux-to-Windows AOTI cross-compilation issue. I do not see an end to end test for that. If it is not possible to exercise this through the CI, we should at least confirm this manually before landing. Not sure if yo...",https://github.com/pytorch/pytorch/pull/185069#pullrequestreview-4471530613,84efefb339344aadf98e34610348bdff0aeb609e01327b85a927d44538e0ee45 review guidance,pr,185701,pr,185701,high,pr.reviews[0].body,needs work,https://github.com/pytorch/pytorch/pull/185701,3588b7bbb529aa6ebd48bbf332d65b0ce7ef378c1680248cbc1737630abbbace closes,pr,186498,issue,120911,high,pr.body,ization such as one static input dim plus another symbolic dim could still pass dynamic-shape coverage. That is the gap called out in issue #120911 and it meant tests like the prod backward coverage from the referenced work could miss real specialization regressions. Add an op...,https://github.com/pytorch/pytorch/pull/186498,0ec6bdc2780eec7d822e34b5dd8807e56c716a09b849ea753794f2515b2f500e references,pr,185918,pr,185917,medium,pr.body,"Stack from ghstack (oldest at bottom): -> #185918 #185917 Summary Add test list registry to torch/testing/_internal/autocast_test_lists.py: register_autocast_test_lists(device_type, cls) for new ba",https://github.com/pytorch/pytorch/pull/185918,5b73a74c6d32470ab2ee54a6b9d66ba9b02dfb3c0787338158abe2958bd05b5f closes,pr,186897,issue,154851,high,pr.body,Stack from ghstack (oldest at bottom): -> #186897 Fixes #154851 Generated by my agent Custom operators defined through torch.library preserve an important user-facing distinction in their schema text: in,https://github.com/pytorch/pytorch/pull/186897,2d42be5010f9245e49f50d3a6d94eba6f16951c02f087d75a9186ffd82b4faee review guidance,pr,186897,issue,154851,high,pr.reviews[0].body,I don't really want to accept this just because of the LOC. I don't want to complicated torch._ops more. I feel like this shouldn't be a difficult change,https://github.com/pytorch/pytorch/pull/186897,2a499e2193b48708d5eb77d235a85fdd8cf70da3a7800bcc888435b7eb5e07a9 review guidance,pr,186897,pr,186897,high,pr.reviews[0].body,I don't really want to accept this just because of the LOC. I don't want to complicated torch._ops more. I feel like this shouldn't be a difficult change,https://github.com/pytorch/pytorch/pull/186897,2706a6d232b797a6c07987582b541943abc1b904727c27bb4b4a6bfe5e80d085 review guidance,pr,186897,issue,154851,high,pr.reviews[1].body,I don't really want to accept this just because of the LOC. I don't want to complicated torch._ops more. I feel like this shouldn't be a difficult change,https://github.com/pytorch/pytorch/pull/186897#pullrequestreview-4519840130,c4f28b8bbca40239c6abdb1d5327caab4ff05f2adbf08837a337ae1213f249ff review guidance,pr,186897,pr,186897,high,pr.reviews[1].body,I don't really want to accept this just because of the LOC. I don't want to complicated torch._ops more. I feel like this shouldn't be a difficult change,https://github.com/pytorch/pytorch/pull/186897#pullrequestreview-4519840130,415b9109d6dad075b457fa534b9230468b958089247d873ccaa40a2553538a94 closes,pr,184380,issue,120381,high,pr.body,nductor from materializing full-size fp32 intermediates in checkpoint backward graphs while leaving normal fusion behavior unchanged. Fixes #120381 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/184380,f45c4a8fa4ff332ba8922d881a6de88b37d4d12b03871f0787373e993b0f6010 review guidance,pr,184380,issue,120381,high,pr.reviews[0].body,can we get a full benchmark run results?,https://github.com/pytorch/pytorch/pull/184380,72d73e9b78aaba7d13c1d0d5e5d15db2ebb43f01478973e49b4a165f94e62d45 review guidance,pr,184380,pr,184380,high,pr.reviews[0].body,can we get a full benchmark run results?,https://github.com/pytorch/pytorch/pull/184380,01bbc73ca377c52e59018eca98b857876852a974df8d745dc28a8f8307966561 review guidance,pr,184380,issue,120381,high,pr.reviews[1].body,Can you do a perf run ? there was a prior attempt at this that made perf wors.e,https://github.com/pytorch/pytorch/pull/184380,645389192b6fc9f522aad436c526da36d025ab46cfc79c993dae4c6e53201e48 review guidance,pr,184380,pr,184380,high,pr.reviews[1].body,Can you do a perf run ? there was a prior attempt at this that made perf wors.e,https://github.com/pytorch/pytorch/pull/184380,8a21aa6f2b7e46d597c7281f05968e60464c6d6dfefbc2648b8014b7ed34289a closes,pr,186500,issue,120375,high,pr.body,"itself, but that would miss the existing StringFormatVariable machinery and would be broader than necessary for f-string formatting. Fixes #120375 Generated by my agent Test Plan: python test/dynamo/test_reorder_logs.py -k test_logger_fstring_debug_resumes_after_graph_break py...",https://github.com/pytorch/pytorch/pull/186500,c74529fedaab45ee9279265a844823fdd1a3953c0c9d824072c9772f46195696 closes,pr,184392,issue,119883,high,pr.body,"ers share Python floor-division and modulo semantics, while preserving the MPS safe_mod workaround for the known Metal compiler case. Fixes #119883 Fixes #187027 Generated by my agent cc @jgong5 @mingfeima @XiaobingSuper @sanchitintel @ashokei @jingxu10 @jerryzh168 @aditew01 @...",https://github.com/pytorch/pytorch/pull/184392,72dcfd6b1672896d7ebdd2c919e62cf7d28431fbe19e366fd90c593199d388f1 closes,pr,184392,issue,187027,high,pr.body,"on floor-division and modulo semantics, while preserving the MPS safe_mod workaround for the known Metal compiler case. Fixes #119883 Fixes #187027 Generated by my agent cc @jgong5 @mingfeima @XiaobingSuper @sanchitintel @ashokei @jingxu10 @jerryzh168 @aditew01 @voznesenskym @...",https://github.com/pytorch/pytorch/pull/184392,10b89945f0f63a6220113971f2b63cf06146878f47719a2428879636557b2c8b closes,pr,186373,issue,129456,high,pr.body,"workloads that explicitly want duck sizing, and the logging test that checks the duck-sizing guard message now enables it explicitly. Fixes #129456 Generated by my agent Benchmark Results: Issue-shaped two-call Dynamo benchmark, backend=CompileCounter, dynamic=True, 10 cold it...",https://github.com/pytorch/pytorch/pull/186373,945bcdb34a71bc2db433acdcc9d8707539eeea6cd857398d46277b02282faa49 closes,pr,186283,issue,136271,high,pr.body,"on is for Dynamo-created graph inputs to tensor constructors, and avoiding public ATen schema/codegen changes keeps the fix narrower. Fixes #136271 Generated by my agent Test Plan: python test/dynamo/test_misc.py -k ""data_ptr"" python test/inductor/test_triton_kernels.py Kernel...",https://github.com/pytorch/pytorch/pull/186283,d1f0eac6ed80b279859b1ee7cbb8d25e13aaa6eeb24bb05d7be80b41f9041300 closes,pr,181720,issue,181374,high,pr.body,"Please note that in the current form it gives no perf benefit over regular CPU memory, it this is a foundation for copy-on-write work Fixes #181374 Fixes #188970 Authored with assistance from Claude Code.",https://github.com/pytorch/pytorch/pull/181720,3e77a49d0b11e8e1bb339954e060b531ae7d022db0646ee2db051e6d90653083 closes,pr,181720,issue,188970,high,pr.body,"at in the current form it gives no perf benefit over regular CPU memory, it this is a foundation for copy-on-write work Fixes #181374 Fixes #188970 Authored with assistance from Claude Code.",https://github.com/pytorch/pytorch/pull/181720,5ea45043d98df0a4ad322085dd827c3bb07e1d3fd75a4445c1295fccd497fc54 references,pr,181720,pr,189256,medium,pr.body,"Stack from ghstack (oldest at bottom): -> #181720 #189256 torch.empty(..., device=""cpu"", pin_memory=True) and tensor.pin_memory() previously returned tensors whose .device was mps:0, because the MP",https://github.com/pytorch/pytorch/pull/181720,9218599874336e8212193eb55de71d086cdb3293ba8287bf133785aac8eee75a review guidance,pr,181720,issue,181374,high,pr.reviews[0].body,Thanks!,https://github.com/pytorch/pytorch/pull/181720,333dcc7a493db6b08298e44f1ec259f3c77454017cc5f42091c38c1684f16019 review guidance,pr,181720,pr,181720,high,pr.reviews[0].body,Thanks!,https://github.com/pytorch/pytorch/pull/181720,a1012f3a4a88bbb21020e2517c66f1ced330a8bf287d587dd330cda29cc86ff3 review guidance,pr,181720,issue,188970,high,pr.reviews[0].body,Thanks!,https://github.com/pytorch/pytorch/pull/181720,b2a313153579148cba707a4d4878e55dec0e0a3060271ccb5b9cb7aba60f57c5 review guidance,pr,181720,pr,189256,high,pr.reviews[0].body,Thanks!,https://github.com/pytorch/pytorch/pull/181720,77ebc4646e156a17755e1ba84536e439b47b9c2a37158dcb5396cf949de3ba3e closes,pr,186502,issue,118747,high,pr.body,nction intact while making traceback frames report the context manager that is actually active. Generator wrappers use the same path. Fixes #118747 Generated by my agent Test Plan: python test/test_utils.py TestTraceback.test_context_decorator_traceback_frame_name TestTracebac...,https://github.com/pytorch/pytorch/pull/186502,595275d72463526cba1bc09a53672003113d4abfaed17128eac33b720016f6df closes,pr,186503,issue,118739,high,pr.body,"s is not expected to be a profitable optimization target, and the old path was observably inconsistent with eager autograd semantics. Fixes #118739 Generated by my agent Test Plan: Original issue repro now passes; compiled unbind outputs have UnbindBackward0 grad_fns and repea...",https://github.com/pytorch/pytorch/pull/186503,72d014b75bc57e6d8b5933268208e5fcbff8f73911517afbf8734e3ce95883f9 closes,pr,186504,issue,118334,high,pr.body,"dedicated HOP context limits the behavior change to framework autograd.Function bodies that HOP speculation already needs to inspect. Fixes #118334 Generated by my agent Benchmark Results: Repeated 10 one-shot torch.compile(fn, backend=""eager"", fullgraph=True) runs on a small...",https://github.com/pytorch/pytorch/pull/186504,4216d5f27bb5f35b9ed2133f16aa9390d9c4d59064340f5f9329e3529d6b5e33 closes,pr,181308,issue,122191,high,pr.closingIssuesReferences,pr #181308 declares a closing reference to issue #122191.,https://github.com/pytorch/pytorch/pull/181308,3ebe84f74e0e9f789699e49c0c960a7dbd110276e2859d231e161d59105e1e36 closes,pr,181308,issue,122191,high,pr.body,"Fixes #122191. Summary torch.set_num_threads(1); x = torch.rand(128, 128, 128) torch.std_mean(x) # ~7 ms on M5 torch.std(x); torch.mean(x) # ~1.8 ms comb",https://github.com/pytorch/pytorch/pull/181308,496afb609aac1acf7cff4360bceabea0dbfa3423513c7b7b85b6344d5f96666e review guidance,pr,181308,pr,43858,high,pr.reviews[0].body,"This looks okay to me, but it's been a long time since I looked at this code. @albanD does this look okay to you too?",https://github.com/pytorch/pytorch/pull/181308,3faf474e8d84df9d95e1a3cc8098f334947b7feede54547c7c0d597674db3d1a review guidance,pr,181308,issue,122191,high,pr.reviews[0].body,"This looks okay to me, but it's been a long time since I looked at this code. @albanD does this look okay to you too?",https://github.com/pytorch/pytorch/pull/181308,e99869a17ae9977e127283ed5bc1a8c7f079b802f3a9d20e8b03a5b857601c06 review guidance,pr,187236,pr,187236,high,pr.reviews[0].body,a couple comments,https://github.com/pytorch/pytorch/pull/187236,314e93b2425275216258940f4a7edde124bb67a60d9f0a2e56592c52b2905956 closes,pr,186506,issue,117555,high,pr.body,loads through .decompose() so the direct py_impl registrations accept schema kwargs such as output/is_target and out0/out1/out2/out3. Fixes #117555 Generated by my agent Test Plan: python test/test_decomp.py HasDecompTest.test_composite_implicit_autograd_out_decompositions_acc...,https://github.com/pytorch/pytorch/pull/186506,a55a8ea473b6412d02ed2ad4cbd0e9820c65d8d22134c433cafb7156d36e12da closes,pr,186507,issue,117265,high,pr.body,"al through RemovableHandle. Watching the hook dictionaries directly keeps the invalidation tied to the state Dynamo actually guarded. Fixes #117265 Generated by my agent Benchmark Results: Command: python - <<'PY' ... PY microbenchmark for first-call torch.compile time and 20,...",https://github.com/pytorch/pytorch/pull/186507,e1742975c4e1e2f5ec1a77b5c00d4d8dd9cc02853338bf2ec23f75a1118aefaa references,pr,186507,issue,117265,medium,pr.comments[2].body,"`torch/_dynamo/guards.py` changes - [x] Review test files - [x] Post review feedback --- ### Summary This PR fixes a real correctness bug (#117265) where Dynamo would reuse compiled graphs after module hooks are registered or removed, because the in-place mutation of hook `Ord...",https://github.com/pytorch/pytorch/pull/186507,e9a3e6694dafa6639fef9ed03afd144e8fb191099ef58a006f0fecf104929bdd references,pr,186507,issue,117265,medium,pr.comments[6].body,eview `torch/_dynamo/guards.py` changes - [x] Review test files - [x] Post review feedback --- ### Summary This PR fixes a correctness bug (#117265) where Dynamo reuses compiled graphs after module hooks are registered/removed in-place. The fix introduces explicit Dynamo cache...,https://github.com/pytorch/pytorch/pull/186507,231c7e01964b06da3822f5e7e2c172bc1e56108e916f91494c0bd09d90955d41 review guidance,pr,186507,issue,117265,high,pr.reviews[0].body,"based on animesh's comment on the issue, it sounds like we just need to flip a few configs on for this to work. Was that not the case?",https://github.com/pytorch/pytorch/pull/186507,3a7ab20cb7fa0b7632771e8d10cd4342cab20311da40d7bfe68169792d3b408b review guidance,pr,186507,pr,186507,high,pr.reviews[0].body,"based on animesh's comment on the issue, it sounds like we just need to flip a few configs on for this to work. Was that not the case?",https://github.com/pytorch/pytorch/pull/186507,af3b103a5a905c492c283d592e7baea85fee4ff63fb53c05524afd381edb02d1 review guidance,pr,186507,issue,117265,high,pr.reviews[1].body,"based on animesh's comment on the issue, it sounds like we just need to flip a few configs on for this to work. Was that not the case?",https://github.com/pytorch/pytorch/pull/186507#pullrequestreview-4500494926,54e21e75fcdef994cf25b00cc78590bafbf1bd50bb087b865c2bfe1c5061ee64 review guidance,pr,186507,pr,186507,high,pr.reviews[1].body,"based on animesh's comment on the issue, it sounds like we just need to flip a few configs on for this to work. Was that not the case?",https://github.com/pytorch/pytorch/pull/186507#pullrequestreview-4500494926,a791f6ae916db906ba091f076d708f50f36f541bc2b78b6a082108fdc256f12d closes,pr,185976,issue,145231,high,pr.body,"The affected models now hit an eager_two_runs_differ baseline in those benchmark jobs, while the graph-break counts improved to zero. Fixes #145231 Generated by my agent Benchmark Results: 20 iterations of torch.compile(..., backend=""eager"") first call on a small custom autogr...",https://github.com/pytorch/pytorch/pull/185976,bc025b0441f6f3e5ab6c0b9cc1062b64e71969e8f95a1afdf2ddd215afa1b9f0 references,pr,185976,issue,186031,medium,pr.comments[6].body,"My agent says I retried the TorchTitan failure and it failed again in the known torchcomms/process-group issue tracked by #186031, outside this PR's custom autograd.Function backward tracing change. I also updated the PR description to document the intentional torch.is",https://github.com/pytorch/pytorch/pull/185976,dc1f6030b96042426a7cbba0ef9fa742d557dff19c79954d98dddb736711b119 review guidance,pr,185976,issue,145231,high,pr.reviews[0].body,"It took me a while to understand this, but here's what I think is happening: this is the case where torch.autograd.grad(create_graph=True) is captured into the graph if so, then we shouldn't emit the torch.no_grad into the autograd.function backward Is this what the PR is actually doing? If so, p...",https://github.com/pytorch/pytorch/pull/185976,70155c90c3c358467567e1016fce3af8c9b3669f15122ebdb92244c3ff4f6541 review guidance,pr,185976,pr,185976,high,pr.reviews[0].body,"It took me a while to understand this, but here's what I think is happening: this is the case where torch.autograd.grad(create_graph=True) is captured into the graph if so, then we shouldn't emit the torch.no_grad into the autograd.function backward Is this what the PR is actually doing? If so, p...",https://github.com/pytorch/pytorch/pull/185976,487d41af9dd6053932a184e11a76bef100ce524ed2f3b81b0418b672d1a32637 review guidance,pr,185976,issue,145231,high,pr.reviews[1].body,"My agent says I updated the PR body to spell out the flaky-test root cause: gradgradcheck calls torch.autograd.grad(..., create_graph=True), and the old generated backward graph embedded torch._C._set_grad_enabled(False), detaching NonDetFunc.backward outputs under create_graph=True. I also pushe...",https://github.com/pytorch/pytorch/pull/185976,0f63258501565e50ff8534d85763489e401a09b9a0dd498ecddb9e0da9156a86 review guidance,pr,185976,pr,185976,high,pr.reviews[1].body,"My agent says I updated the PR body to spell out the flaky-test root cause: gradgradcheck calls torch.autograd.grad(..., create_graph=True), and the old generated backward graph embedded torch._C._set_grad_enabled(False), detaching NonDetFunc.backward outputs under create_graph=True. I also pushe...",https://github.com/pytorch/pytorch/pull/185976,b38930e443b4c2bd4e010964b5aead019e13fa39599758011022c8ad267ef09c review guidance,pr,185976,issue,145231,high,pr.reviews[2].body,"It took me a while to understand this, but here's what I think is happening: - this is the case where torch.autograd.grad(create_graph=True) is captured into the graph - if so, then we shouldn't emit the torch.no_grad into the autograd.function backward Is this what the PR is actually doing? If s...",https://github.com/pytorch/pytorch/pull/185976#pullrequestreview-4500212249,0cb0a60d46b746e20fbd4ebc8b0b46201d436ed58ec01f660addca9f70babd3c review guidance,pr,185976,pr,185976,high,pr.reviews[2].body,"It took me a while to understand this, but here's what I think is happening: - this is the case where torch.autograd.grad(create_graph=True) is captured into the graph - if so, then we shouldn't emit the torch.no_grad into the autograd.function backward Is this what the PR is actually doing? If s...",https://github.com/pytorch/pytorch/pull/185976#pullrequestreview-4500212249,9ee20b35eca6bf565266dbacf97a4a680fd9a6d709872c6b767ccd96d974881d review guidance,pr,185976,issue,145231,high,pr.reviews[3].body,"My agent says I updated the PR body to spell out the flaky-test root cause: gradgradcheck calls torch.autograd.grad(..., create_graph=True), and the old generated backward graph embedded torch._C._set_grad_enabled(False), detaching NonDetFunc.backward outputs under create_graph=True. I also pushe...",https://github.com/pytorch/pytorch/pull/185976#pullrequestreview-4515200216,39ebbd2fa2e3a6998c37e3881bcb72e487eb6bc6eaf188b19209ec7bffafd426 review guidance,pr,185976,pr,185976,high,pr.reviews[3].body,"My agent says I updated the PR body to spell out the flaky-test root cause: gradgradcheck calls torch.autograd.grad(..., create_graph=True), and the old generated backward graph embedded torch._C._set_grad_enabled(False), detaching NonDetFunc.backward outputs under create_graph=True. I also pushe...",https://github.com/pytorch/pytorch/pull/185976#pullrequestreview-4515200216,cd645239a8c3730d141f92c6f1646a4db8f66a0b9a4bf559f2d5defee5d3f4b1 closes,pr,186510,issue,116202,high,pr.body,"ses fixed, remove the Dynamo-specific Adagrad and RMSprop tolerance overrides and add CUDA low-precision addc decomposition coverage. Fixes #116202 Generated by my agent Test Plan: CUDA_VISIBLE_DEVICES=4 python test/dynamo/test_dynamo_decompositions.py -k addc_ops_low_precisio...",https://github.com/pytorch/pytorch/pull/186510,acc9562148b6813e694de184772a5388facea18af35f8ac35f08ed7655acfaa5 references,pr,186510,issue,116202,medium,pr.comments[2].body,"testing/_internal/common_optimizers.py` The tolerance overrides being removed were masking the underlying precision issue (referenced issue #116202). Now that the root cause is fixed, removing them is the right call — it ensures regressions would be caught. ### Summary The fix...",https://github.com/pytorch/pytorch/pull/186510,af7363a4f7341c5565f756093fcce7d3bcd7fc9426a1afbdf04474fdbcd79277 references,pr,186510,issue,116202,medium,pr.comments[4].body,"Removing the two `test_mixed_device_dtype` dynamo tolerance overrides (Adagrad + RMSprop) is the correct follow-through — they were masking #116202, and the test plan confirms the tests now pass under `PYTORCH_TEST_WITH_DYNAMO=1` without them, so regressions will be caught. ##...",https://github.com/pytorch/pytorch/pull/186510,154c5e5efacd4a3a595640b3d7ecafdd7949b1f88002a5283c67302286764c57 closes,pr,186512,issue,115484,high,pr.body,arameters; this patch chooses the narrower correctness boundary instead of trying to encode unstable parameter metadata in the graph. Fixes #115484 Generated by my agent Benchmark Results: Same-shape Parameter.data assignment compile+first-run benchmark with backend=eager over...,https://github.com/pytorch/pytorch/pull/186512,f9215b864a3633665e3e62a384b8526d535290ed73bcfaf72e78194996fa250b closes,pr,186517,issue,114415,high,pr.body,"ugh the subclass wrapper path. This patch keeps the behavior conservative and matches the issue request to at least detect and raise. Fixes #114415 Generated by my agent Test Plan: Reproduced the original issue before the fix with a minimized torch.compile(backend=""aot_eager"",...",https://github.com/pytorch/pytorch/pull/186517,6c1a4b4b9286ba84c0e89689a9da61d0d4a4bdcdfb5f511ee4448c7fbaa50e01 closes,pr,186520,issue,114296,high,pr.body,"TAutograd partitioner / Inductor backward SymInt binding propagation, not fx_minifier input pruning or after_aot repro serialization. Fixes #114296 Generated by my agent Test Plan: python test/functorch/test_minifier.py -v python test/dynamo/test_after_aot.py -k test_save_grap...",https://github.com/pytorch/pytorch/pull/186520,e079e8de84e32b904e1cd381ea874fcfe93f028f0e235151c11ea6f9df2e9bbc closes,pr,184432,issue,113809,high,pr.body,r runtime entry to avoid concurrent pool recording from compiled backward. Add a regression test for overlapping cudagraph tree runs. Fixes #113809 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/184432,d7592d303f1be7dd28ef43d9aee2660a27dd5a3f2e22b2cfc15592cca1ecd5a9 closes,pr,184436,issue,112788,high,pr.body,"names in the cached source passed to async_compile.triton, while preserving the descriptive name for profiling and runtime metadata. Fixes #112788 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...",https://github.com/pytorch/pytorch/pull/184436,5b1a81bdfc3e23009921c20959654335a0f4ed0989f09ef406272accb1c670a8 closes,pr,186529,issue,111414,high,pr.body,"le_grad() contexts. Nested compilation reads the already-preserved user state so compiler-internal state changes do not overwrite it. Fixes #111414 Generated by my agent Benchmark Results: Command: N=200 repeated torch._dynamo.reset(); torch.compile(fn, backend='eager')(x). ma...",https://github.com/pytorch/pytorch/pull/186529,b392c214a1a0dac9a2f37ac15b9b5dfe72fdd4c526c6568a6f37ce18317dc7a9 closes,pr,186530,issue,111385,high,pr.body,"utocast global state in the generated autocast sequence, and the local handler matches the existing patterns for enabled/cache state. Fixes #111385 Generated by my agent Test Plan: python test/dynamo/test_ctx_manager.py -k test__enter__exit_autocast_graph_break_explicit_dtype...",https://github.com/pytorch/pytorch/pull/186530,9cf5091292a306d3ba4080b5af80212af5b457b4e4de9caa535b45d3777e8986 closes,pr,186531,issue,111320,high,pr.body,ort. This keeps the fix scoped to the remaining issue: the test was still being skipped after the underlying OOM stopped reproducing. Fixes #111320 Generated by my agent Test Plan: PYTORCH_TEST_WITH_DYNAMO=1 python test/nn/test_pooling.py TestPoolingNNDeviceTypeCPU.test_max_po...,https://github.com/pytorch/pytorch/pull/186531,3dd632298e0228ce243d9a808c02257d89fd0de82a7dc1efe37563d6c6e5e8ee closes,pr,186533,issue,106135,high,pr.body,schema case and the fallback path that calls an int[] kernel taking const std::vector& through a SymInt-vector typed handle. Fixes #106135 Generated by my agent Test Plan: cmake --build build --target op_registration_test -j 8 build/bin/op_registration_test --gtest_fi...,https://github.com/pytorch/pytorch/pull/186533,c7ce07febc6c203981d72c8d78569106b70d7aef33f29cff5f11161ee02871f0 closes,pr,186534,issue,105768,high,pr.body,"sing a read-count heuristic. The old check only considered reads for one expansion of the expression. For residual blocks like the repro in #105768, a four-read residual add with two downstream users stayed just under the threshold, so it was inlined into both users. That kept...",https://github.com/pytorch/pytorch/pull/186534,18590ebb4986373be7bad002aac9bdbb05559fe90ed38f9a32768b3a2e4ca112 closes,pr,187601,issue,177750,high,pr.body,unner -a (fails only in existing CLANGTIDY whole-file reports in touched C++ files plus unrelated torch/custom_class.h include-cycle) Fixes #177750 Generated by my agent,https://github.com/pytorch/pytorch/pull/187601,c6defde682516aa261fceb207b57f487dfb3cd2d541e2e1b7aa0c3599df2c5ab closes,pr,186538,issue,103417,high,pr.body,"relu() x = torch.randn(16, device='cuda') y = f(x) torch.cuda.synchronize() print(torch.allclose(y, (x + 1).relu())) PY lintrunner -a Fixes #103417 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...",https://github.com/pytorch/pytorch/pull/186538,479026902e29fbd4afddb144f3645836cf38522f922ad7eb85ab0a28f2198983 closes,pr,189224,issue,187810,high,pr.closingIssuesReferences,pr #189224 declares a closing reference to issue #187810.,https://github.com/pytorch/pytorch/pull/189224,efbac98e40edcce96c01c8329538b11dc9fae6e178774eee4ba4a34fa288d826 closes,pr,189224,issue,187811,high,pr.closingIssuesReferences,pr #189224 declares a closing reference to issue #187811.,https://github.com/pytorch/pytorch/pull/189224,49a165237cd7173a4bd2d54a8b53fde88caa764ca58862efb622b20d17782c2a closes,pr,189224,issue,187810,high,pr.body,Fixes #187810 Fixes #187811 Summary Skip test_combo_kernel_dynamic_scale_rblock on XPU for both ComboKernelTests and ComboKernelTestsPerSubkernelBlocks (,https://github.com/pytorch/pytorch/pull/189224,50cad4bd82febe914b5337cddbb529829388ac014873231333f33336c1a6efe9 closes,pr,189224,issue,187811,high,pr.body,Fixes #187810 Fixes #187811 Summary Skip test_combo_kernel_dynamic_scale_rblock on XPU for both ComboKernelTests and ComboKernelTestsPerSubkernelBlocks (inherits the s,https://github.com/pytorch/pytorch/pull/189224,8eaf1f98484db71a48196fac71874f6ba921ddc64eb5a8544a4466afdf8afb2a closes,pr,187354,issue,187332,high,pr.body,ile continuing to use libdevice for finite inputs. This preserves the existing eager semantics without changing the finite fast path. Fixes #187332 Generated by my agent Test Plan: Reproduced the original CUDA mismatch before the fix. Verified CPU/CUDA eager and SciPy return N...,https://github.com/pytorch/pytorch/pull/187354,06a635ab09d94f86f49263445bcc20becdf8f67a9a344c7ae215b0746ab95229 closes,pr,186539,issue,102839,high,pr.body,urrent Dynamo and measures node growth relative to the loop entry so unrelated pre-loop graph nodes are not counted against the loop. Fixes #102839 Generated by my agent Benchmark Results: Repro before fix on starting main: n=500 loop compiled into 1 graph with 502 FX nodes. A...,https://github.com/pytorch/pytorch/pull/186539,dea4a8bca94058a53dd779bd89d399c9f294673cb05dbe048017c5a8dcb100bf closes,pr,186361,issue,130182,high,pr.body,h could show an input like y even though that value is actually grad_y. This made the generated graph hard to read and is the root cause of #130182. Fix this by tagging only the placeholders that correspond to gradients of forward outputs with a grad_ prefix when the existing...,https://github.com/pytorch/pytorch/pull/186361,1912b608b77554f8751065ad536f03a2e89e2069ffa72f6ea54180e1a06fc637 references,pr,186361,issue,186031,medium,pr.comments[4].body,My agent says the repeated torchtitan_features_integration failure is the known TorchTitan torchcomms process-group lookup issue tracked in #186031 (`RuntimeError: Could not resolve the process group registered under the name 20`). This PR only changes Dynamo autograd.Function...,https://github.com/pytorch/pytorch/pull/186361,83ab59405948bd39f82ba41395fb5025559690ab0b5bcba4b9a6cace4d00d960 closes,pr,184345,issue,125236,high,pr.body,grad entered inside a compiled region still trigger freezing optimizations while preserving autograd compilation for training graphs. Fixes #125236 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/184345,99f69e275134f06190b9cdd89136ee69e871911e2f078bedafd79e0472bb037e references,pr,184345,issue,125236,medium,pr.comments[2].body,"est/inductor/test_inductor_freezing.py` The test (`test_folded_conv_bn_with_no_grad_inside_compiled_region`) directly validates the fix for #125236 — grad is enabled at the outer scope, `no_grad` is entered inside `@torch.compile`, and the test asserts that freezing still kick...",https://github.com/pytorch/pytorch/pull/184345,0bc3054dd210d74dc7f60c9b7b3bccaa917aed8b23ab7550f324719bab7d0a8c references,pr,184345,issue,125236,medium,pr.comments[4].body,"euristic that is now replaced by a more precise check (does the graph *actually* need autograd?), which is exactly what enables the fix for #125236 (no_grad inside compiled region should still trigger freezing). ---",https://github.com/pytorch/pytorch/pull/184345,69b4707d76104fedc5e7a1aff2df5c77d45471d3d50db51b5ba0d395f0b30a5e references,pr,184345,issue,125236,medium,pr.comments[7].body,"Post review feedback --- **Summary:** This is a clean, minimal 3-line change (+ test) that correctly addresses the original user story from #125236. The approach is sound and the reviewer concerns have been properly addressed. #### `torch/_inductor/compile_fx.py` — Removing `n...",https://github.com/pytorch/pytorch/pull/184345,aa2d8a809b875e3e4855d5908591084866ba04062887bf96e6a38ffed33869d4 review guidance,pr,184345,issue,125236,high,pr.reviews[1].body,"@claude This doesn't seem right. Freezing is supposed to respect the user specified grad enabled or not. But now, having freezing on seems to force no grad now.",https://github.com/pytorch/pytorch/pull/184345,ad1ac4380f9f97711aac0484f5c81b11448b8748655cb536b0770ccca2adb798 review guidance,pr,184345,pr,184345,high,pr.reviews[1].body,"@claude This doesn't seem right. Freezing is supposed to respect the user specified grad enabled or not. But now, having freezing on seems to force no grad now.",https://github.com/pytorch/pytorch/pull/184345,82924465baa92916178489b18febb57078f9c2a3e457053caeab6e43b8a1c81d review guidance,pr,184345,issue,125236,high,pr.reviews[3].body,"@claude This doesn't seem right. Freezing is supposed to respect the user specified grad enabled or not. But now, having freezing on seems to force no grad now.",https://github.com/pytorch/pytorch/pull/184345#pullrequestreview-4340643502,428954b924088b5f83fb42263cfe72d5cc20ff46ec50c6ae527aa3b6497ae4aa review guidance,pr,184345,pr,184345,high,pr.reviews[3].body,"@claude This doesn't seem right. Freezing is supposed to respect the user specified grad enabled or not. But now, having freezing on seems to force no grad now.",https://github.com/pytorch/pytorch/pull/184345#pullrequestreview-4340643502,9eede189fd09aa87617035e2811d3bd7aaa9fd71be4ef77d7db2b2a63083a32c review guidance,pr,184345,issue,125236,high,pr.reviews[5].body,"My agent says I addressed the grad-mode concern by removing the torch.no_grad() wrapper around the freezing compiler. AOTAutograd still decides when to call the freezing compiler, and the grad-enabled failure is fixed locally by detaching erased meta tensors before re-registering them as parameters.",https://github.com/pytorch/pytorch/pull/184345#pullrequestreview-4342023249,d3cb9c18c2728cc38166b58bb0cbb232bcaf66762caa17d4514739543f4748d4 review guidance,pr,184345,pr,184345,high,pr.reviews[5].body,"My agent says I addressed the grad-mode concern by removing the torch.no_grad() wrapper around the freezing compiler. AOTAutograd still decides when to call the freezing compiler, and the grad-enabled failure is fixed locally by detaching erased meta tensors before re-registering them as parameters.",https://github.com/pytorch/pytorch/pull/184345#pullrequestreview-4342023249,3516ec31ae596cd6e2fe83d7b855265ff821c11f5a042cfce515939677dd1346 review guidance,pr,184345,issue,125236,high,pr.reviews[6].body,needs to be opt in bc memory/semantics change,https://github.com/pytorch/pytorch/pull/184345#pullrequestreview-4342345061,7053ad138d1f4d4ce28a4c734ab09cebe77e831c556c1da32d410680519679a5 review guidance,pr,184345,pr,184345,high,pr.reviews[6].body,needs to be opt in bc memory/semantics change,https://github.com/pytorch/pytorch/pull/184345#pullrequestreview-4342345061,ce1ac56eded6b43e93fbf85689df0529142cc23a39d7a69413a32ed91048eadf review guidance,pr,184345,issue,125236,high,pr.reviews[7].body,"My agent says I addressed the opt-in concern by adding torch._inductor.config.freezing_traced_inference, defaulting it off, and preserving the legacy ambient-grad gate unless that flag is enabled. The regression now checks both default behavior and the opt-in traced-inference freezing path.",https://github.com/pytorch/pytorch/pull/184345#pullrequestreview-4348421870,4e29cf60986461b9a8841311f524133eba90122c70f76dfe0c070348f065907e review guidance,pr,184345,pr,184345,high,pr.reviews[7].body,"My agent says I addressed the opt-in concern by adding torch._inductor.config.freezing_traced_inference, defaulting it off, and preserving the legacy ambient-grad gate unless that flag is enabled. The regression now checks both default behavior and the opt-in traced-inference freezing path.",https://github.com/pytorch/pytorch/pull/184345#pullrequestreview-4348421870,c8705cfb4d96fad702a89e8735f2c5333b5de93c9d267232b9168b4b630d2db2 review guidance,pr,184345,issue,125236,high,pr.reviews[8].body,I feel the PR and the reviewers have lost the plot of the original issues that requested the change. It seems unlikely that a new config flag actually solves any real problem. Let's not invent new config flags for fake problems. Please reread https://github.com/pytorch/pytorch/issues/125236 and h...,https://github.com/pytorch/pytorch/pull/184345#pullrequestreview-4432476185,5b743b1af82b3eb4e1616b19c36badd59034f16f5838e2b34c6991e4a92cff58 review guidance,pr,184345,pr,184345,high,pr.reviews[8].body,I feel the PR and the reviewers have lost the plot of the original issues that requested the change. It seems unlikely that a new config flag actually solves any real problem. Let's not invent new config flags for fake problems. Please reread https://github.com/pytorch/pytorch/issues/125236 and h...,https://github.com/pytorch/pytorch/pull/184345#pullrequestreview-4432476185,0b88222274c6bb9c4891e080300ef371ff77c02324a0fa453bb36027c87db02a review guidance,pr,184345,issue,125236,high,pr.reviews[9].body,"My agent says I reread the original issue and linked reproducer and removed the new opt-in config. The PR now fixes the original user story directly: when Inductor freezing is enabled, AOTAutograd receives the freezing compiler as its inference compiler and decides after tracing whether the graph...",https://github.com/pytorch/pytorch/pull/184345#pullrequestreview-4443749351,e44efd1aa4f2b0abb450d13658fe13da8d7456f5cdb68dd71aa1e66f4ccc7f42 review guidance,pr,184345,pr,184345,high,pr.reviews[9].body,"My agent says I reread the original issue and linked reproducer and removed the new opt-in config. The PR now fixes the original user story directly: when Inductor freezing is enabled, AOTAutograd receives the freezing compiler as its inference compiler and decides after tracing whether the graph...",https://github.com/pytorch/pytorch/pull/184345#pullrequestreview-4443749351,f5d4579ed6dd410e701ef8306c9c5c3e7357ad4135fabd8b6767bdab02791c3d closes,pr,187592,issue,134385,high,pr.body,because the root cause is not regular unsupported custom operators; it is registered HOPs being bypassed by the HOP dispatch branch. Fixes #134385 Generated by my agent Test Plan: python test/test_flop_counter.py TestFlopCounter.test_registered_hop TestFlexAttentionEstimation....,https://github.com/pytorch/pytorch/pull/187592,e98bd8790c137913b38b15e6f8491d1c02f5eb874435cb5919eb81e893f9ec87 closes,pr,185304,issue,160939,high,pr.body,ignificant while_loop stride requirements enforced while allowing irrelevant size-1 stride differences during fake metadata checking. Fixes #160939 Generated by my agent Test Plan: python test/test_fake_tensor.py FakeTensorTest.test_like_preserve_format_preserves_dense_strides...,https://github.com/pytorch/pytorch/pull/185304,db8ea2f690220cff81323a41eb2f6c4dcb1fd5a8f304febad13aa3db7821b6de references,pr,188720,issue,181647,medium,pr.comments[0].body,"o flakiness on trunk: rocm-nightly / linux-noble-rocm-nightly-py3.12-gfx942 / test (default, 1, 6, linux.rocm.gpu.gfx942.1, unstable) (gh) (#181647) test/test_torch.py::TestTorch::test_cxx_flags rocm-nightly / linux-noble-rocm-nightly-py3.12-gfx942 / test (default, 2, 6, linux...",https://github.com/pytorch/pytorch/pull/188720,ae0dd626b9e44676ef8b0fae6d3d7dc77655a952b62894b29fb868bc8798dc42 references,pr,188720,issue,181648,medium,pr.comments[0].body,"able) (gh) (#181649) rocm-nightly / linux-noble-rocm-nightly-py3.12-gfx942 / test (inductor, 1, 2, linux.rocm.gpu.gfx942.1, unstable) (gh) (#181648) test/test_torch.py::TestTorch::test_cxx_flags rocm-nightly / linux-noble-rocm-nightly-py3.12-gfx942 / test (inductor, 2, 2, linu...",https://github.com/pytorch/pytorch/pull/188720,b6ac050399983c71091dc830af3dd9e710bb295322d32320320eebe04616d0b1 references,pr,188720,issue,181649,medium,pr.comments[0].body,"linux-noble-rocm-nightly-py3.12-gfx942 / test (distributed, 1, 3, linux.rocm.gpu.gfx942.4, module:rocm, oncall:distributed, unstable) (gh) (#181649) rocm-nightly / linux-noble-rocm-nightly-py3.12-gfx942 / test (distributed, 2, 3, linux.rocm.gpu.gfx942.4, module:rocm, oncall:di...",https://github.com/pytorch/pytorch/pull/188720,691c6f4dffa385a6d980de7d27369aec01faa30794446d69d441b88b9fde1454 closes,pr,186091,issue,140845,high,pr.body,d export path uses that overload; packed-sequence aten.lstm.data currently hits a separate export/decomposition issue before compile. Fixes #140845 Generated by my agent Test Plan: python -m py_compile torch/_inductor/compile_fx.py torch/_inductor/lowering.py test/inductor/tes...,https://github.com/pytorch/pytorch/pull/186091,2f322328a2731298fbffc7ab2dbd54c9d8933a2eaded6b2e637f88f0210529b4 closes,pr,185280,issue,161119,high,pr.body,"t the root issue is the stale int specialization, so fixing the generic automatic-dynamic handoff covers other size-int uses as well. Fixes #161119 Generated by my agent Benchmark Results: Minimal CPU Dynamo repro mirroring torch.split(..., [0, changing_global_int]) over sizes...",https://github.com/pytorch/pytorch/pull/185280,361d2d0040af04385fcaf0a7795a626b1c009dbe3751c4d86c622ce3769d84a7 closes,pr,188049,issue,188048,high,pr.closingIssuesReferences,pr #188049 declares a closing reference to issue #188048.,https://github.com/pytorch/pytorch/pull/188049,5608d590d7171a301631d468a8d8e918d4b45063425797f5d7758feb8ff07329 closes,pr,188049,issue,188048,high,pr.body,"idends become NaN, and finite nonzero dividends divided by opposite-signed infinities become -1. Integer floor division is unchanged. Fixes #188048 Test Plan: python test/inductor/test_torchinductor.py -k test_div_floor_float_nonfinite_cuda lintrunner -a torch/_inductor/loweri...",https://github.com/pytorch/pytorch/pull/188049,b92871106dffcd66bd007513f8bc1e11163fec263a2426ae6d7bbd5490f04c47 review guidance,pr,188049,issue,188048,high,pr.reviews[0].body,address automated feedback from claude,https://github.com/pytorch/pytorch/pull/188049,1abe3c2882831c8dc1b08e1a3139a4898d684f4e6d8f3f2c07e9e1f195bd5158 closes,pr,186546,issue,97750,high,pr.body,-raise the original exception so BackendCompilerFailed reports the useful root error. Add focused regression coverage for both paths. Fixes #97750 Generated by my agent Test Plan: python test/dynamo/test_minifier.py TestDynamoMinifierBackend python test/dynamo/test_minifier.py...,https://github.com/pytorch/pytorch/pull/186546,d0ca7c467a5f67034ec289d0f0ca7851fce185629a35e58d25c2df62d8ae716d closes,pr,186549,issue,97078,high,pr.body,ustom path to module mapping subclasses fixes the root cause directly while preserving normal deepcopy semantics for the common case. Fixes #97078 Benchmark Results: Command: compared previous helper behavior (copy.deepcopy under FakeCopyMode) against patched deepcopy_to_fake_...,https://github.com/pytorch/pytorch/pull/186549,c0609583181e375925fde9a0f6780dfa60d93cb2417951803ba41fe5178b4152 closes,pr,186553,issue,93501,high,pr.body,"ing the public wrapper skip only, but that leaves direct lower-op calls and _VF tracing paths with the same fake metadata root cause. Fixes #93501 Generated by my agent Test Plan: python test/dynamo/test_misc.py MiscTests.test_pack_padded_sequence_graph_breaks MiscTests.test_p...",https://github.com/pytorch/pytorch/pull/186553,81b5839e16489dfe5ecf5e406134b9aceebeca969746e70925617aa447725ff2 references,pr,189253,pr,186918,medium,pr.body,"he UCC backend and >=2 GPUs. Backend-specific, not device-generic. Test Plan: python test/distributed/test_c10d_spawn_ucc.py -v Depends on: #186918 @mansiag05 @vishalgoyal316 @RiyaP2508",https://github.com/pytorch/pytorch/pull/189253,77d80a701767179b42e97dd8055217adba798947e69aff2db557be68f9f51a00 closes,pr,186551,issue,93367,high,pr.body,he existing placeholder metadata keeps the repro faithful to the live graph while preserving real input bytes for accuracy debugging. Fixes #93367 Generated by my agent Test Plan: python test/dynamo/test_after_aot.py TestAfterAot.test_get_compile_args_symbolic_tracing TestAfte...,https://github.com/pytorch/pytorch/pull/186551,d0b8defe1b6a2c35cc28ec7855311d8e80f07cb2130f7b6204def81d6b710514 review guidance,pr,186551,issue,93367,high,pr.reviews[0].body,"cc @claude please review this. Also, can you add tests ?",https://github.com/pytorch/pytorch/pull/186551,34a93515b6056b2d9175645ac923d6746ec3b2db390d9489b5e527674c20a57a review guidance,pr,186551,pr,186551,high,pr.reviews[0].body,"cc @claude please review this. Also, can you add tests ?",https://github.com/pytorch/pytorch/pull/186551,79995d4265c427880d767c2a10f4d07f683156d72cc23b21282ffb0b9b958f90 closes,pr,186561,issue,90507,high,pr.body,"in real differentiable views bypass that shortcut and use the intermediate alias reconstruction path, preserving their view metadata. Fixes #90507 Generated by my agent Test Plan: ninja -C build torch_python python repro for issue #90507 before/after; after fix compiled output...",https://github.com/pytorch/pytorch/pull/186561,3e61f8779700f8f14b6037fd9cd247fcc76d53baef4ea54d58b3f77a393c8e92 closes,pr,187057,issue,187041,high,pr.body,egression test covers a value-opaque object stored in a dict and returned as traceable wrapper subclass metadata under torch.compile. Fixes #187041 Generated by my agent Test Plan: python test/test_opaque_obj_v2.py -k test_value_type_graph_output_subclass_metadata_side_effect...,https://github.com/pytorch/pytorch/pull/187057,5f7258ca0b15a31119f75fbe4fb5bb75e7228923a22dc730054d1f997b231249 closes,pr,186435,issue,127111,high,pr.body,"nCPU kernel, because that broader approach accepted eager-invalid mixed-layout/scalar programs and produced incoherent fake metadata. Fixes #127111 Generated by my agent Test Plan: python test/test_fake_tensor.py -k mkldnn python -m py_compile torch/_subclasses/fake_tensor.py...",https://github.com/pytorch/pytorch/pull/186435,d1c8e7351955af391932a20d478a302dcd5f34c6c9bb83a36eb8ce31b8d55328 closes,pr,186397,issue,129327,high,pr.body,"ndle literal bool predicates explicitly so torch._check(False, ...) is preserved rather than passed through symbolic metadata lookup. Fixes #129327 Generated by my agent Test Plan: python test/dynamo/test_misc.py -k test_check_raises_at_runtime_when_recompile_predicate_false_c...",https://github.com/pytorch/pytorch/pull/186397,834d1a4010fe9e64e58bce6813f379fb4df536562e72f0cc91535bc79f31740b review guidance,pr,186397,issue,129327,high,pr.reviews[0].body,Rebase - we recently made some fixes to torch._check handling in Dynamo,https://github.com/pytorch/pytorch/pull/186397,893ff6bafcc0edf4943fcf8d99877f4654a567ca0d6a61ff4bfc12fde6ee3e41 review guidance,pr,186397,pr,186397,high,pr.reviews[0].body,Rebase - we recently made some fixes to torch._check handling in Dynamo,https://github.com/pytorch/pytorch/pull/186397,ca001173ba544da26e37c0b2297416e03d697aa49aa4133fa65aa9a788c70065 review guidance,pr,186397,issue,129327,high,pr.reviews[1].body,"My agent says I rebased the PR onto current main as requested and adapted the constant-false torch._check path to the newer Dynamo message handling. The rebased diff is now just the Dynamo runtime-check behavior plus the regression test; local validation passed with the focused regression, broade...",https://github.com/pytorch/pytorch/pull/186397,8c3a06e143cad84fea7e76838335956b70f8c769ea1ed673d23a34b3fee6a650 review guidance,pr,186397,pr,186397,high,pr.reviews[1].body,"My agent says I rebased the PR onto current main as requested and adapted the constant-false torch._check path to the newer Dynamo message handling. The rebased diff is now just the Dynamo runtime-check behavior plus the regression test; local validation passed with the focused regression, broade...",https://github.com/pytorch/pytorch/pull/186397,97514aa08aa37b17fd08ed4e2ab51a130e50cf9a0858c3d466cbac0a9962167d closes,pr,186395,issue,186084,high,pr.body,py torch/_inductor/scheduler.py test/distributed/test_overlap_bucketing_unit.py test/distributed/test_aten_comm_compute_reordering.py Fixes #186084 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/186395,c7dcaaad2ae3385d465ceb20b467f04dd409fba9d05feeb4a1e4a268fe4c85fc review guidance,pr,186395,issue,186084,high,pr.reviews[0].body,I think #182419 is responsible for this. id rather revert and fix the underlying issues.,https://github.com/pytorch/pytorch/pull/186395,ea1fa55f87cf97362c8f2a0ac8b59226ddb7879c7593dbdcbb243190cba1f609 review guidance,pr,186395,pr,186395,high,pr.reviews[0].body,I think #182419 is responsible for this. id rather revert and fix the underlying issues.,https://github.com/pytorch/pytorch/pull/186395,ade2d341468c940b9a34b2678de2afa33c7701a6e6c7575f6c145e9bcf8a2956 closes,pr,186409,issue,128649,high,pr.body,ints. This keeps the change scoped to subclass view fakeification rather than weakening the general symbolic-shape replacement logic. Fixes #128649 Generated by my agent Test Plan: python test/dynamo/test_subclasses.py -k test_subclass_view git diff --check lintrunner -a --ski...,https://github.com/pytorch/pytorch/pull/186409,64d079a244b82ed3fba210711481095d8a52cba716b7baa7937ff3d0a864446d review guidance,pr,186409,issue,128649,high,pr.reviews[0].body,cc @laithsakka for dynamic shapes review,https://github.com/pytorch/pytorch/pull/186409,3d9bf51b460450a6c8fdf825837985798023ca76545467636cbcb043f08c6c2a review guidance,pr,186409,pr,186409,high,pr.reviews[0].body,cc @laithsakka for dynamic shapes review,https://github.com/pytorch/pytorch/pull/186409,d5e56b0085e1ee9564e91f11d9c557345b6778eb25558ff13078fbd92208a0a0 closes,pr,185861,issue,185630,high,pr.body,nto the stored metadata. Update the SerializableTuple example so future hand-written ViewMeta classes do not repeat the lifetime bug. Fixes #185630 Generated by my agent Test Plan: python test/test_functionalization.py TestFunctionalization.test_special_view_meta_pickle_roundt...,https://github.com/pytorch/pytorch/pull/185861,366120659fdd4b06444073e6ffc91817a5eb3c4dd71d929e96d84fa8ccc46a28 closes,pr,188053,issue,187759,high,pr.closingIssuesReferences,pr #188053 declares a closing reference to issue #187759.,https://github.com/pytorch/pytorch/pull/188053,0f88f9c032c5fc38c09ee23b39e25990357fcae45010ae1e2c00521634c0e604 closes,pr,188053,issue,187759,high,pr.body,"s a test verifying that torch.linalg.svdvals() correctly handles NaN inputs, fixing the inconsistency with torch.linalg.svd() documented in #187759. Problem svdvals() silently swallows NaN in some backend configs, returning finite singular values for a NaN matrix. This is a si...",https://github.com/pytorch/pytorch/pull/188053,5d5d62061ace6c04a705054891ac54fe74ed3a901f6662f9be5813283e7d6f05 closes,pr,186420,issue,128531,high,pr.body,"skipIf directly instead of swapping in NoTest, so local runs without torchrec report explicit skipped tests rather than NO TESTS RAN. Fixes #128531 Generated by my agent Test Plan: python test/dynamo/test_torchrec.py -v python test/dynamo/test_torchrec.py TorchRecTests.test_bu...",https://github.com/pytorch/pytorch/pull/186420,34c129297b19b406c3071784e9194960db89a98b95fb9e8f87ad480fe5d491c1 review guidance,pr,186420,issue,128531,high,pr.reviews[0].body,can we figure out if torchrec is still a thing that we care about? if it isn't then let's just delete the tests,https://github.com/pytorch/pytorch/pull/186420,526b806cd3d14917142624fa1554e1bd18d49ee288d32b19ed40576f957d4021 review guidance,pr,186420,pr,186420,high,pr.reviews[0].body,can we figure out if torchrec is still a thing that we care about? if it isn't then let's just delete the tests,https://github.com/pytorch/pytorch/pull/186420,139e283b1c7d49c9aa892c7d0113a844a4743c06a7d7a902cfff9b7e6f6543ef review guidance,pr,186420,issue,128531,high,pr.reviews[1].body,"This only triggers torchrec tests on Dynamo 3.11, 3.12, 3.13 tests. Dynamo tests run on the default shards of 3.10 and 3.14 seem to be excluded by this change. My claude suggests adding if [[ ""$TEST_CONFIG"" == 'default' ]] && ls dist/fbgemm_gpu/*.whl >/dev/null 2>&1; then # Only build jobs that o...",https://github.com/pytorch/pytorch/pull/186420,45806ae8de0329066725fefb282102b97c5dd09dd2c9f296eb95713cd64d550d review guidance,pr,186420,pr,186420,high,pr.reviews[1].body,"This only triggers torchrec tests on Dynamo 3.11, 3.12, 3.13 tests. Dynamo tests run on the default shards of 3.10 and 3.14 seem to be excluded by this change. My claude suggests adding if [[ ""$TEST_CONFIG"" == 'default' ]] && ls dist/fbgemm_gpu/*.whl >/dev/null 2>&1; then # Only build jobs that o...",https://github.com/pytorch/pytorch/pull/186420,e859cab295e51aae7af2ebc755a20fe172da14cb5a93f967d6004802a85c334d closes,pr,185871,issue,185510,high,pr.body,ntime was available for that failing case. Test Plan: python test/inductor/test_inductor_utils.py -k symbolic_stride_order Issue repro from #185510 before fix: reproduced InductorError wrapping TypeError: cannot determine truth value of Relational: 1 < s53 Same issue repro aft...,https://github.com/pytorch/pytorch/pull/185871,269af03f1efd26e4959e2d968c6fbc53b1dbb512aec02fba525ef97224ab3e24 review guidance,pr,185871,issue,185510,high,pr.reviews[0].body,LGTM,https://github.com/pytorch/pytorch/pull/185871,02cae811047b45eac4bcf49890ee193ffec10c2542a8aca9a4e5cfe3e617ba3b review guidance,pr,185871,pr,185871,high,pr.reviews[0].body,LGTM,https://github.com/pytorch/pytorch/pull/185871,fdc9416add476546738d350331d5d0d824fb6e700eb19aef3c61aeec8380c0c8 review guidance,pr,185871,issue,185510,high,pr.reviews[1].body,LGTM,https://github.com/pytorch/pytorch/pull/185871#pullrequestreview-4519999604,65360d1467abe45f3ef3956e9f13f2c6164590393c356071b17f0d173efc36d0 review guidance,pr,185871,pr,185871,high,pr.reviews[1].body,LGTM,https://github.com/pytorch/pytorch/pull/185871#pullrequestreview-4519999604,e28e596a70449960604badc1ef55d8874490333c587b161210698e5c431c2e88 closes,pr,185873,issue,185509,high,pr.body,"l2d backward or the specific fusion, but the dependency loss is scheduler-level state and should be fixed at the recomputation point. Fixes #185509 Generated by my agent Test Plan: Reproduced original issue on CPU before the fix: forward max diff 1.9073486328125e-06; gradient...",https://github.com/pytorch/pytorch/pull/185873,dc599d90e14ba01f2f4f33589c601ee30790b374381ed7315649b61c9200c11e closes,pr,185324,issue,160399,high,pr.body,"g generated-wrapper device assertions, but that only checks the already-bad compile-time devices and does not address the root cause. Fixes #160399 Generated by my agent Test Plan: python test/test_fake_tensor.py -k test_conv_rejects_mismatched_fake_devices python test/inducto...",https://github.com/pytorch/pytorch/pull/185324,035028d0fd1fc56765302ab5aa91e30e78da55fe6194278bbc44521a54890511 review guidance,pr,185324,issue,160399,high,pr.reviews[0].body,"I think this is more of a point fix for the github issue. Should the proper behavior be that when the model model is moved to cuda, we should do a recompilation? Should we have a generic device check?",https://github.com/pytorch/pytorch/pull/185324,29048c2fa2824023b4272c71bf8c4ab8a32da6506ec562f9e4d990b1015d5bfe review guidance,pr,185324,pr,185324,high,pr.reviews[0].body,"I think this is more of a point fix for the github issue. Should the proper behavior be that when the model model is moved to cuda, we should do a recompilation? Should we have a generic device check?",https://github.com/pytorch/pytorch/pull/185324,7ff942d3d3faba722f02c6bf9dc9b791de129f5571ba349bee501430a0dc0091 closes,pr,188965,issue,147170,high,pr.closingIssuesReferences,pr #188965 declares a closing reference to issue #147170.,https://github.com/pytorch/pytorch/pull/188965,b972d2a2c35d830150f6727b47fd60110546ae50de6b736a1722a0ed4677fe8f closes,pr,188965,issue,147170,high,pr.body,"Fixes #147170 Summary get_source_partitions() in torch/fx/passes/utils/source_matcher_utils.py could return input_nodes, output_nodes, and params in non-",https://github.com/pytorch/pytorch/pull/188965,f3ffdb7be3693b9cb98091ed3555c5f70d2f1e9a3b90caf7797a4d4c83aa6e0a review guidance,pr,188965,issue,147170,high,pr.reviews[0].body,"💡 Codex Review Here are some automated review suggestions for this pull request. Reviewed commit: 35dbf6f4e3 ℹ️ About Codex in GitHub Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you Open a pull request for review Mark a draft as ready Comment ""@code...",https://github.com/pytorch/pytorch/pull/188965,e439a7e1c96741931db1d9181c3304b267be0d9fa7a82d043473869a6700459f closes,pr,186476,issue,122512,high,pr.body,"cleanup cannot both delete the same resume state, and generated boxed-resume local names are collision-resistant against user locals. Fixes #122512 Generated by my agent Benchmark Results: Manual CUDA memory repro from the issue-derived graph-break/resume case, n=6000, backend...",https://github.com/pytorch/pytorch/pull/186476,a33740985858f77113b832ced95492740b604556ffb5e6c7d1fd7261c01ff88a references,pr,186476,issue,174929,medium,pr.comments[0].body,"possibly due to flakiness on trunk: inductor / inductor-cpu-test / test (cpu_inductor_torchbench, 1, 2, mt-l-x86iamx-8-64, unstable) (gh) (#174929) detectron2_maskrcnn_r_50_fpn inductor / inductor-cpu-test / test (cpu_inductor_torchbench, 2, 2, mt-l-x86iamx-8-64, unstable) (gh...",https://github.com/pytorch/pytorch/pull/186476,e8cf7875ced4b6bc26061c050ab4f916d401644cd17936b18d348c496e98912c closes,pr,185876,issue,185497,high,pr.body,s keeps the fix local to the activation-saving decision rather than changing mutation replay or disabling recomputation more broadly. Fixes #185497 Generated by my agent Test Plan: Original issue-style CPU/CUDA reproducer: eager and compiled pass after the fix python test/indu...,https://github.com/pytorch/pytorch/pull/185876,8efdd5d21da87642a4dd0b73dc7f9f7e78c7a2b2b01db58ca5907656c739fc51 closes,pr,186246,issue,185382,high,pr.body,g and FMA parity. Narrowing only the dynamic scalar path fixes the stale specialization root cause with the smallest behavior change. Fixes #185382 Generated by my agent Test Plan: python - <<'PY' ... issue adam_step/train_loop reproducer ... PY: Eager vs Inductor 0.000000e+00...,https://github.com/pytorch/pytorch/pull/186246,6537535e2cecd1ba33fce57588f8f68389ed6f4e458e11ea17782f82aa5425ac review guidance,pr,188653,pr,188653,high,pr.reviews[0].body,"Could you please check first if fix from intel/torch-xpu-ops#3990 is relevant to test cases that would be fixed by this PR? Or if they are totally unrelated and doesn't change the results? It seems related on the first sight, but of course doesn't have to be.",https://github.com/pytorch/pytorch/pull/188653,0ffa1c3469c4f23361a49047ddf1b8102dc0c37ad9bcb4a5a9be816c5cb017f6 closes,pr,184508,issue,184229,high,pr.body,"d buffer placeholders before export builds user input metadata, and fill missing nn_module_stack metadata for DTensor dispatch nodes. Fixes #184229 Generated by my agent",https://github.com/pytorch/pytorch/pull/184508,5955acdd4ff8f02628bb2ab97c2693c4a27831d2e0217abf8a9b052149552a33 closes,pr,185897,issue,183957,high,pr.body,Stack from ghstack (oldest at bottom): -> #185897 Fixes #183957 torch.library.define accepts string schemas and passes them to the dispatcher schema parser. Generic enum.Enum worked because it is registe,https://github.com/pytorch/pytorch/pull/185897,1c4dba6c3d69a4aaeee5616baad73bf3859224d96b479908e3df4901cb200009 closes,pr,186564,issue,93437,high,pr.body,behavior. The skipped targets cover ATen empty/new_empty variants and PrimTorch empty factories that can appear after decomposition. Fixes #93437 Generated by my agent Test Plan: python test/functorch/test_minifier.py -k test_skip_uninitialized_suffix_outputs python test/funct...,https://github.com/pytorch/pytorch/pull/186564,d446cbbc801afaca96b21a42e90ae395dbad2e79110c036af78a23359c4bf712 closes,pr,185939,issue,145687,high,pr.body,Stack from ghstack (oldest at bottom): -> #185939 Fixes #145687 Generated by my agent The LLaMA-style int8 weight-only quantized MLP tail has a down-projection GEMM followed by residual add and RMSNorm.,https://github.com/pytorch/pytorch/pull/185939,c54a29a46acc084690b8752fb770ccf16b57634c82c103541e591cd0417d07aa references,pr,185939,issue,145687,medium,pr.comments[2].body,the constraint. **Test (`test_cpu_select_algorithm.py:1703-1768`)** The test is comprehensive and directly models the problem scenario from #145687. Asserting both `cpp_templated_kernel_counter=3` and `cpp_epilogue_fusion_counter=3` validates that all three GEMM projections (g...,https://github.com/pytorch/pytorch/pull/185939,94099b4648d3721197ccccc6f99bf2f1da606c2b8735f982a4af6231043901f0 references,pr,185939,issue,145687,medium,pr.comments[5].body,ng constraint for those same templates (which already handle reindexing in codegen). The test directly models the LLaMA-style MLP tail from #145687. ### Detailed Feedback **`torch/_inductor/ir.py:10035-10050` — Force realization for int8 WoQ GEMM template consumers** The logic...,https://github.com/pytorch/pytorch/pull/185939,39765013f2663544efeb13794b9051b5a283cd4da9d664e49fccd3d94e5b7eac closes,pr,172074,issue,170648,high,pr.closingIssuesReferences,pr #172074 declares a closing reference to issue #170648.,https://github.com/pytorch/pytorch/pull/172074,f71010039dc834bb010f09dd9c6b8b76a36ae52eab113fb555ae343d60ca6f8b closes,pr,172074,issue,170648,high,pr.body,"Fixes #170648 Problem When fully_shard() is applied, it creates new nn.Parameter objects to hold the sharded data. The original parameter's grad_dtype pr",https://github.com/pytorch/pytorch/pull/172074,3f78f605d23c9cb0d324c5a61ad53d147365c2f4dac15779964a1e3a3370556d references,pr,189240,pr,186918,medium,pr.body,"d broadcast error logging, requires multi-device process group setup. Test Plan: python test/distributed/test_c10d_logger.py -v Depends on: #186918 cc: @mansiag05 @vishalgoyal316 @RiyaP2508",https://github.com/pytorch/pytorch/pull/189240,fa79a7482e4aa7cc23a76d0231ae09b69300571d68699aadb345570f8d9259b2 closes,pr,186380,issue,129418,high,pr.body,"mpositions must not lower to mutable ops. Running the decompositions under functionalization fixes the root ordering problem instead. Fixes #129418 Generated by my agent Benchmark Results Measured make_fx(dispatch_functionalize(...)) tracing with timeit.repeat(number=50, repea...",https://github.com/pytorch/pytorch/pull/186380,31aa5e0dcca62d43b5b78462d2cf74730046d46066fc7f66859493904d4dd93d references,pr,186380,issue,129418,medium,pr.comments[11].body,"t(func) is decomp_fn` (identity against the global table), so a user's `neg_decomp` is *not* default → runs under functionalization (fixing #129418's reproducer), while `core_aten`/prims default decomps keep their old proxy-path lowering → existing AOT/export graphs are preser...",https://github.com/pytorch/pytorch/pull/186380,ed45cf514e2da87a7707d6c2c2e2d31338a2d113cda367545d2d7cdafb52ddee closes,pr,186446,issue,126474,high,pr.body,ly. A broad generic check would incorrectly hard-error on ops like masked_fill_ where native only warns for internal overlap of self. Fixes #126474 Generated by my agent Benchmark Results: Measured compile time for a small torch.compile(backend='eager') function with 20 inplac...,https://github.com/pytorch/pytorch/pull/186446,d9b7921dd3fd0f31dbd7c8efeb12e35699041c9a1ae92b7714727a5f92ba3612 closes,pr,186505,issue,118332,high,pr.body,Stack from ghstack (oldest at bottom): -> #186505 The remaining actionable part of #118332 is that guards like 1 % divisor != 0 stay as modular guards even when range reasoning already knows the divisor is positive. For an integer,https://github.com/pytorch/pytorch/pull/186505,41fdba8be6e085da1b45a187e1a692b3a6458b491cac3a7bd639d30e039ef5d0 references,pr,187457,pr,189146,medium,pr.body,"leaf (cumulative-leaf fork on main), so the GitHub diff shows the softmax core+perf PRs + this one. To see only what this PR adds on top of #189146, use this fork compare (renders as a normal diff of just this delta): anagnorisis2peripeteia/pytorch@mps-softmax-perf...mps-logso...",https://github.com/pytorch/pytorch/pull/187457,8bb54e6d37643f565dfa7373228f42544242bcb370db049f8dc74edfc7b766c6 closes,pr,186462,issue,123972,high,pr.body,he actual condition we need. Checking export mode at the HOP support gate keeps the scope local to the behavior export can represent. Fixes #123972 Generated by my agent Test Plan: python test/export/test_export.py TestExport.test_cond_export_allows_branch_input_alias_with_gra...,https://github.com/pytorch/pytorch/pull/186462,31b9f1f0cbcabad23e3e46b63c780c117be0c4f5cada8356f9a57f1acb6d9e82 closes,pr,186923,issue,185141,high,pr.body,Ten OpOverload/OpOverloadPacket targets and the two missing fake/meta error patterns instead of adding torchvision-specific handling. Fixes #185141 Generated by my agent Test Plan: python test/dynamo/test_misc.py MiscTests.test_custom_op_missing_fake_impl_graph_breaks MiscTest...,https://github.com/pytorch/pytorch/pull/186923,ead5f3d13e6df15c23b99b2a192ebb54062f928464949e1d01dcbcad86710d57 closes,pr,186485,issue,122294,high,pr.body,namic_shapes contract users wrote. Marking only the declared Dim dimensions unbacked is narrower and keeps invalid examples rejected. Fixes #122294 Generated by my agent Test Plan: python test/export/test_export.py -k test_repeated_named_dim_invalid_zero_one_examples_fail pyth...,https://github.com/pytorch/pytorch/pull/186485,de09e2dd43bb0fc39da286d6e7316932abc2cedbbdb3e95a27ad70948d970e7a references,pr,186485,issue,122294,medium,pr.comments[2].body,apes.py` - [x] Review `torch/fx/passes/runtime_assert.py` - [x] Post review summary --- ### Summary This PR fixes a real user-facing issue (#122294) where declaring `dynamic_shapes` with `Dim(...)` on dimensions that happen to have size 0 or 1 in the example inputs causes expo...,https://github.com/pytorch/pytorch/pull/186485,ab590e33aa81805e21c38532d228d83559bb35c73e65d3c614288e1ac747bd9a closes,pr,186485,issue,122294,high,pr.comments[5].body,nsor.py` - [x] Review `torch/fx/passes/runtime_assert.py` - [x] Review test changes - [x] Post review summary --- ### Summary This PR fixes #122294 where `dynamic_shapes` with `Dim(...)` on dimensions that happen to have size 0 or 1 in the example inputs causes export to fail...,https://github.com/pytorch/pytorch/pull/186485,c763fd2f221554b4f8f18d470635a2d40df5caf6a8e34cade859d678bfdc3aab closes,pr,187493,issue,187340,high,pr.body,o match eager. This keeps the fix scoped to nextafter's dtype semantics rather than treating non-contiguous layout as the root cause. Fixes #187340 Generated by my agent Test Plan: Inline CUDA repro before fix: compiled fp16 output equaled input bits while eager moved by one f...,https://github.com/pytorch/pytorch/pull/187493,92e8c63eb4599f98c98d3e84c2bff39aecbd19e37b65a6cc031a5970a2aba20f closes,pr,184694,issue,178680,high,pr.body,"including static traces and AOTAutograd cache hits, so alias-sensitive synthetic-base graphs are not reused for incompatible inputs. Fixes #178680 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...",https://github.com/pytorch/pytorch/pull/184694,d99a0cfc0ff5917e7e9569f0a0f6b7e21ba54aebaa303203c4d60c79942207e8 references,pr,184694,issue,178680,medium,pr.comments[2].body,ges - [x] Review `guards.cpp` changes - [x] Review test changes - [x] Post review feedback --- ### Summary This PR fixes a correctness bug (#178680) where AOTAutograd's synthetic-base graphs could be incorrectly reused for inputs with incompatible storage-overlap relationships...,https://github.com/pytorch/pytorch/pull/184694,4d052aeec8e71bd3e0363e459770f092b28fffe61588a841ea92cf14b49eedb0 review guidance,pr,184694,issue,178680,high,pr.reviews[0].body,"(Reviewed by me, assisted by AI) [blocker] The five benchmark-baseline flips from pass to eager_two_runs_differ -- mobilenetv2_100 and tf_efficientnet_b0 in inductor_timm_training.csv, and mnasnet1_0, mobilenet_v2, shufflenet_v2_x1_0 in inductor_torchbench_training.csv -- diverge from current mai...",https://github.com/pytorch/pytorch/pull/184694,e1a6ce0ac8da148cea1fd813723af8a6b6ea236c2c9af2de6814dcafcd52c42c review guidance,pr,184694,pr,184694,high,pr.reviews[0].body,"(Reviewed by me, assisted by AI) [blocker] The five benchmark-baseline flips from pass to eager_two_runs_differ -- mobilenetv2_100 and tf_efficientnet_b0 in inductor_timm_training.csv, and mnasnet1_0, mobilenet_v2, shufflenet_v2_x1_0 in inductor_torchbench_training.csv -- diverge from current mai...",https://github.com/pytorch/pytorch/pull/184694,2dea2960c2622885f95e339db473e6646f86a6616044f4dba5e8abca35729f7a review guidance,pr,184694,issue,178680,high,pr.reviews[2].body,"(Reviewed by me, assisted by AI) [blocker] The five benchmark-baseline flips from `pass` to `eager_two_runs_differ` -- `mobilenetv2_100` and `tf_efficientnet_b0` in `inductor_timm_training.csv`, and `mnasnet1_0`, `mobilenet_v2`, `shufflenet_v2_x1_0` in `inductor_torchbench_training.csv` -- diverg...",https://github.com/pytorch/pytorch/pull/184694#pullrequestreview-4566363664,936735237d01ea51f99efca8d9f549476fa34039b9513b051a622f0fa7c22ddd review guidance,pr,184694,pr,184694,high,pr.reviews[2].body,"(Reviewed by me, assisted by AI) [blocker] The five benchmark-baseline flips from `pass` to `eager_two_runs_differ` -- `mobilenetv2_100` and `tf_efficientnet_b0` in `inductor_timm_training.csv`, and `mnasnet1_0`, `mobilenet_v2`, `shufflenet_v2_x1_0` in `inductor_torchbench_training.csv` -- diverg...",https://github.com/pytorch/pytorch/pull/184694#pullrequestreview-4566363664,d50c948f195016346c29dbc3564da78d0a3b4c214bba226b1c303ab3af09af1f closes,pr,187311,issue,187284,high,pr.body,"pported, and the issue's hessian repro also fails with cudagraphs. This change targets the Inductor-specific forward AD tangent loss. Fixes #187284 Generated by my agent Benchmark Results: Command: python /tmp/bench_187284.py Workload: 20 repeated FX graph cache hits for a sma...",https://github.com/pytorch/pytorch/pull/187311,0687c23a71fe5c5388c63496b174a13e0f4367e755c1cfcc9e2a0427a50cdf7e references,pr,184355,issue,177259,medium,pr.body,Stack from ghstack (oldest at bottom): #184678 #184356 -> #184355 Fixes parts of #177259 cc @mruberry,https://github.com/pytorch/pytorch/pull/184355,bf94f3c8ea86fd65b89017f34befbd39042aba7b30ebb96a3dacca2298a82620 references,pr,184355,pr,184356,medium,pr.body,Stack from ghstack (oldest at bottom): #184678 #184356 -> #184355 Fixes parts of #177259 cc @mruberry,https://github.com/pytorch/pytorch/pull/184355,12aba60a089e4c3471ec4c1d4c97c7746dea6eb8ae76637ba0ee4cfaa60c7bb9 references,pr,184355,pr,184678,medium,pr.body,Stack from ghstack (oldest at bottom): #184678 #184356 -> #184355 Fixes parts of #177259 cc @mruberry,https://github.com/pytorch/pytorch/pull/184355,79d3071be5ee2fec14b0ac1c8685feeadb2d54af23dcbbc7cd40a7ec004499ac review guidance,pr,184355,issue,177259,high,pr.reviews[0].body,"SGTM, thank you very much.",https://github.com/pytorch/pytorch/pull/184355,e4eda10a25963e8f41c4a359dc730276acbbf331780291ec6345b442a7ac72ba review guidance,pr,184355,pr,184355,high,pr.reviews[0].body,"SGTM, thank you very much.",https://github.com/pytorch/pytorch/pull/184355,eb5eb6be82941faeb4549893614f8f0d994593ece6a9b3329c5b6026e71babb1 review guidance,pr,184355,pr,184356,high,pr.reviews[0].body,"SGTM, thank you very much.",https://github.com/pytorch/pytorch/pull/184355,1f3bf20a960e9a1a1ecfeba186c5550d3c2eecf915768fe58588e6067fe581ff review guidance,pr,184355,pr,184678,high,pr.reviews[0].body,"SGTM, thank you very much.",https://github.com/pytorch/pytorch/pull/184355,1029453ea764daf6ccfdda03812c1ef4e1244604b3bf5543932c7c74d260ab09 closes,pr,184323,issue,127863,high,pr.body,"ch subkernel bodies and emit a single shared Triton body with per-branch setup, reducing repeated code that slows Triton compilation. Fixes #127863 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...",https://github.com/pytorch/pytorch/pull/184323,1d823ff8e1ecac8ddd8a72e58ec125a976fcbf9e3a6946e0e6d45a4ed20f07f5 closes,pr,184573,issue,183396,high,pr.body,"hile capturing the public tensor-length CTC op for Dynamo/Inductor fallback, so torch.compile preserves runtime backend selection.\n\nFixes #183396\nGenerated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nr...",https://github.com/pytorch/pytorch/pull/184573,3173dd60fb66e5171f6e747d72058a74526b3935c247b02476fc078c44c0fa68 review guidance,pr,184573,issue,183396,high,pr.reviews[0].body,It shouldn't be necessary to special case the Dynamo code,https://github.com/pytorch/pytorch/pull/184573,8d77033b7a784a7e288296591394737222348632b9d83986205771fc466e99a9 review guidance,pr,184573,pr,184573,high,pr.reviews[0].body,It shouldn't be necessary to special case the Dynamo code,https://github.com/pytorch/pytorch/pull/184573,29bb39ed975e404eff196d6bb407ec2e2abde510f93cafe4eb4665555336e154 review guidance,pr,184573,issue,183396,high,pr.reviews[2].body,It shouldn't be necessary to special case the Dynamo code,https://github.com/pytorch/pytorch/pull/184573#pullrequestreview-4340621506,056c36ee61146d567d4693a2cb0e31cb90d3ad0539cf39ef91d8296d596747b4 review guidance,pr,184573,pr,184573,high,pr.reviews[2].body,It shouldn't be necessary to special case the Dynamo code,https://github.com/pytorch/pytorch/pull/184573#pullrequestreview-4340621506,3d7f5f78aea3bba90d6a44593fa44df259e6cb695c676fd5b737dc497ae6c825 references,pr,184678,issue,177259,medium,pr.body,Stack from ghstack (oldest at bottom): -> #184678 #184356 #184355 Fixes parts of #177259 cc @mruberry,https://github.com/pytorch/pytorch/pull/184678,5523a6b17f1c2c0a885da78487b3b8a4bfa668d6b3340514be66137cdde6cc4a references,pr,184678,pr,184355,medium,pr.body,Stack from ghstack (oldest at bottom): -> #184678 #184356 #184355 Fixes parts of #177259 cc @mruberry,https://github.com/pytorch/pytorch/pull/184678,83bd827eae68f77033eda2c7b33cff85b69ebc2c05e1adcc98acf282578467f6 references,pr,184678,pr,184356,medium,pr.body,Stack from ghstack (oldest at bottom): -> #184678 #184356 #184355 Fixes parts of #177259 cc @mruberry,https://github.com/pytorch/pytorch/pull/184678,cacb5dd1388acf31fe1080518d3ad25625ae514947e516f348a2ea9682c7d404 closes,pr,183628,issue,183249,high,pr.body,ose samples appear in mixed-random Inductor graphs. Also fix fractional_max_pool3d sample dimension ordering to match native kernels. Fixes #183249 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/183628,a008da772e0cc4fe49bb9d20771a671a48c3d01d0c61a1c5cfb79bece5d1dc55 closes,pr,183654,issue,181696,high,pr.body,"the weight-only, bias-only, and weight+bias cases, in both the _refs decomposition and Inductor's CPU affine lowering. Fixes #183120 Fixes #181696 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...",https://github.com/pytorch/pytorch/pull/183654,ac357bce43218b9c9ac80538ed396d246e3830bede3acde97f0298f5c7984b7b closes,pr,183654,issue,183120,high,pr.body,"utation across the weight-only, bias-only, and weight+bias cases, in both the _refs decomposition and Inductor's CPU affine lowering. Fixes #183120 Fixes #181696 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzhe...",https://github.com/pytorch/pytorch/pull/183654,de1e3205357670158add7b7a80a9037354bd60c731e5d90b1f0da634cbfccd7b review guidance,pr,183654,issue,181696,high,pr.reviews[1].body,agree with your comment,https://github.com/pytorch/pytorch/pull/183654,c329ab3166736f6dbfad44398ab410f68b2e35faa2dee725b2094ef4642e3934 review guidance,pr,183654,issue,183120,high,pr.reviews[1].body,agree with your comment,https://github.com/pytorch/pytorch/pull/183654,90d1f85c0db4999e885aa13518fb130d61e33e89c4552dbee8bc3705c65e95b0 review guidance,pr,183654,pr,183654,high,pr.reviews[1].body,agree with your comment,https://github.com/pytorch/pytorch/pull/183654,e4c7b4094f7a0aa0a6a3e3ba09f562ffe17fb47a8ba04b598d35daf5661f398d closes,pr,184595,issue,183015,high,pr.body,y allow non-fake functorch wrapper tensors while fake-propagating _autograd_grad so compiled vmap(hessian) works across CPU and CUDA. Fixes #183015 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/184595,db1050adfb0590845aa60d6dc2b6afb1f7e3862a8dd916009792480530f3683a closes,pr,187486,issue,187415,high,pr.body,te. Targeted invalidation at old binding sites keeps the existing rebind invariant explicit and local to the code paths that need it. Fixes #187415 Generated by my agent Test Plan: Original issue repro before fix: failed during compilation with torch._inductor.exc.InductorErro...,https://github.com/pytorch/pytorch/pull/187486,6d038bf7fb8bac02c52deebfdb52943d3987b69140545b8cc4dd6e60976f04c0 references,pr,187456,issue,187455,medium,pr.body,[MPS] Migrate softmax and _softmax_backward to native Metal kernels (core) What / why First of a small series (Part of #187455) migrating MPS softmax off MPSGraph. This core PR lands the correctness foundation: native Metal kernels for _softmax / _softmax_backward_d,https://github.com/pytorch/pytorch/pull/187456,a22f9c18a179c3580935624be66ee786c9ddb3038e5bec77fd7017111b33be10 closes,pr,184602,issue,182940,high,pr.body,s frozen and keep fake tensor cache entries distinct across guard-mutation state so eager fake execution preserves symbolic metadata. Fixes #182940 Generated by my agent,https://github.com/pytorch/pytorch/pull/184602,240d0ebc760944d2dde2a311767b10c2b9bec74d2462316ca0266287a3a80625 review guidance,pr,184602,issue,182940,high,pr.reviews[0].body,"Reorganize the diff so that it is more obvious from diff review that the original logic hasn't changed, and we're only adding carveouts for frozen",https://github.com/pytorch/pytorch/pull/184602,610f44c8ace9614f9eb16adf5b3cc627016a355aafb531980af0442c06b2905a review guidance,pr,184602,pr,184602,high,pr.reviews[0].body,"Reorganize the diff so that it is more obvious from diff review that the original logic hasn't changed, and we're only adding carveouts for frozen",https://github.com/pytorch/pytorch/pull/184602,1aee03106ea57e030594a9c3b0b26a11dfd00619d039c6623e9afbb6917a99d1 closes,pr,184632,issue,182200,high,pr.body,"ch path and the fast binary path, and covers the real-valued promotion ops including copysign and transposed complex tensor variants. Fixes #182200 Fixes #184101 Generated by my agent Test Plan: lintrunner -a python test/test_meta.py -k test_cuda_mixed_dtype_binary_pointwise_f...",https://github.com/pytorch/pytorch/pull/184632,5e85a0e3178a8bca4f885f17140ad8d3a7431b3d3562a550cc1b557c356c8e4d closes,pr,184632,issue,184101,high,pr.body,"e fast binary path, and covers the real-valued promotion ops including copysign and transposed complex tensor variants. Fixes #182200 Fixes #184101 Generated by my agent Test Plan: lintrunner -a python test/test_meta.py -k test_cuda_mixed_dtype_binary_pointwise_fake_stride pyt...",https://github.com/pytorch/pytorch/pull/184632,d6eaac129cbba7d07df343a38814dc7bc4f8b83cd2af6e7f1802d509a6bfeb36 review guidance,pr,184632,issue,182200,high,pr.reviews[0].body,"Instead, have the C++ meta implementation sniff for a fake device, and then fix the striding correctly in that case.",https://github.com/pytorch/pytorch/pull/184632,3bd4746305c7843bcc5fb5ffa78f054525dad271b31f924f8fd0ee311f7cc7ec review guidance,pr,184632,issue,184101,high,pr.reviews[0].body,"Instead, have the C++ meta implementation sniff for a fake device, and then fix the striding correctly in that case.",https://github.com/pytorch/pytorch/pull/184632,db181397c05cc83bfcf1a6d38e97f877fa8eccaea023b318e7cf289269047438 review guidance,pr,184632,pr,184632,high,pr.reviews[0].body,"Instead, have the C++ meta implementation sniff for a fake device, and then fix the striding correctly in that case.",https://github.com/pytorch/pytorch/pull/184632,343ad7fe768e8db5eaea30eb2a1ac9607dd98b3e393dda497779ac58302f439e closes,pr,184639,issue,181891,high,pr.body,"accepting None constants through Dynamo, CUDA graph conditional capture, and Inductor while guarding unsupported differing constants. Fixes #181891 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...",https://github.com/pytorch/pytorch/pull/184639,967361283427c3494c450ea428c4a1758c870bde13176d85cab5728855517fbf review guidance,pr,184639,issue,181891,high,pr.reviews[7].body,"claude asked for the following: ``` Recommendation Request Changes — implementation looks correct (Python-wrapper path verified to already be output-type-agnostic; constant slots don't shift MultiOutput indices), but add (1) a positive Inductor test for an equal constant output including cpp_wrap...",https://github.com/pytorch/pytorch/pull/184639#pullrequestreview-4403652238,81fbf64f323deec42b2738f4366a61c3ff0d75ba7d2b3effd11f182c68e0d81a review guidance,pr,184639,pr,184639,high,pr.reviews[7].body,"claude asked for the following: ``` Recommendation Request Changes — implementation looks correct (Python-wrapper path verified to already be output-type-agnostic; constant slots don't shift MultiOutput indices), but add (1) a positive Inductor test for an equal constant output including cpp_wrap...",https://github.com/pytorch/pytorch/pull/184639#pullrequestreview-4403652238,f7e5456282d63341c1549e5c356ec76aabbc685fd7b865363973185fd999ea24 closes,pr,185469,issue,155856,high,pr.body,"ow use EagerAndRecordGraphs for both TMA descriptor variants and assert one graph is captured, so a future zero-graph fallback fails. Fixes #155856 Generated by my agent Test Plan: python test/dynamo/test_reconstruct.py -k tma (passed: OK, skipped=1) python test/dynamo/test_re...",https://github.com/pytorch/pytorch/pull/185469,a328d034242f8446fa4825a1cb309bff28fcad50e841e9ede0da61e8f0b8d890 closes,pr,183701,issue,181690,high,pr.body,duction suffix and emitted full-width non-contiguous stores. Reuse the tail vector kernel suffix so tail_size bounds the final store. Fixes #181690 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/183701,2df88be6f262a8b14cded8ff05e8e8097653dd85c596b8b4d6456a9cc7181894 closes,pr,185496,issue,146719,high,pr.body,"alone, handles dynamic-shape SymInt metadata, and falls back rather than guessing if the exact alias key is ambiguous. Fixes #155044 Fixes #146719 Generated by my agent Test Plan: python test/export/test_export.py TestExport.test_autograd_grad_reuses_saved_forward_tensor TestE...",https://github.com/pytorch/pytorch/pull/185496,e3ae2ccaffc6b8b0447cf251c1ab378a7f65c976c003d44a47b5356c5842a65c closes,pr,185496,issue,155044,high,pr.body,"ted by storage alone, handles dynamic-shape SymInt metadata, and falls back rather than guessing if the exact alias key is ambiguous. Fixes #155044 Fixes #146719 Generated by my agent Test Plan: python test/export/test_export.py TestExport.test_autograd_grad_reuses_saved_forwa...",https://github.com/pytorch/pytorch/pull/185496,1e58def40867a555dcdabba7b6c1f02511a98042e89b0a181a9bb76da2001538 review guidance,pr,185496,issue,146719,high,pr.reviews[0].body,make_fx should handle this case already (full forward+backward tracing). I would expect no code changes to proxy_tensor.py. Please do a more detailed root cause and explanation.,https://github.com/pytorch/pytorch/pull/185496,81f0d3a992487f660eed1bdc0103b657cc65aea86c06b8eef544d2d3f748aeb1 review guidance,pr,185496,issue,155044,high,pr.reviews[0].body,make_fx should handle this case already (full forward+backward tracing). I would expect no code changes to proxy_tensor.py. Please do a more detailed root cause and explanation.,https://github.com/pytorch/pytorch/pull/185496,c28c0ef60545ab657abd0c20ce3c493f246256b23da33d115e41578b01a041dd review guidance,pr,185496,pr,185496,high,pr.reviews[0].body,make_fx should handle this case already (full forward+backward tracing). I would expect no code changes to proxy_tensor.py. Please do a more detailed root cause and explanation.,https://github.com/pytorch/pytorch/pull/185496,7c17853dbc80f140765affca79c7fc8d373fbd42a10fed779c13959b957e75ed review guidance,pr,185496,issue,146719,high,pr.reviews[1].body,"My agent says I dug into the make_fx distinction you called out. Default post-dispatch make_fx does handle the repro, but non-strict export intentionally uses fake pre-dispatch make_fx; that path is where the saved sqrt result reappears as a fresh FakeTensor wrapper and misses proxy lookup. I upd...",https://github.com/pytorch/pytorch/pull/185496,54d498028a5b8bc3146fdeb89b2556a39e8aa1f79c2313611ccce524d7c120de review guidance,pr,185496,issue,155044,high,pr.reviews[1].body,"My agent says I dug into the make_fx distinction you called out. Default post-dispatch make_fx does handle the repro, but non-strict export intentionally uses fake pre-dispatch make_fx; that path is where the saved sqrt result reappears as a fresh FakeTensor wrapper and misses proxy lookup. I upd...",https://github.com/pytorch/pytorch/pull/185496,c74393507a3c80d973c6b462697515f7820e29ae191803c1c830255602248309 review guidance,pr,185496,pr,185496,high,pr.reviews[1].body,"My agent says I dug into the make_fx distinction you called out. Default post-dispatch make_fx does handle the repro, but non-strict export intentionally uses fake pre-dispatch make_fx; that path is where the saved sqrt result reappears as a fresh FakeTensor wrapper and misses proxy lookup. I upd...",https://github.com/pytorch/pytorch/pull/185496,9cabb553a6a90b2b8c68080d859cb1ef7282b3b3836fcfe10256ef2eb7a55036 review guidance,pr,185496,issue,146719,high,pr.reviews[2].body,make_fx should handle this case already (full forward+backward tracing). I would expect no code changes to proxy_tensor.py. Please do a more detailed root cause and explanation.,https://github.com/pytorch/pytorch/pull/185496#pullrequestreview-4500364747,5941781cf13fd87a268de8cdbff78c4ee29e6a46890212415edf602ca8e1439c review guidance,pr,185496,issue,155044,high,pr.reviews[2].body,make_fx should handle this case already (full forward+backward tracing). I would expect no code changes to proxy_tensor.py. Please do a more detailed root cause and explanation.,https://github.com/pytorch/pytorch/pull/185496#pullrequestreview-4500364747,604c65667f29faaf9b9ad64f95f2b414a89a76f26ea337090220255757912f5b review guidance,pr,185496,pr,185496,high,pr.reviews[2].body,make_fx should handle this case already (full forward+backward tracing). I would expect no code changes to proxy_tensor.py. Please do a more detailed root cause and explanation.,https://github.com/pytorch/pytorch/pull/185496#pullrequestreview-4500364747,3bb1771e6ab62bad2fce2dcec07d4622cbfc11d229346fccfe4cb2d54af58ec5 review guidance,pr,185496,issue,146719,high,pr.reviews[3].body,"My agent says I dug into the make_fx distinction you called out. Default post-dispatch make_fx does handle the repro, but non-strict export intentionally uses fake pre-dispatch make_fx; that path is where the saved sqrt result reappears as a fresh FakeTensor wrapper and misses proxy lookup. I upd...",https://github.com/pytorch/pytorch/pull/185496#pullrequestreview-4519936395,bbc8b38e7bd9be2b523b6cc64a400f61406c232e971ffa5197fe475fe77caafc review guidance,pr,185496,issue,155044,high,pr.reviews[3].body,"My agent says I dug into the make_fx distinction you called out. Default post-dispatch make_fx does handle the repro, but non-strict export intentionally uses fake pre-dispatch make_fx; that path is where the saved sqrt result reappears as a fresh FakeTensor wrapper and misses proxy lookup. I upd...",https://github.com/pytorch/pytorch/pull/185496#pullrequestreview-4519936395,40c871f72f214733eb7c27402342a5803fbd9ad20a404e41e4bf3b4cb9850d07 review guidance,pr,185496,pr,185496,high,pr.reviews[3].body,"My agent says I dug into the make_fx distinction you called out. Default post-dispatch make_fx does handle the repro, but non-strict export intentionally uses fake pre-dispatch make_fx; that path is where the saved sqrt result reappears as a fresh FakeTensor wrapper and misses proxy lookup. I upd...",https://github.com/pytorch/pytorch/pull/185496#pullrequestreview-4519936395,28399588766c3db7134df639b2bc836652d8a118c162ccde1f022fc476b486b5 references,pr,184356,issue,177259,medium,pr.body,Stack from ghstack (oldest at bottom): #184678 -> #184356 #184355 Fixes parts of #177259 cc @mruberry,https://github.com/pytorch/pytorch/pull/184356,dd57bb06b2588f3922dd1c27a69540f3d6caa9b3b72193ccb6eea268dc71ce1d references,pr,184356,pr,184355,medium,pr.body,Stack from ghstack (oldest at bottom): #184678 -> #184356 #184355 Fixes parts of #177259 cc @mruberry,https://github.com/pytorch/pytorch/pull/184356,b8beffec0e0cbfd1ea21de888482eb8f2f064271dad9448bb3e8d09c5c1052b0 references,pr,184356,pr,184678,medium,pr.body,Stack from ghstack (oldest at bottom): #184678 -> #184356 #184355 Fixes parts of #177259 cc @mruberry,https://github.com/pytorch/pytorch/pull/184356,2b42b57705bd105dafb25baceb0cd011894c9b304befeb033b845c20553b9c2c review guidance,pr,184356,issue,177259,high,pr.reviews[0].body,SGTM.,https://github.com/pytorch/pytorch/pull/184356,827a274ff80bdf0633241a776d85e0be09ecdef948ac26c0bc19e3258ab4cbea review guidance,pr,184356,pr,184355,high,pr.reviews[0].body,SGTM.,https://github.com/pytorch/pytorch/pull/184356,bb6e24e32f0d6f2ae54c978ac7ab8310458a5c5a1c289103100a944c0c59b289 review guidance,pr,184356,pr,184356,high,pr.reviews[0].body,SGTM.,https://github.com/pytorch/pytorch/pull/184356,18dad65b4fc7e29004dcf221aa8eb4e9829eaf21292084701795fc0039400408 review guidance,pr,184356,pr,184678,high,pr.reviews[0].body,SGTM.,https://github.com/pytorch/pytorch/pull/184356,bec0efb024a31246c4250b5563b247e51c27b4e9191feb55e95d057d801b8f52 review guidance,pr,184356,issue,177259,high,pr.reviews[1].body,Thank you!,https://github.com/pytorch/pytorch/pull/184356,501b87b78cdc5d84cb0d061a5faa90f94c12e01c5c19650ed1106753ccf28d5a review guidance,pr,184356,pr,184355,high,pr.reviews[1].body,Thank you!,https://github.com/pytorch/pytorch/pull/184356,2ec129b64c4aad886b96471e4b390480c0b55ad33ca816e0b7e48ada3c5d0cac review guidance,pr,184356,pr,184356,high,pr.reviews[1].body,Thank you!,https://github.com/pytorch/pytorch/pull/184356,62d05a3d16c1ec76d35b2549e8d57a7702ba110d0cdf63f7399eb901a3223abd review guidance,pr,184356,pr,184678,high,pr.reviews[1].body,Thank you!,https://github.com/pytorch/pytorch/pull/184356,9877410ac6eac0bd109785e3da0752ef9eb367badccd70066d3c26df397c722c closes,pr,185804,issue,147326,high,pr.body,ct_export_len_tensor_pytree_nn_module_input python - <<'PY' ... PY # adapted issue reproducer lintrunner -a git diff --cached --check Fixes #147326 Generated by my agent,https://github.com/pytorch/pytorch/pull/185804,8dc19a8bf66882b8b26dca5d34e368b74935e51098dd95111c360904f49397da closes,pr,183511,issue,178040,high,pr.body,"owering that validates dtype/device and matmul metadata before dropping the ignored input, including zero-output and out_dtype cases. Fixes #178040 Generated with AI cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv @ji...",https://github.com/pytorch/pytorch/pull/183511,65a34e4e32e841c0792a9aa489d427960c05d80b5e0038b8c1b5fe50c3a05b72 references,pr,183511,issue,188721,medium,pr.comments[0].body,"akiness on trunk: B200 Smoke Tests / linux-jammy-cuda13.0-py3.12-gcc11-sm100 / test-osdc (smoke_b200, 1, 1, mt-l-x86iamx-22-225-b200) (gh) (#188721) test/inductor/test_flex_flash.py::TestFlexFlashCUDA::test_cutedsl_captured_alias_views_keep_distinct_layouts_case_offset_cuda Th...",https://github.com/pytorch/pytorch/pull/183511,92dac3b3ca7ecabec08782a9ba78ad0d7199745f950ae72efd7873be306ede0f closes,pr,186388,issue,186354,high,pr.body,Stack from ghstack (oldest at bottom): -> #186388 Fixes #186354 The native override eager router is installed as the real backend kernel for aten overrides. Normally that path should evaluate eager-only,https://github.com/pytorch/pytorch/pull/186388,60f447b6c09f9733db9af0759b5ab05c16204291b3ec569ae62b1f7726e4a4dd closes,pr,188990,issue,188989,high,pr.closingIssuesReferences,pr #188990 declares a closing reference to issue #188989.,https://github.com/pytorch/pytorch/pull/188990,6505409ce81aa109f2bad640eda47e20c7956616bfda97f60e379ebc859873db closes,pr,188990,issue,188989,high,pr.body,"Fix #188989 If flat_param is already on CPU, summon_full_params(offload_to_cpu=True) should not attempt to move it to CPU or free the unsharded paramet",https://github.com/pytorch/pytorch/pull/188990,e680d728a17e0cfeff1351ffcc67433cf43588d3ee69e859e509d48bb1103cf9 closes,pr,186656,issue,185814,high,pr.body,rdinate_descent_tuner.py TestCoordinateDescentTuner.test_mix_order_reduction_xblock_must_divide_rsplit git diff --check lintrunner -a Fixes #185814 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/186656,bbec1d3da6b9f23932ea619b53e3f0025c8bab6dff1dcf943d46004653e475bc review guidance,pr,186656,issue,185814,high,pr.reviews[0].body,"In max-autotune mode, we already autotune XBLOCK and RSPLIT_SIZE for mix order reduction. This PR seems mainly add more choices when max-autotune is not on? I'm curious about the impact on compilation time.",https://github.com/pytorch/pytorch/pull/186656,05c648f62192eec86fdc02c55cbe4bbf4ec19b2a8373b4bd7ab5641d6fd79ccf review guidance,pr,186656,pr,186656,high,pr.reviews[0].body,"In max-autotune mode, we already autotune XBLOCK and RSPLIT_SIZE for mix order reduction. This PR seems mainly add more choices when max-autotune is not on? I'm curious about the impact on compilation time.",https://github.com/pytorch/pytorch/pull/186656,85b674eb8035bd348e97614116b520c339d2f6a196e8ace799e152772889613d review guidance,pr,186656,issue,185814,high,pr.reviews[1].body,"We should not autotune XBLOCK by default, which will blow up compile time significantly",https://github.com/pytorch/pytorch/pull/186656,acf4fa89212853a7cd33ec455e4cc05fdf3313d36d38f0afe2ec75a96f2d31bb review guidance,pr,186656,pr,186656,high,pr.reviews[1].body,"We should not autotune XBLOCK by default, which will blow up compile time significantly",https://github.com/pytorch/pytorch/pull/186656,13529146d098ef6082ad5193b4053bb8e81fb36ba14f5e0ce061f78fd25a4530 closes,pr,184656,issue,181177,high,pr.body,backward derivative in compiled torch.func.grad. Add regression coverage for the compile path and direct op mask/logsumexp behavior. Fixes #181177 Generated by my agent,https://github.com/pytorch/pytorch/pull/184656,9ff834a86037c2640442b4a07aef548affa26644ac4d62133e845f628552c4b1 references,pr,184656,issue,181177,medium,pr.comments[5].body,"rrectness. **6. Test coverage is thorough** - `test_compile_grad_cpu_sdpa_higher_order`: End-to-end regression test for the original issue (#181177) - `test_cpu_flash_sdpa_autograd_aux_output_contract`: Validates shape, dtype, stride, and value correctness of logsumexp includi...",https://github.com/pytorch/pytorch/pull/184656,e1e61c0a1c26cd98c94aed78ae06be2fc4b8f3cddcb7dbe7b29277efedbd03b9 closes,pr,186857,issue,180651,high,pr.body,"r mask shapes. Using the actual num_blocks tensors preserves the data-dependent distribution without choosing a universal fake split. Fixes #180651 Generated by my agent Benchmark Results CUDA causal flex_attention workload: Z=1, H_Q=32, H_KV=8, Q=512, KV=1664, D=128, fp16, ma...",https://github.com/pytorch/pytorch/pull/186857,1362c189d59aab822135c0576bdcba98324ac0d6abcbc73254676241191173d3 references,pr,186857,pr,181096,medium,pr.body,"akeTensor inputs are skipped explicitly to keep autotune benchmark inputs concrete. I considered a partial/full heuristic like the stale PR #181096, but that bakes in causal-mask assumptions and was already shown to regress other mask shapes. Using the actual num_blocks tensor...",https://github.com/pytorch/pytorch/pull/186857,9be566e83ac875e9a9fa6425106873375925cc52aaf2df2c2ab235f12b8326ac closes,pr,185878,issue,185480,high,pr.body,"rics, GELU backward opmath, and frac signed-zero behavior respectively; none changes the expm1/ELU/CELU path or would fix this issue. Fixes #185480 Generated by my agent Test Plan: Reproduced pre-fix CELU and ELU mismatch on CUDA bf16 subnormal input: eager returned bf16 bits...",https://github.com/pytorch/pytorch/pull/185878,c68907a0ae9e56951a12dd64ac859d34202c767e24195fb2a799e583b067cd01 closes,pr,183744,issue,180447,high,pr.body,"epts BF16 CpuFullyConnected validation, so AArch64 BF16 linear can use the optimized ACL inner-product path under freezing. Partially fixes #180447 Generated by my agent cc @gujinghui @PenghuiCheng @XiaobingSuper @jianyuh @jgong5 @mingfeima @sanchitintel @ashokei @jingxu10 @mi...",https://github.com/pytorch/pytorch/pull/183744,d4b62e5100425b1fec268e77d885b5682ba9bdb9489fe5a76eaa606932028a00 review guidance,pr,183744,issue,180447,high,pr.reviews[0].body,"Thanks for taking the initiative to do this. We (Arm and Intel) usually run a comprehensive sets of acceptance tests before bumping oneDNN version, so we shouldn't be updating oneDNN version here. cc: @aditew01 @leslie-fang-intel",https://github.com/pytorch/pytorch/pull/183744,9ffce1101e0629c9dd837298945955176ef61143740b7c5651bfdb28f6e59864 review guidance,pr,183744,pr,183744,high,pr.reviews[0].body,"Thanks for taking the initiative to do this. We (Arm and Intel) usually run a comprehensive sets of acceptance tests before bumping oneDNN version, so we shouldn't be updating oneDNN version here. cc: @aditew01 @leslie-fang-intel",https://github.com/pytorch/pytorch/pull/183744,9ce4002f0c74837d923793c7be2e7558384170b5fa5a61e0724aa070941dd492 review guidance,pr,183744,issue,180447,high,pr.reviews[2].body,"Thanks for taking the initiative to do this. We (Arm and Intel) usually run a comprehensive sets of acceptance tests before bumping oneDNN version, so we shouldn't be updating oneDNN version here. cc: @aditew01 @leslie-fang-intel",https://github.com/pytorch/pytorch/pull/183744#pullrequestreview-4318339371,504205895a18a9df1ebb6df57a4c9ad7e67c8f9a69501fd7d0831b2b0dd01421 review guidance,pr,183744,pr,183744,high,pr.reviews[2].body,"Thanks for taking the initiative to do this. We (Arm and Intel) usually run a comprehensive sets of acceptance tests before bumping oneDNN version, so we shouldn't be updating oneDNN version here. cc: @aditew01 @leslie-fang-intel",https://github.com/pytorch/pytorch/pull/183744#pullrequestreview-4318339371,00cfd3e23d2c0c0c8fa4ab872a6aeeb5c0bf93566b957138177b71da06016bbb review guidance,pr,183744,issue,180447,high,pr.reviews[6].body,"LGTM, let's just wait for CI to pass. Once https://github.com/pytorch/pytorch/pull/181222 which updates oneDNN version to v3.12 is merged, https://github.com/pytorch/pytorch/issues/180447 will be fully fixed.",https://github.com/pytorch/pytorch/pull/183744#pullrequestreview-4417177872,4bc603afff1b4c06cea7109ef93e27b5753c0514ef9bd8bb07a904859366a54a review guidance,pr,183744,pr,183744,high,pr.reviews[6].body,"LGTM, let's just wait for CI to pass. Once https://github.com/pytorch/pytorch/pull/181222 which updates oneDNN version to v3.12 is merged, https://github.com/pytorch/pytorch/issues/180447 will be fully fixed.",https://github.com/pytorch/pytorch/pull/183744#pullrequestreview-4417177872,4b3e07ea6a7ce19f554fa82e70df570633bb89e98697e93c3c870ae571c05bce closes,pr,184667,issue,180354,high,pr.body,on HOP subgraphs while retracing and keep branch-local tensor constants out of the top-level decomposed export signature/state_dict. Fixes #180354 Generated by my agent,https://github.com/pytorch/pytorch/pull/184667,8dcaf2636b399607e0500cfb4ccadbf794a8962dacc68748dfa3033bc729b123 review guidance,pr,184667,issue,180354,high,pr.reviews[0].body,"I don't think this is the right fix, but I don't care enough about this problem to think about the right fix right now. For example, I don't care about torch.export very much at the moment, if you can come up with a minimal repro that does not involve export I may care about this more",https://github.com/pytorch/pytorch/pull/184667,56666e48ba50f1ef39962cc7ad0dfb71ebf9c9a39ce4cfa3200af62a522e3ae8 review guidance,pr,184667,pr,184667,high,pr.reviews[0].body,"I don't think this is the right fix, but I don't care enough about this problem to think about the right fix right now. For example, I don't care about torch.export very much at the moment, if you can come up with a minimal repro that does not involve export I may care about this more",https://github.com/pytorch/pytorch/pull/184667,7404603f9ce2b9778645600c106d3c796510c8929e529fc78c21eeaad6fffe9c review guidance,pr,184667,issue,180354,high,pr.reviews[1].body,"I'm also not too confident in the changes in this PR. What do you think about fixing it in this direction -- after the first export, can we unlift the tensor constants from the cond branch subgraphs and lift them as top-level tensor constants in the ExportedProgram. The subgraphs would then recei...",https://github.com/pytorch/pytorch/pull/184667,424f393b7c3c3a38175b0dd042a5c7aea5024c00661bc769e477f189f077f06a review guidance,pr,184667,pr,184667,high,pr.reviews[1].body,"I'm also not too confident in the changes in this PR. What do you think about fixing it in this direction -- after the first export, can we unlift the tensor constants from the cond branch subgraphs and lift them as top-level tensor constants in the ExportedProgram. The subgraphs would then recei...",https://github.com/pytorch/pytorch/pull/184667,6fb3cd71ab2c81de8aec63bb6dd7be37103bac8b4cfc5ecf6b99041b4ab489c4 references,pr,182987,pr,183028,medium,pr.body,Stack from ghstack (oldest at bottom): #183028 -> #182987 CachingAutotuner._combo_sequential_autotune runs a search over Phase 1 (block sizes per sub-kernel group) and Phase 2 (warps/sta,https://github.com/pytorch/pytorch/pull/182987,2d25fe3999e062eda81b433e791712d60e3d98a0fb209b1595078059991a077d review guidance,pr,182987,pr,182987,high,pr.reviews[0].body,"Hmm, we can consider accepting this, but, I'm surprised that this is only handling combo kernels. Is there any reason not to do a more general parallelization of configs to processes ? The same problem applies if we have a Reduction with hint Default. We compile 6 separate configs, all in one sub...",https://github.com/pytorch/pytorch/pull/182987,47cb58cba83d5f9e1d044b029af1b32b886381d145fa4b543c3bb0ab4da43752 review guidance,pr,182987,pr,183028,high,pr.reviews[0].body,"Hmm, we can consider accepting this, but, I'm surprised that this is only handling combo kernels. Is there any reason not to do a more general parallelization of configs to processes ? The same problem applies if we have a Reduction with hint Default. We compile 6 separate configs, all in one sub...",https://github.com/pytorch/pytorch/pull/182987,b5e69cf036308a304b7e2568ba717e2cb7e7a9418bc8887f1422a415d1b83789 review guidance,pr,183854,pr,179777,high,pr.reviews[0].body,Can we fix the underlying issue instead ?,https://github.com/pytorch/pytorch/pull/183854,6423094aef0c0b82ed2de999fc249a8e485efc43ac5f59d5dc218412c673ca43 review guidance,pr,183854,pr,183854,high,pr.reviews[0].body,Can we fix the underlying issue instead ?,https://github.com/pytorch/pytorch/pull/183854,f094b74ae353e15fc6d65a90407257d1acf21772d9dc28b43cc8b467645be03d review guidance,pr,183854,pr,179777,high,pr.reviews[1].body,issue closed?,https://github.com/pytorch/pytorch/pull/183854,f5cfebbb27ba9a491d9203cbe88d4b5a0bf5ac2a399e5b349fb51da8f0568c3b review guidance,pr,183854,pr,183854,high,pr.reviews[1].body,issue closed?,https://github.com/pytorch/pytorch/pull/183854,ccb060ec7ed3fafc1485219d78633583666fd2855f0d86e4fe12a0c3138a9576 closes,pr,187609,issue,185470,high,pr.body,"so output dtype semantics are unchanged. A prior stale PR, #185846, made the same root-cause fix but has not landed on current main. Fixes #185470 Generated by my agent Test Plan: python test/inductor/test_torchinductor.py GPUTests.test_threshold_low_precision_boundary_cuda (P...",https://github.com/pytorch/pytorch/pull/187609,5dc8b2530883ddd295a344e4f6b71b54847125f4c33cc6750d04a6ce819ca287 references,pr,187609,pr,185846,medium,pr.body,"2 comparison behavior. The selected result still uses the original input tensor, so output dtype semantics are unchanged. A prior stale PR, #185846, made the same root-cause fix but has not landed on current main. Fixes #185470 Generated by my agent Test Plan: python test/indu...",https://github.com/pytorch/pytorch/pull/187609,6c2f19bc01c2718a9c004bac5c1aecf19a58180f179015fa1e58a511864da513 closes,pr,185501,issue,185396,high,pr.closingIssuesReferences,pr #185501 declares a closing reference to issue #185396.,https://github.com/pytorch/pytorch/pull/185501,34b973865e5a135fac4d2cb2f26e90b644917cd3fea94f8d983170c58746cef7 closes,pr,185501,issue,185396,high,pr.body,Fixes #185396 This updates the PR to fix the underlying issue rather than only improve the diagnostic. Root cause: Inductor was generating a Triton point,https://github.com/pytorch/pytorch/pull/185501,46a0abf51e1f19c4205d05e565b1fc419b9687c7ab1236619f86446f3270074e closes,pr,186444,issue,126514,high,pr.body,differentiable paths use the numerator form because correctness there requires avoiding underflow or a differentiable divide by zero. Fixes #126514 Generated by my agent Benchmark Results: Command: python /data/users/jansel/pytorch-issue-fixer/state/126514/bench_capturable_ada...,https://github.com/pytorch/pytorch/pull/186444,a9f8bc9bfd2415f3fdbc8acd7c82f9fd445b2f81029eab8fe13097095bdad5e7 closes,pr,185082,issue,165073,high,pr.body,ivergence. The source-based unlift keeps the Dynamo graph lifted internally and makes the public ExportedProgram conversion explicit. Fixes #165073 Generated by my agent Test Plan: python test/export/test_export.py TestDynamismExpression.test_strict_export_lifts_symint_in_torc...,https://github.com/pytorch/pytorch/pull/185082,25dae014785fa2cd56582827300a03a94c1007110eafa5912cfa7a3213e755e2 closes,pr,184695,issue,178677,high,pr.body,sor sparsity conversion for the all-masked src_key_padding_mask case so eager follows the dense path that torch.compile already uses. Fixes #178677 Generated by my agent,https://github.com/pytorch/pytorch/pull/184695,b2033edb98f37d775d854cbb17992442951142ea23b0ba87466f48b4b30be7fd closes,pr,186150,issue,140171,high,pr.body,cking storage sizes or offsets; the tradeoff is extra recompilation for these rare storage-observable as_strided/view-scatter graphs. Fixes #140171 Generated by my agent Test Plan: python test/dynamo/test_repros.py -k as_strided_storage_metadata_recompile_inductor python test/...,https://github.com/pytorch/pytorch/pull/186150,5b81eeecda3174a4b4acdff7b10845ebda1b33ad092d6ebf0d4803eba38d5745 competes with,pr,186150,issue,140171,medium,pr.comments[5].body,"compile_fx, reinplace) - [x] Review test changes - [x] Post comprehensive review feedback --- ### Summary This PR fixes a correctness bug (#140171) where compiled `as_strided_scatter` on view inputs rejected valid storage offsets because the meta check compared against `input....",https://github.com/pytorch/pytorch/pull/186150,e1dce18a73930f28dec1a344b43c858bfb43d1b8870cacb677912165524177dc references,pr,186150,issue,140171,medium,pr.comments[8].body,"compile_fx, reinplace) - [x] Review test changes - [x] Post comprehensive review feedback --- ### Summary This PR fixes a correctness bug (#140171) where compiled `as_strided_scatter` on view inputs rejected valid storage offsets because the prim meta check compared against `i...",https://github.com/pytorch/pytorch/pull/186150,065732f34015e85dd228848bd6e722ee969343ffe30b6320bb36a5832ea9d8b8 references,pr,186150,issue,140171,medium,pr.comments[11].body,"compile_fx, reinplace) - [x] Review test changes - [x] Post comprehensive review feedback --- ### Summary This PR fixes a correctness bug (#140171) where compiled `as_strided_scatter` on view inputs rejected valid storage offsets because the prim meta check compared against `i...",https://github.com/pytorch/pytorch/pull/186150,952e2319f7f59b52bf7aea0a1294b0716927fcbd93315aa408cd048f1cea60df closes,pr,184709,issue,178388,high,pr.body,nit side effects are replayed correctly. Add coverage for module attribute assignment and preserved non-constructor wrapper mutation. Fixes #178388 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/184709,38b3d828f7f1965c54378486f2cd3ea4f76ead901d15543f7342e5d53070d700 closes,pr,183866,issue,178317,high,pr.body,ed consumers. Add an explicit low-precision dtype boundary for native matmul reductions and cover remove_no_ops/native bmm precision. Fixes #178317 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/183866,751822cc42c5f92797cc9f018a54b0fa845d2b2105c01dd6816eddfc98cdea43 references,pr,183866,issue,178317,medium,pr.comments[2].body,ontext and read changed files - [x] Analyze the implementation - [x] Post review feedback --- **Summary:** This PR fixes a correctness bug (#178317) where native `tl.dot`-based matmul lowerings accumulate in fp32 but `aten.mm`/`aten.bmm` should expose fp16/bf16 outputs to down...,https://github.com/pytorch/pytorch/pull/183866,4b8d21dc64818976d6f4f6782abfe8405500ff40d7e32fdcb43aa7c3e475ae76 closes,pr,184713,issue,178248,high,pr.body,mutable nn.Module hook dictionaries so forward-hook registration inside compiled forward no longer fails on RemovableHandle.next_id. Fixes #178248 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/184713,bf12cb913e536dba5aa0a284626abd8acbfb20f07f8d3d69fb001370b0cbfbed closes,pr,184716,issue,178127,high,pr.body,e transposed convolution validation in the meta shape path so invalid output_padding combinations raise instead of producing a shape. Fixes #178127 Generated by my agent,https://github.com/pytorch/pytorch/pull/184716,b80717baa531eef6814f6ee8767c85a2fdb1460ae6292110e2c023376b87bba9 closes,pr,185585,issue,153204,high,pr.body,ace alias-observable graphs soundly. Avoiding that would need a layout-parametric contiguous lowering rather than relaxing the guard. Fixes #153204 Generated by my agent Test Plan: python test/test_dynamic_shapes.py TestUbackedOps.test_unbacked_reshape1 TestUbackedOps.test_unb...,https://github.com/pytorch/pytorch/pull/185585,3a4afdbd175b8fb852bc224a0a13fd1164d997fc807efc5042557cf194474697 review guidance,pr,185585,issue,153204,high,pr.reviews[0].body,.,https://github.com/pytorch/pytorch/pull/185585,eb0130bcdfa3cabcc81ab74b7426aedc54a5bdb1bdd9410effc28fccbaf530de review guidance,pr,185585,pr,185585,high,pr.reviews[0].body,.,https://github.com/pytorch/pytorch/pull/185585,fda6d706bfafc349c18468978b8c33f60c6655e8b35ba2157bb97ce3540ffd26 review guidance,pr,185585,issue,153204,high,pr.reviews[1].body,.,https://github.com/pytorch/pytorch/pull/185585#pullrequestreview-4467896171,279e3b957cf1cbc74c1dd89c26e06eb70a2ad7430f33ade8c4d74327f8c0bb75 review guidance,pr,185585,pr,185585,high,pr.reviews[1].body,.,https://github.com/pytorch/pytorch/pull/185585#pullrequestreview-4467896171,d15c7998233ac28b9c59c1f04704515e414611039000064d99088f025106f5a8 closes,pr,183871,issue,178096,high,pr.body,le addcmul chains to binary folding. Adds CUDA regressions for the dynamic BatchNorm+Conv accuracy case and shared-conv no-fold path. Fixes #178096 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/183871,e6a6d4fe62e900072a6887d7ea1880e6eba1ad1868a461f5611517e9fb071e9f references,pr,183871,issue,178096,medium,pr.comments[2].body,"dynamic_conv_accuracy`: Uses `atol=0, rtol=0` — a strong assertion confirming exact numeric match with CUDA eager. Good regression test for #178096. - `test_shared_conv_bn_keeps_addcmul`: Validates that when the conv output is shared, `addcmul` is preserved (no binary folding...",https://github.com/pytorch/pytorch/pull/183871,eaf52dfdbdffce17bd82d5f2eff49b6f5c74665d5ae1140c02edd043daed1476 references,pr,183871,issue,178096,medium,pr.comments[3].body,"l freezing work, restoring non-CUDA eval BN ordering, and preserving conv-first ordering for the addcmul decomposition. I also clarified on #178096 that the remaining TF32-on repro gap is expected TF32 conv behavior, while this PR fixes the CUDA BN ordering mismatch found duri...",https://github.com/pytorch/pytorch/pull/183871,2eefe92336afe703583b6a12612b99ca5f68eacbdf2f804c4d62900588285f5e closes,pr,183871,issue,178096,high,pr.reviews[0].body,(Claude Review) [cleanup] `Fixes #178096` will auto-close an issue whose reported number won't change. The issue's ~9.5e-2 gap comes from the `conv2` path (eager cuDNN with TF32 vs,https://github.com/pytorch/pytorch/pull/183871,eecab4b9764dedf3dfa2e2dfd559f4c3de59580409e2e6934e64d6a275a7af0b review guidance,pr,183871,issue,178096,high,pr.reviews[0].body,"(Claude Review) [cleanup] Fixes #178096 will auto-close an issue whose reported number won't change. The issue's ~9.5e-2 gap comes from the conv2 path (eager cuDNN with TF32 vs the compiled conv), not from BatchNorm: the repro leaves TF32 on, and this PR's own regression test disables it (@with_t...",https://github.com/pytorch/pytorch/pull/183871,ec5a237b6be69b5d54c86971333c281a105301bfb5372346f19d097dd77e2958 review guidance,pr,183871,pr,183871,high,pr.reviews[0].body,"(Claude Review) [cleanup] Fixes #178096 will auto-close an issue whose reported number won't change. The issue's ~9.5e-2 gap comes from the conv2 path (eager cuDNN with TF32 vs the compiled conv), not from BatchNorm: the repro leaves TF32 on, and this PR's own regression test disables it (@with_t...",https://github.com/pytorch/pytorch/pull/183871,925efa8458cb80b696b9491e0d93365f0019d97bba5d5b122e73327a88f73fad review guidance,pr,183871,issue,178096,high,pr.reviews[2].body,"(Claude Review) [cleanup] `Fixes #178096` will auto-close an issue whose reported number won't change. The issue's ~9.5e-2 gap comes from the `conv2` path (eager cuDNN with TF32 vs the compiled conv), not from BatchNorm: the repro leaves TF32 on, and this PR's own regression test disables it (`@w...",https://github.com/pytorch/pytorch/pull/183871#pullrequestreview-4432551775,565b938dd18a671da8ce4564f1e3188c28f8514a74fb4ac91de067489fd15ef1 review guidance,pr,183871,pr,183871,high,pr.reviews[2].body,"(Claude Review) [cleanup] `Fixes #178096` will auto-close an issue whose reported number won't change. The issue's ~9.5e-2 gap comes from the `conv2` path (eager cuDNN with TF32 vs the compiled conv), not from BatchNorm: the repro leaves TF32 on, and this PR's own regression test disables it (`@w...",https://github.com/pytorch/pytorch/pull/183871#pullrequestreview-4432551775,eafc5e596baa48ee8e82ae208648eeabd80caef8594b439dbd9e26c2cf4d7ce7 closes,pr,188967,issue,137285,high,pr.closingIssuesReferences,pr #188967 declares a closing reference to issue #137285.,https://github.com/pytorch/pytorch/pull/188967,06b785e276949d60fe58b3afdb4d98c1970f57a4bd86d2592409dfc199f3b8eb closes,pr,188967,issue,137285,high,pr.body,"Fixes #137285 Summary Updates docs/source/logging.md to match the current PyTorch logging API: Added native_dsl component Added autotuning_inputs, cachin",https://github.com/pytorch/pytorch/pull/188967,39895f268bf245dfb85c53fd5bcf04fab49b7fc989a8847026fc9d67f66f26a4 review guidance,pr,188967,issue,137285,high,pr.reviews[0].body,"💡 Codex Review Here are some automated review suggestions for this pull request. Reviewed commit: f0ff13c74d ℹ️ About Codex in GitHub Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you Open a pull request for review Mark a draft as ready Comment ""@code...",https://github.com/pytorch/pytorch/pull/188967,9a31d6932d2a1e433a6dc5ae774876764bdc6986a0ba890677438202c11476b6 review guidance,pr,188967,issue,137285,high,pr.reviews[1].body,Looks good,https://github.com/pytorch/pytorch/pull/188967,fff0fb1c99b3035e887732ddbcc67642ab3728c4e9e2a4ef0e32fe280aac990d closes,pr,184793,issue,176211,high,pr.body,"ing failed sources, but reporting them keeps the recompile reason actionable and preserves information about the stale guard sources. Fixes #176211 Generated by my agent Test Plan: Reproduced the original issue script before the fix; it failed with TypeError: 'NoneType' object...",https://github.com/pytorch/pytorch/pull/184793,9920c9f2b930c37fdc448d92e88363e40d466ace2c97e4477a97c2e01f4b1560 closes,pr,184917,issue,173816,high,pr.body,all traceback preservation state in one TLS layer fixes the root cause and matches the existing regional-inductor TLS pattern nearby. Fixes #173816 Generated by my agent Test Plan: python test/fx/test_fx_traceback_tls.py python test/dynamo/test_aot_autograd.py -k test_split_wi...,https://github.com/pytorch/pytorch/pull/184917,5fd3f1b6dd34d51adf3c8c3949956b1a8ec7e72bcb30c613707424e8f5d576cc closes,pr,184928,issue,173052,high,pr.body,"after the fact. This patch instead fixes the broadcast shape before the comparison, which matches the documented tolerance semantics. Fixes #173052 Generated by my agent Test Plan: python test/dynamo/test_repros.py --use-pytest -k linalg_pinv_tensor_tolerances_compile lintrunn...",https://github.com/pytorch/pytorch/pull/184928,2fafac760cc9a101f7a8db6aad70636e29e0ac3a83fb4c1dbecafe46a77f258a closes,pr,184009,issue,167070,high,pr.body,erage for AOT C++ wrapper handling so SymBool inputs use their runtime graph input name instead of leaking internal unbacked symbols. Fixes #167070 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/184009,6454c8b3ec0a1b8ba1ed865320f91501f773ffcb2fa312ced971ce4b4e2264a3 review guidance,pr,184009,issue,167070,high,pr.reviews[0].body,How are the bools used ? would it be simpler to just codegen it as an int?,https://github.com/pytorch/pytorch/pull/184009,f635fa2ea2f64f4b43f28bacdc1719b5ffe19d1a08c75dc4961f4037b28c7ded review guidance,pr,184009,pr,184009,high,pr.reviews[0].body,How are the bools used ? would it be simpler to just codegen it as an int?,https://github.com/pytorch/pytorch/pull/184009,6d4f19d785ccd897df95db7d01d5373c5f8c8a7f12da624f522ce0fe1e7ca264 closes,pr,183862,issue,179368,high,pr.body,"al fake propagation, so layout-optimized SDPA/conv backward graphs do not fail before lowering can apply fallback layout constraints. Fixes #179368 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...",https://github.com/pytorch/pytorch/pull/183862,1fae16f5899ac20d40f85f4b74aaefe5a58434e3f7373c80216014d955839abd review guidance,pr,183862,issue,179368,high,pr.reviews[0].body,"Hmm, this is a general issue. For operators that require particular strides, we capture the operators with correct input strides, then modify the graph such that their input strides change, and then re-propagate and error. It's only an error within the FakeTensorPropagation. Within the inductor l...",https://github.com/pytorch/pytorch/pull/183862,40806f9ed24a8f45ec26a8b11f8f2e994d6eec91f201aa607ac00f5d525ec0ae review guidance,pr,183862,pr,183862,high,pr.reviews[0].body,"Hmm, this is a general issue. For operators that require particular strides, we capture the operators with correct input strides, then modify the graph such that their input strides change, and then re-propagate and error. It's only an error within the FakeTensorPropagation. Within the inductor l...",https://github.com/pytorch/pytorch/pull/183862,cf1730d34dab2997b48cedab684b4477d0bd8276de39dd9e0b7a208b87718b2f review guidance,pr,183862,issue,179368,high,pr.reviews[1].body,"Is the problem that when FakeTensorUpdator runs the strides can now be wrong? If so I think the principled solution is for FakeTensorUpdator to take into account the inductor stride requirements. If the stride says requires_contiguous, then we force the inputs to be contiguous and during FakeTens...",https://github.com/pytorch/pytorch/pull/183862,d3a13fa1509d27046ef5954077449a779cc5e90c40a3e52492b58c674bf0325b review guidance,pr,183862,pr,183862,high,pr.reviews[1].body,"Is the problem that when FakeTensorUpdator runs the strides can now be wrong? If so I think the principled solution is for FakeTensorUpdator to take into account the inductor stride requirements. If the stride says requires_contiguous, then we force the inputs to be contiguous and during FakeTens...",https://github.com/pytorch/pytorch/pull/183862,9994c006f6ec93a205c679790622345e2c60f6b6b441395e62ee27806a5cfd8c closes,pr,185059,issue,167007,high,pr.body,"ubclass-defined attrs from dir(instance), so this patch preserves that behavior and extends it with the canonical inner tensor names. Fixes #167007 Generated by my agent Test Plan: python test/export/test_export.py TestExport.test_nonstrict_subclass_attr_access_in_torch_functi...",https://github.com/pytorch/pytorch/pull/185059,73b9a548c3ba2da1b684ee32f5672e6f5e5b4c3532192b9ac78eddeeb0db2034 closes,pr,184015,issue,166604,high,pr.body,n the scalar was optimized out. This avoids crashing the randperm indexing pattern when slice_shape is absent from the initial match. Fixes #166604 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/184015,a2da20f4f2082eb0e22dba73ea7fd4cf3d20526e32807948bb9af0cfe09738df review guidance,pr,189183,pr,189183,high,pr.reviews[0].body,Add a test,https://github.com/pytorch/pytorch/pull/189183,80e674b3a2745b5b5b25e1f11a50fed050aa8e801b72a2b2b9fef13505fde683 closes,pr,185063,issue,166525,high,pr.body,"Stack from ghstack (oldest at bottom): -> #185063 Fixes #166525 WhileLoop.create assumed every carried and additional input from the outer higher order op node was an FX node with meta[""val""]. A Python s",https://github.com/pytorch/pytorch/pull/185063,235adae8587262e889ad919bf89c29ad5006e1f26a3a14a2b3238b298be90a07 closes,pr,188966,issue,116396,high,pr.closingIssuesReferences,pr #188966 declares a closing reference to issue #116396.,https://github.com/pytorch/pytorch/pull/188966,4bc21183adac2073b0f4f0a6c539009e7cfc66f33a7bb69af2e9b33bcdc795c7 closes,pr,188966,issue,116396,high,pr.body,"Fixes #116396 Summary operator.concat and operator.iconcat were missing from Dynamo's builtin operator registries, causing graph breaks when these operat",https://github.com/pytorch/pytorch/pull/188966,a2fd89e7deaf64ddfe9c395655e12d946f967aaf5f9c24a83bd72fa07d97fd55 review guidance,pr,188966,issue,116396,high,pr.reviews[0].body,"💡 Codex Review Here are some automated review suggestions for this pull request. Reviewed commit: 846e7cb836 ℹ️ About Codex in GitHub Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you Open a pull request for review Mark a draft as ready Comment ""@code...",https://github.com/pytorch/pytorch/pull/188966,a19e70aef75233841faaab1ea07de395a49b3749bff40c249ddf0182da9190c2 review guidance,pr,188966,issue,116396,high,pr.reviews[1].body,Add tests to test/dynamo and check if the CI failures are related,https://github.com/pytorch/pytorch/pull/188966,6ce1b89483b8ec0e3ad0fd69877a8c6aa89bb468fec0e1e8c4d392bca62d3930 closes,pr,185064,issue,166516,high,pr.body,Stack from ghstack (oldest at bottom): -> #185064 Fixes #166516 The root cause of the reported scan performance gap was that scan had no way to request unrolled code generation. Dynamo always captured sc,https://github.com/pytorch/pytorch/pull/185064,3d036d75e25f906b7919ec4ffeec49e3af9c07ba01fad962cf564fe25566bf70 references,pr,185064,issue,166516,medium,pr.review_comments[5].body,"o the single-step `while_loop` for `unroll=2`/`10`/`True`, every assertion in this test would still pass. Since the point of this PR (issue #166516) is the unrolled codegen, consider also asserting on the generated structure, e.g. `run_and_get_code` and count `while_loop` call...",https://github.com/pytorch/pytorch/pull/185064,66f56af6b17a61d18ba707f7d60cc2b055f326b956fc7b2bffeae9508960b242 review guidance,pr,185064,issue,166516,high,pr.reviews[0].body,"(Claude Review) [blocker] The unroll tests only assert eager-vs-compiled numerical equality, which stays green even if unrolling silently no-ops -- for a performance-only change that needs a codegen-structure check before merge (see the comment on test_control_flow.py). The rest are cleanups: the...",https://github.com/pytorch/pytorch/pull/185064,0b47644eef53b6ecca721fb09af1fafffd934f40553d30b9656d3dced54e8a28 review guidance,pr,185064,pr,185064,high,pr.reviews[0].body,"(Claude Review) [blocker] The unroll tests only assert eager-vs-compiled numerical equality, which stays green even if unrolling silently no-ops -- for a performance-only change that needs a codegen-structure check before merge (see the comment on test_control_flow.py). The rest are cleanups: the...",https://github.com/pytorch/pytorch/pull/185064,b555fe0523f803939816b8565af34db0fb4198bb60646caae9ddfe2a8c348f5e review guidance,pr,185064,issue,166516,high,pr.reviews[2].body,"(Claude Review) [blocker] The unroll tests only assert eager-vs-compiled numerical equality, which stays green even if unrolling silently no-ops -- for a performance-only change that needs a codegen-structure check before merge (see the comment on test_control_flow.py). The rest are cleanups: the...",https://github.com/pytorch/pytorch/pull/185064#pullrequestreview-4438510731,25a671d5c5236973478bdfc07ad7ef80e8277505e86b761f5ef0612d33bfcf2b review guidance,pr,185064,pr,185064,high,pr.reviews[2].body,"(Claude Review) [blocker] The unroll tests only assert eager-vs-compiled numerical equality, which stays green even if unrolling silently no-ops -- for a performance-only change that needs a codegen-structure check before merge (see the comment on test_control_flow.py). The rest are cleanups: the...",https://github.com/pytorch/pytorch/pull/185064#pullrequestreview-4438510731,e3dfe3145039dab5df11c2df84134aa6d25a910bfa77a65ac448cbace25ea0c7 closes,pr,185800,issue,147486,high,pr.body,eeps its existing behavior. This addresses the root cause in the logging setup rather than special-casing remote_cache's atexit hook. Fixes #147486 Generated by my agent Test Plan: pytest test/inductor/test_remote_cache.py -k test_dump_cache_stats_after_stderr_capture_closed p...,https://github.com/pytorch/pytorch/pull/185800,c80556f33e508c8fc3fb45362784f045d07a4ba560234d1da669c786a4b2e6f3 closes,pr,185908,issue,149583,high,pr.body,Stack from ghstack (oldest at bottom): -> #185908 Fixes #149583 EasyDict constructs objects by iterating self.__class__.__dict__.keys(). Dynamo previously either failed to trace the mappingproxy keys cal,https://github.com/pytorch/pytorch/pull/185908,22fe94e850aa849d8eaf7763d8c4afbe17236529fd00799a8f83e5e92ef052ce closes,pr,185085,issue,164922,high,pr.body,"eeps the change in Dynamo, avoids backend-specific lowering work, and matches the maintainer direction from the abandoned earlier PR. Fixes #164922 Generated by my agent Test Plan: python test/dynamo/test_unspec.py -k datetime_now python test/dynamo/test_unspec.py -k random li...",https://github.com/pytorch/pytorch/pull/185085,4e8e2be2aa9e1f565f7113df0003eeb46340af5ee2f9126f78fa71a699e00830 closes,pr,185088,issue,164559,high,pr.body,ver dropout. The test verifies that both the forward and backward compiler graphs no longer contain graphsafe RNG state placeholders. Fixes #164559 Generated by my agent Test Plan: python test/functorch/test_aot_joint_with_descriptors.py -k test_export_compile_rng_hop_has_no_g...,https://github.com/pytorch/pytorch/pull/185088,18827327289f164412d42fabcb5a17b582071f51544df4759c95b33d91a34281 closes,pr,186801,issue,186796,high,pr.body,n-TensorWeakRef weakrefs. Ordinary user weakrefs remain blocked because they can observe a tensor object whose contents were swapped. Fixes #186796 Generated by my agent Test Plan: python test/test_torch.py -k test_swap_allows_tensor_weakref python test/test_torch.py TestTorch...,https://github.com/pytorch/pytorch/pull/186801,4f80592727eb5327e54b32b67387fbece9acd1b21155cd22cba7c3bdb079e0df closes,pr,160585,issue,156075,high,pr.closingIssuesReferences,pr #160585 declares a closing reference to issue #156075.,https://github.com/pytorch/pytorch/pull/160585,b774591d49024d78aee8375dcbc3f3b130241a6729c91591720761b0a2084fcd closes,pr,160585,issue,156075,high,pr.body,"Fixes #156075 Update: Rebased and brought to date. In order to support dim=None for torch.logsumexp, I followed the pattern used by torch.sum and torch.m",https://github.com/pytorch/pytorch/pull/160585,36421f8b00b0cd91e9997a4ced1052b96998948c12c1061be3e11a0f739b8e94 references,pr,160585,issue,174929,medium,pr.comments[0].body,"possibly due to flakiness on trunk: inductor / inductor-cpu-test / test (cpu_inductor_torchbench, 1, 2, linux.2xlarge.amx, unstable) (gh) (#174929) detectron2_maskrcnn_r_50_fpn This comment was automatically generated by Dr. CI and updates every 15 minutes.",https://github.com/pytorch/pytorch/pull/160585,43eb37b0d77207536028e4f16487853a4a377b806b02e3df320fa4c2b9237fef review guidance,pr,160585,issue,156075,high,pr.reviews[0].body,Amazing work! Thanks for going through all the different systems and updating them!,https://github.com/pytorch/pytorch/pull/160585,4efaa046ccc24091a5a13a2a04706c57fe6ba58ca79a69afc0ee97313d9d8600 review guidance,pr,188480,pr,188480,high,pr.reviews[0].body,Can you add tests for the linter? Like the GB_REGISTRY linter.,https://github.com/pytorch/pytorch/pull/188480,1a8eefca299d93e28973aaf2d87fca88840d0d0c65202ae616a07c5123a31ebf review guidance,pr,188480,pr,188480,high,pr.reviews[1].body,"LGTM, thanks.",https://github.com/pytorch/pytorch/pull/188480,eb33bd83c0faeb202fca3db84db96105eef1b9bbf41b3a8dc06fbd7cae86ed3d closes,pr,183994,issue,172942,high,pr.body,CPU tensor inputs. Preserve NumPy dtype semantics for tensor factories and keep CUDA graphs free of scalar DeviceCopy partitions.\n\nFixes #172942\nGenerated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/183994,d00a4acb9b7d36ad048b76ee35442a3555cb505edd8c8dad3bf84c730a228281 review guidance,pr,183994,issue,172942,high,pr.reviews[0].body,Would it be better to copy these to pinned memory and convert to cuda ? We do already have a path for handling this for 0-dim tensors in post grad,https://github.com/pytorch/pytorch/pull/183994,70dba5d647eb05b656e09ef2a7d3d1fda019f4ba5906397c943bb6bc91712915 review guidance,pr,183994,pr,183994,high,pr.reviews[0].body,Would it be better to copy these to pinned memory and convert to cuda ? We do already have a path for handling this for 0-dim tensors in post grad,https://github.com/pytorch/pytorch/pull/183994,91e24540691658af88b1479d2e5c439e374f4bf358f0460ce40aac47dcc4523c review guidance,pr,183994,issue,172942,high,pr.reviews[1].body,The problem is that we're not giving different behavior from without cudagraphs. and we can also partition out what we dont use.,https://github.com/pytorch/pytorch/pull/183994,184e1c33e9750f54737d18004dc621b3761db33d0abf9f796689c53c81105ceb review guidance,pr,183994,pr,183994,high,pr.reviews[1].body,The problem is that we're not giving different behavior from without cudagraphs. and we can also partition out what we dont use.,https://github.com/pytorch/pytorch/pull/183994,ce5ec1561635f9f0e0ec0e30b5804bb6a05bb1692c09385eb84098e00badb6ed closes,pr,186764,issue,171905,high,pr.closingIssuesReferences,pr #186764 declares a closing reference to issue #171905.,https://github.com/pytorch/pytorch/pull/186764,f718dcbdc0ef8cfc79b4582b34bfad222cbce26426b381fb631a8b9befc3361a closes,pr,186764,issue,171905,high,pr.body,"ved the ""can't be public as is because setting the module directly doesn't work"" comment, as that limitation no longer applies. Issue Fixes #171905 Issue: #171905 Diffstat torch/ao/quantization/__init__.py | 5 +++-- torch/ao/quantization/utils.py | 4 +--- 2 files changed, 4 in...",https://github.com/pytorch/pytorch/pull/186764,8d037bb9d277ccc48a0ca168f83a3eb742c1d6862b0c072862b4cc64ea277205 review guidance,pr,186764,issue,171905,high,pr.reviews[0].body,.,https://github.com/pytorch/pytorch/pull/186764,affcfc04f49f0362d9fbe1099d6ea5d83a3f61af460aa91ab1d95789a64cd627 closes,pr,186913,issue,183901,high,pr.body,"Stack from ghstack (oldest at bottom): -> #186913 Fixes #183901 Inductor already distinguishes value-producing symbolic expressions from pure indexing expressions through value_expr, but Triton still pri",https://github.com/pytorch/pytorch/pull/186913,500b24421d4a8763d7247fd283dbd0fa0fce48a3c6598b1c1b853458bfaed221 review guidance,pr,186913,issue,183901,high,pr.reviews[0].body,cc @laithsakka who said he was interested in doing typed sympy expressions. can you review this?,https://github.com/pytorch/pytorch/pull/186913,34f2537bda3f949d7a326d800b8a88e34804dc244bc775886921f0d9974c4761 review guidance,pr,186913,pr,186913,high,pr.reviews[0].body,cc @laithsakka who said he was interested in doing typed sympy expressions. can you review this?,https://github.com/pytorch/pytorch/pull/186913,a35c5205a6c6b6a6023c6f9a36beaf17b2f8863ae694ec144e60d2ca5dabf576 closes,pr,188327,issue,174183,high,pr.closingIssuesReferences,pr #188327 declares a closing reference to issue #174183.,https://github.com/pytorch/pytorch/pull/188327,529feb83ee920a335825518508565714f881d2fd656e37c23445f3dc1803aa27 closes,pr,188327,issue,174183,high,pr.body,"> ValueError: in_channels must be divisible by groups Follows the pattern already used by BatchNorm, GroupNorm, and RNN cell modules. Fixes #174183 Test Plan python test/test_nn.py -k test_module_error_inputs_torch_nn_Conv2d",https://github.com/pytorch/pytorch/pull/188327,6bbae0a78cb2c89464b2cf21de625e0e602480223ddd200be28bd81f1c87fb05 references,pr,182376,issue,181801,medium,pr.body,For #181801 authored with codex cc @ptrblck @msaroufim @jerryzh168 @tinglvv @nWEIdia @jgong5 @mingfeima @XiaobingSuper @sanchitintel @ashokei @jingxu10,https://github.com/pytorch/pytorch/pull/182376,3d53f80663d6eae61461e238d67b3e0b0c3cc0e5950e99e7d186964587f930be review guidance,pr,182376,issue,181801,high,pr.reviews[0].body,give me my testsssssss,https://github.com/pytorch/pytorch/pull/182376,641510e580aa02e68e91adf868465a9eb8d9bef3bb4f40e89a72d97c283abec2 closes,pr,185751,issue,184845,high,pr.closingIssuesReferences,pr #185751 declares a closing reference to issue #184845.,https://github.com/pytorch/pytorch/pull/185751,305f0fe61ea458a4d109a8cf37f560c53114d206c8430950fed5252453bc562b closes,pr,185751,issue,184845,high,pr.body,runs after the torch_function dispatch path so custom tensor subclasses are unaffected. constant mode is unconstrained and untouched. Fixes #184845 Checklist Passes lint (lintrunner torch/nn/functional.py test/test_nn.py) Added/updated tests Updated documentation (if applicabl...,https://github.com/pytorch/pytorch/pull/185751,ab50b7bed40d48384a415ed4a1ea6712a504840f60c8df565f691f8769517bd9 closes,pr,187908,issue,187429,high,pr.closingIssuesReferences,pr #187908 declares a closing reference to issue #187429.,https://github.com/pytorch/pytorch/pull/187908,47b7f38d3693859f537a1585307e242b7f5f0c8d21afb923dfe06b9487bdcdf0 closes,pr,187908,issue,187429,high,pr.body,"Fixes #187429 Root cause Commit 0eae6b68f42 (""Unify torch.tensor and torch.ops.aten.scalar_tensor behavior"") intentionally relaxed the cast in fill_inpla",https://github.com/pytorch/pytorch/pull/187908,02b459deec06d55832348b63636e92651c6b44bcb16a4ae24e6cbc6976040da9 review guidance,pr,187908,issue,187429,high,pr.reviews[0].body,"Please make sure this also works for float8/float4 by including a test: - aten/src/ATen/native/TensorCompare.cpp:653 and :666 — Regression for float8/float4 result types (must fix). The guard is isReducedFloatingType(result_type), which (c10/core/ScalarType.h:105-108) returns true for Half, BFloa...",https://github.com/pytorch/pytorch/pull/187908,442e3c61cf272a69901b8b0e5daf2cc9645bddbadb3e763bc49a1f15650fbebb closes,pr,185756,issue,171356,high,pr.closingIssuesReferences,pr #185756 declares a closing reference to issue #171356.,https://github.com/pytorch/pytorch/pull/185756,1e12e2587620a6d2b341ad3dc420c0226de0a07f8d361082aeeabe9d30fcc531 closes,pr,185756,issue,171356,high,pr.body,"the same bug for their backend, confirming that a CUDA-only fix is insufficient. Our meta-function approach fixes XPU automatically. Fixes #171356 Checklist Passes lint (lintrunner aten/src/ATen/native/TensorCompare.cpp test/test_shape_ops.py) Added/updated tests Updated docum...",https://github.com/pytorch/pytorch/pull/185756,9f0e291c03080a07462a730998a6d3d657ad1efc2854245e6fb8c690e36be918 references,pr,185756,pr,187908,medium,pr.body,so limiting the guard lets those dtypes surface their normal clamp not implemented kernel error. This matches the same narrowing applied in #187908. bfloat16 with max=65507 correctly does not raise since 65507 is representable in bfloat16 (same exponent range as float32). Prio...,https://github.com/pytorch/pytorch/pull/185756,a4ecd973ed4cc926159d6b7074e543f94090368124b59316461ded1ae04575f5 closes,pr,187614,issue,186028,high,pr.body,d. This keeps the existing linalg implementation and avoids adding a Dynamo or Inductor workaround for a native operator shape query. Fixes #186028 Generated by my agent Test Plan: ninja -C build torch_cpu issue repro before fix: dynamic=True failed with Cannot call numel() on...,https://github.com/pytorch/pytorch/pull/187614,98d2ee39522b8cc11e546d7d0ba4f55f05bc73054cef192c6ea2df2110134b38 closes,pr,173948,issue,173252,high,pr.closingIssuesReferences,pr #173948 declares a closing reference to issue #173252.,https://github.com/pytorch/pytorch/pull/173948,d42e0416994e9b2584e56e9cbe8d9af1b585998f29829b206293ad481bbdf10d closes,pr,173948,issue,173252,high,pr.body,Fixes #173252 The issue was that lazy modules tried to use SymInt dimensions directly for materialization which requires concrete integers. Now converts,https://github.com/pytorch/pytorch/pull/173948,4b844abb534f56f45f3b53f40ad17d6a48a0e08be71a1ff55b95942ce6c894e0 review guidance,pr,176072,pr,146018,high,pr.reviews[0].body,"Pull request overview Improves TorchInductor’s Triton launcher invocation diagnostics by validating launcher argument counts up-front, preventing confusing/incorrect TypeError failures when argument mismatches occur. Changes: Add CachingAutotuner._validate_launcher_args() and invoke it before lau...",https://github.com/pytorch/pytorch/pull/176072,818a7b710cc1c2f13f24a44df7c3709f55a38a264268164d38a21fc1b62079e5 review guidance,pr,176072,pr,174534,high,pr.reviews[0].body,"Pull request overview Improves TorchInductor’s Triton launcher invocation diagnostics by validating launcher argument counts up-front, preventing confusing/incorrect TypeError failures when argument mismatches occur. Changes: Add CachingAutotuner._validate_launcher_args() and invoke it before lau...",https://github.com/pytorch/pytorch/pull/176072,650194f1e9966ea2753a52b2b77fd6e5a04ad22a79faa4e384ec4e66c7461b02 review guidance,pr,176072,pr,175228,high,pr.reviews[0].body,"Pull request overview Improves TorchInductor’s Triton launcher invocation diagnostics by validating launcher argument counts up-front, preventing confusing/incorrect TypeError failures when argument mismatches occur. Changes: Add CachingAutotuner._validate_launcher_args() and invoke it before lau...",https://github.com/pytorch/pytorch/pull/176072,dfe4026039b7bcbf7716f37d923d8fd54a46820dc2bbcc4bd09ffd108bd21835 review guidance,pr,176072,pr,146018,high,pr.reviews[1].body,Test failures?,https://github.com/pytorch/pytorch/pull/176072,ceaaf383b799904ef88d3108e0735dc1d2ebf3832001d0838b17b5444aaab09b review guidance,pr,176072,pr,174534,high,pr.reviews[1].body,Test failures?,https://github.com/pytorch/pytorch/pull/176072,ac5772b0a48ef681847a797af39711cf1d23ca62b7aa660a8fb8556539a9b212 review guidance,pr,176072,pr,175228,high,pr.reviews[1].body,Test failures?,https://github.com/pytorch/pytorch/pull/176072,4e8bdedc1592f9ac7d3a01dcb2b54dde852825e02c2a9f783d4d5526f4c7a4e3 closes,pr,183652,issue,183259,high,pr.body,"Inductor views without requiring stride metadata, and fall back for unresolved slice_scatter bounds so eager validation is preserved. Fixes #183259 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...",https://github.com/pytorch/pytorch/pull/183652,f1b4c0f1f2a34195325aeaa5a2df066512881a6418deab41db118d4da90cd190 review guidance,pr,183652,issue,183259,high,pr.reviews[1].body,legit mosty except for _slice_size_from_meta,https://github.com/pytorch/pytorch/pull/183652,fa52a6f33b9a872dca46768bdfd000a3ac69362807267652265b9282108af00f review guidance,pr,183652,pr,183652,high,pr.reviews[1].body,legit mosty except for _slice_size_from_meta,https://github.com/pytorch/pytorch/pull/183652,d0915240fc4a604b4f372083a66527b96561402454e17faef105dbacfdb5cfc0 references,pr,181109,issue,141550,medium,pr.body,"3: ""Generate and maintain test_times.json for RISC-V CI sharding"" — accepted RISC-V enablement tracking issue). Prior context: RFC #171659, #141550. cc @mruberry @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv @jiayisunx...",https://github.com/pytorch/pytorch/pull/181109,5f7f3466934e3da31df313ed42806bf41c75ff6d9f82aec5bc8a84238cb9fb3f references,pr,181109,issue,171659,medium,pr.body,"(Phase 1.3: ""Generate and maintain test_times.json for RISC-V CI sharding"" — accepted RISC-V enablement tracking issue). Prior context: RFC #171659, #141550. cc @mruberry @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv @...",https://github.com/pytorch/pytorch/pull/181109,c69655ae408e372747c51c724a0020bed511663a6c7b9c1370e28885567f8221 references,pr,181109,issue,180975,medium,pr.body,"l skip), missing `time` attribute, nested subdirectory edge case, CLI exit codes, and full end-to-end JSON structure validation. Related to #180975 (Phase 1.3: ""Generate and maintain test_times.json for RISC-V CI sharding"" — accepted RISC-V enablement tracking issue). Prior co...",https://github.com/pytorch/pytorch/pull/181109,81efeaa0a657200f19eee1853c6a0f035a888c2bec591c087a2b7068db8b5d08 review guidance,pr,181109,issue,141550,high,pr.reviews[0].body,@CodersAcademy006 Thank you for your work. It's been very helpful for the RISC-V CI. I've raised a few minor issues—hope they are helpful to you.,https://github.com/pytorch/pytorch/pull/181109,8c3b764a3409f7300fdcef44a52ad23cc0fc4239840d5a8e2e8a25ffe32dc6d8 review guidance,pr,181109,issue,171659,high,pr.reviews[0].body,@CodersAcademy006 Thank you for your work. It's been very helpful for the RISC-V CI. I've raised a few minor issues—hope they are helpful to you.,https://github.com/pytorch/pytorch/pull/181109,f797d82618396b5e26c058ded96b21c3b0e14c494480b658c5116fb14c2f02d7 review guidance,pr,181109,issue,180975,high,pr.reviews[0].body,@CodersAcademy006 Thank you for your work. It's been very helpful for the RISC-V CI. I've raised a few minor issues—hope they are helpful to you.,https://github.com/pytorch/pytorch/pull/181109,3396c3cf29e1080bff0b10bef7d8490fd7f314904956f59afd45766770d68499 review guidance,pr,181109,issue,141550,high,pr.reviews[1].body,All my feedback has been addressed. Nice work — LGTM.,https://github.com/pytorch/pytorch/pull/181109,5bf2679d163dc78d7e060a2faddc50d16f07d5a9da2b900dcd7230921c0ec08c review guidance,pr,181109,issue,171659,high,pr.reviews[1].body,All my feedback has been addressed. Nice work — LGTM.,https://github.com/pytorch/pytorch/pull/181109,2f23fc0e8835bf31881a56f09883a39c9b685147b07658284da823257842bf32 review guidance,pr,181109,issue,180975,high,pr.reviews[1].body,All my feedback has been addressed. Nice work — LGTM.,https://github.com/pytorch/pytorch/pull/181109,4a31fab3bf65656b474ad52585b98b2564989886c2e68e024cc03417260ca004 references,pr,182377,issue,182382,medium,pr.body,"s an opt-in path for DTensor to allocate local tensor shards using SymmetricMemory. This is part of the one-sided DTensor work described in #182382. Allocating local tensor shards using SymmetricMemory makes them remotely addressable, enabling one-sided communication operation...",https://github.com/pytorch/pytorch/pull/182377,de6b10e2bd8d037667e71011c244c6698309345ee54e03db19e0983500f00f8e closes,pr,184934,issue,173046,high,pr.body,"d rnn_relu, gru, lstm, and quantized packed variants, so validating at the common packed entry points is the narrower root-cause fix. Fixes #173046 Generated by my agent Test Plan: PYTHONPATH=/data/users/jansel/pytorch-issue-fixer/pytorch ninja -C build PYTHONPATH=/data/users/...",https://github.com/pytorch/pytorch/pull/184934,81b3b27033f5a9cafed13c23ecc56d469a2490faaecc60c25b547aaeac99c796 closes,pr,184938,issue,172956,high,pr.body,t-backward mutation as the forward output. Keeping the forward value and the backward side effect separate preserves eager semantics. Fixes #172956 Generated by my agent Test Plan: python test/dynamo/test_autograd_function.py -k test_backward_mutation_saved_tensor_returned_as_...,https://github.com/pytorch/pytorch/pull/184938,df8607a963b45f65bb874f4320083255f1f6138e033db51d78d5e26afa343c70 references,pr,184938,issue,172956,medium,pr.reviews[0].body,"er would help future readers. [cleanup] Test coverage for the new functionality is also narrow. The new test covers the canonical case from #172956. Worth adding: An aliased-output case (return saved, saved or return saved[:k], saved[k:]) that exercises what the new loop does...",https://github.com/pytorch/pytorch/pull/184938,aa3dc13f3d02a9b3a85fbce873393e57f98a3ca73a0c2f433316d300a5909d54 review guidance,pr,184938,issue,172956,high,pr.reviews[0].body,"(Claude Review) [blocker] CI is red on the AOTAutograd input-mutation suite, and despite Dr. CI's ""44 unrelated failures (unstable)"" framing the failures look caused by this PR. Across the last 100 commits on main, the following four tests fail in zero commits: test/functorch/test_aotdispatch.py:...",https://github.com/pytorch/pytorch/pull/184938,63290552f7b091217eb857b4f96c09b05aea55f81f7071e3e6148636b9d380c4 review guidance,pr,184938,pr,184938,high,pr.reviews[0].body,"(Claude Review) [blocker] CI is red on the AOTAutograd input-mutation suite, and despite Dr. CI's ""44 unrelated failures (unstable)"" framing the failures look caused by this PR. Across the last 100 commits on main, the following four tests fail in zero commits: test/functorch/test_aotdispatch.py:...",https://github.com/pytorch/pytorch/pull/184938,e81a37ccc8a952169f080a947d1c3e66b5f91ad9299c48d7b4a79b0a73b8111f closes,pr,184950,issue,172711,high,pr.body,"e tensors before calling the generated extern kernel. Updating both paths keeps the fake layout and runtime/meta behavior consistent. Fixes #172711 Generated by my agent Test Plan: ninja -C build lib/libtorch_cpu.so TORCHINDUCTOR_FX_GRAPH_CACHE=0 python standalone issue repro,...",https://github.com/pytorch/pytorch/pull/184950,2c496d1bf8b506b0d11da320148904f61a86247c381b11237c7bfb75722ab244 references,pr,184950,issue,172711,medium,pr.comments[2].body,py`) **`test_conv_transpose_pointwise_regular_weight_shape`:** Good end-to-end regression test that directly reproduces the original issue (#172711). Tests both shape correctness and numerical closeness between eager and compiled. **`test_conv_transpose_pointwise_fake_meta_sha...,https://github.com/pytorch/pytorch/pull/184950,7e03754ae2169123eba6da8f4d3640ab0dc163ca871d75ddf916ac45d530008d closes,pr,184747,issue,177712,high,pr.body,tch. Canonicalize structurally proven all-true boolean masks and zero additive masks to None before eager/Inductor backend execution. Fixes #177712 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/184747,9b642900b3c8d45318b82485be672ccbc3db6a2a57a4b705f36baa7f588c038b closes,pr,183879,issue,177327,high,pr.body,"tion and int8 WOQ weights, and use portable generated int8 load/prefetch helpers so the CPU GEMM template can be selected on AArch64. Fixes #177327 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...",https://github.com/pytorch/pytorch/pull/183879,063156135d4c7e734e8a2cc8376bfcf9c65e3fc245b20f88d7b4f8d0d3686f85 closes,pr,183881,issue,177118,high,pr.body,"contiguous pointwise reads are compatible recomputations of reduction inputs, while preserving fully non-contiguous softmax behavior. Fixes #177118 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...",https://github.com/pytorch/pytorch/pull/183881,eb72c9385788b9e5e799eb7c6a4c5193dac7590376c4c6116cfd1fc6a85b4eab review guidance,pr,183881,issue,177118,high,pr.reviews[4].body,defering review on perf speedup until ive done the other reviews.,https://github.com/pytorch/pytorch/pull/183881#pullrequestreview-4377317826,0bdaf2b4c92992d4698c669c6856c17aeef48b76ec4acb7ee7516b9542023366 review guidance,pr,183881,pr,183881,high,pr.reviews[4].body,defering review on perf speedup until ive done the other reviews.,https://github.com/pytorch/pytorch/pull/183881#pullrequestreview-4377317826,2a6ca21d7a494083345c5c81dc9c09589c0abb275a33d894c519a948593097e5 closes,pr,183908,issue,154146,high,pr.body,"d of producing Python with undefined names, while missing unbacked symbols continue to defer until their binding point. Fixes #176770 Fixes #154146 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...",https://github.com/pytorch/pytorch/pull/183908,0d3ee97558ac4cee37e4b72de76f7077ccbf225c73b9d01f3465a26cf7db240b closes,pr,183908,issue,176770,high,pr.body,"skipped instead of producing Python with undefined names, while missing unbacked symbols continue to defer until their binding point. Fixes #176770 Fixes #154146 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzhe...",https://github.com/pytorch/pytorch/pull/183908,b304b1baa4005bbe7e908cf1b14ed72bedc14d8c99d1af4cbf6757b4f05814a7 review guidance,pr,183944,pr,176112,high,pr.reviews[1].body,This is extremely deeply nested branching code. I do not understand it. Even the agent review sees potential issues. Are you happy with this code?,https://github.com/pytorch/pytorch/pull/183944,b64ff56611326f46ddf9b40ee6bc7257b26386c396ae21c0ee36b13ba86e3344 review guidance,pr,183944,pr,183944,high,pr.reviews[1].body,This is extremely deeply nested branching code. I do not understand it. Even the agent review sees potential issues. Are you happy with this code?,https://github.com/pytorch/pytorch/pull/183944,55b7a632fe8cefcae26c3be943467fac8a50004e13e586f26fb9a42306264dfc closes,pr,183954,issue,175968,high,pr.body,"stants remain explicit: opaque modules/class singletons, Generators used by factory kwargs, and compiler-owned _LeafCallable objects. Fixes #175968 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...",https://github.com/pytorch/pytorch/pull/183954,bacb96dad2d39666e1f9ca8100349e86ab7bf964c83745e8558ab9d5e5988c6b references,pr,183954,issue,175968,medium,pr.comments[2].body,"as an argument""). Good. 4. **Test coverage** — `test_reference_opaque_errors_on_custom_op_dispatch_creation` covers the exact scenario from #175968: a `__torch_dispatch__` handler creating a reference opaque mid-trace. The test structure (tensor subclass → custom op dispatch →...",https://github.com/pytorch/pytorch/pull/183954,45d72a26f0c8bc1a4626672ca237ef73c730009350fd13d7d9f017142c1f96cf references,pr,183954,issue,175968,medium,pr.comments[12].body,iven the internal usage. 3. **Test coverage is good** — `test_reference_opaque_errors_on_custom_op_dispatch_creation` faithfully reproduces #175968 (tensor subclass → custom op `__torch_dispatch__` → opaque creation mid-trace). `test_reference_opaque_subclass_creation_errors`...,https://github.com/pytorch/pytorch/pull/183954,c7ed61b3cb28a53be29a195b74e574678639c0db0603de0c1c9b43c6937c2e33 references,pr,183954,issue,175968,medium,pr.comments[15].body,check. 5. **Test coverage** — Five new tests cover the key scenarios: - `test_reference_opaque_errors_on_custom_op_dispatch_creation` — the #175968 reproduction - `test_reference_opaque_input_is_tracked` — positive case (tracked input is fine) - `test_reference_opaque_closure_...,https://github.com/pytorch/pytorch/pull/183954,dd062ac1bc4bd8de943caeab862dc76557e0562dcdf71834a140904de239e053 references,pr,183954,issue,175968,medium,pr.comments[18].body,"s), but worth noting as a theoretical concern. ### Test coverage The 9 new tests are well-designed and cover the key scenarios: - The exact #175968 reproduction (custom op dispatch → opaque creation) - Tracked input positive case - Closure capture without passing as input - Id...",https://github.com/pytorch/pytorch/pull/183954,a77326bd02a5d5185071af8af3d4524df18e91ced61b5b0e50c59321320afc71 references,pr,183954,issue,175968,medium,pr.review_comments[5].body,My agent says agreed. I kept this PR scoped to the correctness fix for #175968; the monkey-patch/perf cleanup is a reasonable follow-up.,https://github.com/pytorch/pytorch/pull/183954,86537768b0ea2dc92f997ec38c48849f0fe6710c66ea68313174ecfe213053de review guidance,pr,183954,issue,175968,high,pr.reviews[0].body,Approving w/ comments,https://github.com/pytorch/pytorch/pull/183954,b34dfa73f7afb08359104ef33fb3a84cf01061027ddb8ae96214f6138955e25f review guidance,pr,183954,pr,183954,high,pr.reviews[0].body,Approving w/ comments,https://github.com/pytorch/pytorch/pull/183954,847cee860f7cebe4ffcde011a3d67e36cf013f258084f7d2472e4f03b8975462 review guidance,pr,183954,issue,175968,high,pr.reviews[2].body,Approving w/ comments,https://github.com/pytorch/pytorch/pull/183954#pullrequestreview-4311951668,66c1fd8afddc576dbe9e7856d0d20a93a6f30db7a55803809c0b2163340c7453 review guidance,pr,183954,pr,183954,high,pr.reviews[2].body,Approving w/ comments,https://github.com/pytorch/pytorch/pull/183954#pullrequestreview-4311951668,1da23611668a7bb0c3b08e063e230f2a44d4e406a033c8ad7df1a80e42bed315 review guidance,pr,183954,issue,175968,high,pr.reviews[7].body,"requesting changes to litigate the implementation strategy. I did not understand why this implementation was taken, more details in the PR body would have helped",https://github.com/pytorch/pytorch/pull/183954#pullrequestreview-4323584981,ba78cb484ea987ed79287b7ffbb76b4d550f11818e76f0ec98e816f44267cd6d review guidance,pr,183954,pr,183954,high,pr.reviews[7].body,"requesting changes to litigate the implementation strategy. I did not understand why this implementation was taken, more details in the PR body would have helped",https://github.com/pytorch/pytorch/pull/183954#pullrequestreview-4323584981,25b4a8de1d8227d1727ac3b62c261d414552e326e13b74fb91ae3886d7ea5017 review guidance,pr,183954,issue,175968,high,pr.reviews[9].body,"I talked to my codex, it suggested the followign as one of the options: ``` 2. Reject untracked reference opaques in FX arg creation In the tracer path that turns Python objects into FX args, if a reference opaque has no tracked proxy/input source/reconstruct rule, error instead of emitting get_a...",https://github.com/pytorch/pytorch/pull/183954#pullrequestreview-4392941419,6df0b2e479a547ecb12ed927d0447d602435add1e20efd555fcf2cb73e4fa48d review guidance,pr,183954,pr,183954,high,pr.reviews[9].body,"I talked to my codex, it suggested the followign as one of the options: ``` 2. Reject untracked reference opaques in FX arg creation In the tracer path that turns Python objects into FX args, if a reference opaque has no tracked proxy/input source/reconstruct rule, error instead of emitting get_a...",https://github.com/pytorch/pytorch/pull/183954#pullrequestreview-4392941419,0201877b6c00ecbb4393d19a249a8ba5f0c9a6ccd705441ef2e9cfd45c6cc463 review guidance,pr,183954,issue,175968,high,pr.reviews[10].body,"your tests are very failing, are you looking for feedback from me?",https://github.com/pytorch/pytorch/pull/183954#pullrequestreview-4402367940,19e77fbc08f15e7db92bf07b40b5bfc5fc15ad35fb9578c48cfe68dea7fe72d5 review guidance,pr,183954,pr,183954,high,pr.reviews[10].body,"your tests are very failing, are you looking for feedback from me?",https://github.com/pytorch/pytorch/pull/183954#pullrequestreview-4402367940,610880ebed1ad4889d77c8566b0a238fe1299b3e296005a138b69d29312679d2 closes,pr,184803,issue,175858,high,pr.body,p checks skip guards. This version keeps setitem cheap while making pop/popitem conservative where replay needs the original key set. Fixes #175858 Generated by my agent Test Plan: python test/dynamo/test_dicts.py -k setitem python test/dynamo/test_dicts.py -k pop python test/...,https://github.com/pytorch/pytorch/pull/184803,ea29ec6e793156fab7a14e4ea508709ce1eb63c7c1730e30b4eecc55b3c467d0 closes,pr,183958,issue,175833,high,pr.body,s copies when shared K/V tensors hit the same layout requirement. Enable it by default and add a regression test for the issue repro. Fixes #175833 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/183958,21c2e26046ec2c481c40b4b5d6b1fe5aa9a5378f68f2502c12ed1db38046ba30 references,pr,183958,issue,175833,medium,pr.comments[1].body,"ONSTRAINT=0. test/inductor/test_fused_attention.py The new test _test_cache_sdpa_constraint_shared_kv_default is a good regression test for #175833. It verifies that with the default config, a shared K/V tensor (via torch.relu(kv) passed as both key and value) only produces a...",https://github.com/pytorch/pytorch/pull/183958,f58d37caeedac8d20fba1b0686c13054ee82612d7298bb1f9b46ff917fd8fbde closes,pr,186366,issue,129845,high,pr.body,"e old test commented out, but that would preserve the broken autograd.Function composition instead of fixing the shared capture path. Fixes #129845 Generated by my agent Test Plan: pytest test/dynamo/test_autograd_function.py -k 'test_mark_single_output_non_differentiable or t...",https://github.com/pytorch/pytorch/pull/186366,e97d95d075eabc310d8cfd2371368c639dc2b4512204c37a9552f2582c23dc29 references,pr,186366,issue,129845,medium,pr.comments[5].body,"ion The previously commented-out `self.run_test(score_mod, q.dtype, device)` is now enabled — this is the original motivating use case from #129845. #### Potential edge case The `is_under_vmap` detection in `autograd_function.py` at line 831–835 iterates `fwd_args` checking `i...",https://github.com/pytorch/pytorch/pull/186366,de7d6b8da13ea3b41fdb7e6244c7877207c6d5847609a14b5b081f53e807e351 review guidance,pr,186366,issue,129845,high,pr.reviews[0].body,"Pull request overview This PR fixes an interaction between torch.compile, torch.vmap, and “new-style” torch.autograd.Function (i.e., with setup_context), by ensuring the internal Dynamo-lowered ApplyTemplate Function correctly participates in generated vmap rules and correctly handles ctx.mark_no...",https://github.com/pytorch/pytorch/pull/186366,916bb9fcda0be1c6ed1a8f68545dc8abd77da766b75028c538e01f32f38d6aa7 review guidance,pr,186366,pr,186366,high,pr.reviews[0].body,"Pull request overview This PR fixes an interaction between torch.compile, torch.vmap, and “new-style” torch.autograd.Function (i.e., with setup_context), by ensuring the internal Dynamo-lowered ApplyTemplate Function correctly participates in generated vmap rules and correctly handles ctx.mark_no...",https://github.com/pytorch/pytorch/pull/186366,630b38c7f2d81463cc65cd30aac456037eb2840e8defe2daac83243ff1465947 review guidance,pr,186366,issue,129845,high,pr.reviews[2].body,"## Pull request overview This PR fixes an interaction between `torch.compile`, `torch.vmap`, and “new-style” `torch.autograd.Function` (i.e., with `setup_context`), by ensuring the internal Dynamo-lowered `ApplyTemplate` Function correctly participates in generated vmap rules and correctly handle...",https://github.com/pytorch/pytorch/pull/186366#pullrequestreview-4517491739,4084c4e0a0668f0ca1145f65ea4bf1a492b344ae72a157077d808b15aa318b9f review guidance,pr,186366,pr,186366,high,pr.reviews[2].body,"## Pull request overview This PR fixes an interaction between `torch.compile`, `torch.vmap`, and “new-style” `torch.autograd.Function` (i.e., with `setup_context`), by ensuring the internal Dynamo-lowered `ApplyTemplate` Function correctly participates in generated vmap rules and correctly handle...",https://github.com/pytorch/pytorch/pull/186366#pullrequestreview-4517491739,50b0ba786f84cec6f6f1d5ac7542bb9cd5b5336e0b0f1f5dd8a1e2f073f2d4b4 review guidance,pr,186366,issue,129845,high,pr.reviews[7].body,"(Reviewed by me, assisted by AI) [blocker] The runtime vmap-enforcement block added to `AutogradFunctionApply.__call__` in `torch/_functorch/autograd_function.py` (the `is_under_vmap` check that raises `RuntimeError`) is new, reachable, user-facing behavior with no test. The two trace-time tests...",https://github.com/pytorch/pytorch/pull/186366#pullrequestreview-4555451503,aef8864b7a8c73e4422328dde7b8d640c6e13dfb47da8e3bd5d5df9ab46b206e review guidance,pr,186366,pr,186366,high,pr.reviews[7].body,"(Reviewed by me, assisted by AI) [blocker] The runtime vmap-enforcement block added to `AutogradFunctionApply.__call__` in `torch/_functorch/autograd_function.py` (the `is_under_vmap` check that raises `RuntimeError`) is new, reachable, user-facing behavior with no test. The two trace-time tests...",https://github.com/pytorch/pytorch/pull/186366#pullrequestreview-4555451503,9f03a3551adcbf766258f42a292fc760f95a0bfe9e276a7c8e4edce1e151185b review guidance,pr,183681,pr,182093,high,pr.reviews[0].body,Looks good although I don't think we should be overriding default gil handling for sub process,https://github.com/pytorch/pytorch/pull/183681,83fda7b5955b6738527d203d7b30ca9789523ad99bcd763a5f76bbd2084346bd review guidance,pr,183681,pr,183681,high,pr.reviews[0].body,Looks good although I don't think we should be overriding default gil handling for sub process,https://github.com/pytorch/pytorch/pull/183681,0645fca61cb0186753211811e89942e77574aae9bb172398763a6ee9bb553edb closes,pr,186845,issue,175477,high,pr.body,"eplayed and users take a second derivative through it, those detach nodes cut gradient flow and can produce silently wrong gradients, as in #175477. Add the existing aten.detach.default -> nop_decomposition mapping to _MakefxTracer's default decomposition table for non-pre-dis...",https://github.com/pytorch/pytorch/pull/186845,c58e7d1a6c81ebfbf5bc1b454d457a86a9ec4403ddde6595f37281537b7106eb closes,pr,184838,issue,175264,high,pr.body,"ferent, incomplete path. Centralizing the type-to-VT decision in the builders fixes both sourceful and sourceless descriptor objects. Fixes #175264 Generated by my agent Test Plan: python test/dynamo/test_functions.py FunctionTests.test_class_dict_python_descriptor_get Functio...",https://github.com/pytorch/pytorch/pull/184838,4bf4242ac31011287010a61e05b245954f31c7b60ece5f113d015b7199d649e0 closes,pr,187104,issue,175370,high,pr.closingIssuesReferences,pr #187104 declares a closing reference to issue #175370.,https://github.com/pytorch/pytorch/pull/187104,183ba294062c78ab2d94e9b9978129c647ff6b6b7875274b82bfae652d70cbe4 closes,pr,187104,issue,175370,high,pr.body,"Fixes #175370 F.embedding_bag on CPU segfaults when offsets is empty and indices is not. Root cause With zero bags, make_offset2bag fills offset2bag with",https://github.com/pytorch/pytorch/pull/187104,a94eaaa6e94070ee4f0753fbb7017d1fc18b9d7dcf6ac139daf397df86c74bd9 closes,pr,184848,issue,175154,high,pr.body,"or/clamp, which does not fix the same-size shortcut case. This keeps the native shortcuts scoped to the paths that actually use them. Fixes #175154 Generated by my agent Test Plan: python test/dynamo/test_repros.py -k interpolate_nearest python heredoc exact issue repro with b...",https://github.com/pytorch/pytorch/pull/184848,7787c358acce044e91eefcb8f990b027b04163fae05c29e7fa424c0135d36723 competes with,pr,189181,pr,188042,medium,pr.body,"t-advisor.yml) and cleanly separates the three ""no-revert, not-the-suspect"" verdicts. This is a minimal alternative to the prompt change in #188042 — no new principle paragraph and no enumerated error-string signatures; it mostly moves infra out of garbage and tightens the def...",https://github.com/pytorch/pytorch/pull/189181,0e15dd9f39824965a54b2da3b18508356096dfa937e9f2b70313fbea4da7b546 references,pr,189181,pr,188042,medium,pr.reviews[1].body,"LGTM! After I re-read and update #188042, the 2 PRs seem to follow the same line of thought now IMO. The remaining different is that #188042 adds a new reasoning principal for infr",https://github.com/pytorch/pytorch/pull/189181,80471492eba0ebe36a15983f4b3212536b26f03b8aa414734b5db0559943d586 review guidance,pr,189181,pr,8213,high,pr.reviews[1].body,"LGTM! After I re-read and update #188042, the 2 PRs seem to follow the same line of thought now IMO. The remaining different is that #188042 adds a new reasoning principal for infra_issue that looks ok to me, bit it might just be optional if AI advisor can figure out that part itself. Also, do we...",https://github.com/pytorch/pytorch/pull/189181,b0be62afb4bae86590a13164a5ba56bfc8fd07feb57042e1a51f874dd1edda7c review guidance,pr,189181,pr,188042,high,pr.reviews[1].body,"LGTM! After I re-read and update #188042, the 2 PRs seem to follow the same line of thought now IMO. The remaining different is that #188042 adds a new reasoning principal for infra_issue that looks ok to me, bit it might just be optional if AI advisor can figure out that part itself. Also, do we...",https://github.com/pytorch/pytorch/pull/189181,44527b0929f441027a4a6dd419b1fd65d3b8d9ffae03a32b8830518ef0ed8bd6 closes,pr,185764,issue,149565,high,pr.body,os.getpid() under fullgraph compilation. This keeps the fix scoped to the emitted guidance instead of changing graph-break behavior. Fixes #149565 Generated by my agent Test Plan: python test/dynamo/test_decorators.py DecoratorTests.test_untraceable_builtin_recommends_nonstric...,https://github.com/pytorch/pytorch/pull/185764,6146c20eb2a7c3f5cadd3258a545e41149dd771f7e806679b2d6d9c813af6490 review guidance,pr,185764,issue,149565,high,pr.reviews[4].body,requesting changes for discussion,https://github.com/pytorch/pytorch/pull/185764#pullrequestreview-4499840137,d805c0809bc1cd364c9e3f050ec86633d9273f8b1f99a2aea714483bb007a13e review guidance,pr,185764,pr,185764,high,pr.reviews[4].body,requesting changes for discussion,https://github.com/pytorch/pytorch/pull/185764#pullrequestreview-4499840137,e1976e317d12e50a079154999e0afc7041611aee9d5dfeb8350edc75f1be184f closes,pr,183969,issue,174472,high,pr.body,"revious duplicate result. Add regression coverage for unrelated writes, input aliases, unknown mutations, and result mutation safety. Fixes #174472 Generated by my agent cc @mlazos",https://github.com/pytorch/pytorch/pull/183969,ea5f2f077a6b19e5bf9a06ec867ded64512a5a28919cb58bea7c2a65d3ff012c closes,pr,184864,issue,174371,high,pr.body,as gapped and empty views. The torchgen ViewMeta flag is also extended to as_strided_ so inplace-view metadata follows the same path. Fixes #174371 Generated by my agent Test Plan: PYTHONPATH=/data/users/jansel/pytorch-issue-fixer/pytorch ninja -C build torch_python PYTHONPATH...,https://github.com/pytorch/pytorch/pull/184864,0954c54570c0c7138000162a8076d8c621682a26eb462ddced95d76fe7c2b88a closes,pr,185108,issue,163189,high,pr.body,and auto-initializing None buffers would invent user state. Restoring the module state fixes the side effect at the export boundary. Fixes #163189 Generated by my agent Test Plan: python test/export/test_export.py -k none_initialized_buffer_mutation_no_fake_tensor_leak python...,https://github.com/pytorch/pytorch/pull/185108,329d0a5d23c304ec70c929feca785508567d46fa3fd32493de48081af8e10c74 closes,pr,184948,issue,172839,high,pr.body,or dependencies inside static pytree context and risk incorrect exports. The replay path keeps those values as graph outputs instead. Fixes #172839 Generated by my agent Test Plan: python test/export/test_experimental.py -k test_export_out_spec_tensor_closure_without_flat_outp...,https://github.com/pytorch/pytorch/pull/184948,f9cccc5d6beb945b820927480d6fc6cdc236c5fa6487ee49b6b826c0e989e437 closes,pr,183731,issue,180807,high,pr.body,ernels fail fast before downstream static Triton kernels can consume mismatched tensor dtypes and silently compute incorrect results. Fixes #180807 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/183731,416d8148aeb907e98fc4da85e445b5f36f12c594f1c18c12998302f8c4f43214 references,pr,188747,pr,187479,medium,pr.body,"Built on top of #187479 Summary Cleans up dead code in the override dispatch. PyBackend objects are created via c10::make_intrusive in wrap(), which bypasses pybin",https://github.com/pytorch/pytorch/pull/188747,d6c4a4e0878db776cbb027839f9857919279b7995e3b5712f62369f86384dbea closes,pr,184958,issue,172026,high,pr.body,"ly graph-breaks after observing an escaped wrapped output, which matches the existing Dynamo handling for other unsafe graph outputs. Fixes #172026 Generated by my agent Test Plan: python test/dynamo/test_higher_order_ops.py -k vjp_returning -k vjp_consumed_in_compiled_region_...",https://github.com/pytorch/pytorch/pull/184958,8d37e042a21ef6bb0fb4e4974b74c18e46adaa69805e945fbfb9a1ca49c12817 references,pr,184958,issue,172026,medium,pr.comments[4].body,"ng. 3. **`backend=""aot_eager""`** in fullgraph test — ensures the validation fires *before* AOTAutograd, matching the actual crash path from #172026. 4. **New multi-vjp test** (`test_vjp_returning_graph_breaks_at_offending_call_site`) — directly tests the scenario from the orig...",https://github.com/pytorch/pytorch/pull/184958,d8f2bb1196a6da2f7724a3a74d0a536bb0e7d4691f8bb796dc7b8eea001a57ed references,pr,184958,issue,172026,medium,pr.review_comments[2].body,"[cleanup] `backend=""eager""` exercises the new `_validate_no_vjp_wrapped_outputs` guard but doesn't reproduce the original crash. The bug in #172026 is an AOTAutograd `NotImplementedError: Cannot access storage of TensorWrapper` raised from `_aot_autograd/collect_metadata_analy...",https://github.com/pytorch/pytorch/pull/184958,b739c600317f3a1a445095d5f2c4544542fe0827d7a84bc451255c689f4a9776 review guidance,pr,184958,issue,172026,high,pr.reviews[6].body,Aaron and I discussed the desired behavior before and we aligned on that so I'm deferring this review to him,https://github.com/pytorch/pytorch/pull/184958#pullrequestreview-4403872479,6e53c00fc090061b0a99eb7c02042c8c13b3f297dc5ed0da96210774f555ca9d review guidance,pr,184958,pr,184958,high,pr.reviews[6].body,Aaron and I discussed the desired behavior before and we aligned on that so I'm deferring this review to him,https://github.com/pytorch/pytorch/pull/184958#pullrequestreview-4403872479,2d8700a1d2adb9488e6ff8f8c90cb4505c4632e4f1ff231a4aae1e89bd4eef6f references,pr,181357,issue,116396,medium,pr.body,=True Tests added alongside the existing test_operator_concat / test_operator_iconcat tests in test/dynamo/test_functions.py. Fixes part of #116396 cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv @jiayisunx @kadeng @c...,https://github.com/pytorch/pytorch/pull/181357,56da9ab4ab819c75b32230b21d3fd235e1c0781806ef063a7d730f64f5a01736 review guidance,pr,181357,issue,116396,high,pr.reviews[0].body,See test failures,https://github.com/pytorch/pytorch/pull/181357,23372ccfdbfd3745636c4ec5c3405e5dce764695e5c3e0c79dd0d1fcbadf2071 review guidance,pr,180689,pr,180017,high,pr.reviews[0].body,fix merge conflicts,https://github.com/pytorch/pytorch/pull/180689,834956dce90eb0c2ef9fabaec7e25ac2faeab9e8d08de955c1f8d850c85cf2b1 review guidance,pr,180689,pr,180018,high,pr.reviews[0].body,fix merge conflicts,https://github.com/pytorch/pytorch/pull/180689,14ac8ca163db307eee450da0b2460efb5810c8b8a65e900b23d5b3aff7b47992 review guidance,pr,180689,pr,180019,high,pr.reviews[0].body,fix merge conflicts,https://github.com/pytorch/pytorch/pull/180689,e447e6a131a1d156203e85d4634a4be522643a7391952a325afbdb1f901fa1e5 closes,pr,184961,issue,171978,high,pr.body,"nd would still allow wrong shapes for max(lengths) < padded_length, so this patch handles the packed-sequence metadata path directly. Fixes #171978 Generated by my agent Test Plan: python test/export/test_export.py -k test_export_packed_sequence python test/export/test_export....",https://github.com/pytorch/pytorch/pull/184961,72553dffc687ef2a94b5371b8428c4ed50bd79d83c035febdbad54faf47c98c3 closes,pr,184964,issue,171974,high,pr.body,ndexing mask semantics. Scoping conversion to slice bounds fixes the reported cache update without changing direct indexing behavior. Fixes #171974 Generated by my agent Test Plan: python test/export/test_export.py TestExport.test_export_stateful_cache_tensor_scalar_slice_assi...,https://github.com/pytorch/pytorch/pull/184964,63f6a6e053c4a09be9717e8c152387a572cf129f4f658eddfdacf2816c4a8342 closes,pr,184966,issue,171964,high,pr.body,l inputs as the original module. This avoids generating a new forward body with exec; only the inspectable signature needs to change. Fixes #171964 Generated by my agent Test Plan: python test/export/test_unflatten.py -k test_unflatten_reexport_dynamic_shapes python test/expor...,https://github.com/pytorch/pytorch/pull/184966,62553e3751f8c74902d174663fbe2ec42f4e7d0ed577c17e5b8029fdf8e4d7fb closes,pr,184967,issue,171897,high,pr.body,patch through CompositeImplicitAutograd. Restricting this to ops with runtime backend kernels targets the bypass that caused the bug. Fixes #171897 Generated by my agent Test Plan: python test/export/test_export.py -k run_decompositions_decomposes_cia_op_with python test/expor...,https://github.com/pytorch/pytorch/pull/184967,f53cb0e5fa48c535384a2778a216c52eba27b12db0a9c7d134bd868c0ddc4e98 closes,pr,184970,issue,171894,high,pr.body,I checked the abandoned #173517 attempt and kept the same root cause direction without adding a hard-coded lift_fresh_copy exception. Fixes #171894 Generated by my agent Test Plan: ninja -C build lib/libtorch_cpu.so python test/test_functionalization.py TestFunctionalization.t...,https://github.com/pytorch/pytorch/pull/184970,7b88618da15c576f367cbbc7e5c9587e7643248206dcf9c54a2527c9dcf20a11 review guidance,pr,184970,issue,171894,high,pr.reviews[0].body,cc @aorenste do you know who the best person is to review functionalization related changes? there are some here,https://github.com/pytorch/pytorch/pull/184970,49fa6697dd43acdd3be8811e6afc1ab55c83266639285aad3e3aade241b5848b review guidance,pr,184970,pr,173517,high,pr.reviews[0].body,cc @aorenste do you know who the best person is to review functionalization related changes? there are some here,https://github.com/pytorch/pytorch/pull/184970,4ab10fd876b3df46b6b66b934aecff7eabf97c3ec4061f249cba27ce618f19ed review guidance,pr,184970,pr,184970,high,pr.reviews[0].body,cc @aorenste do you know who the best person is to review functionalization related changes? there are some here,https://github.com/pytorch/pytorch/pull/184970,70d7168ef3124186eed1edadb397fe25495a0db08cfe52416c751002ff786905 closes,pr,185815,issue,146951,high,pr.body,"than attempting to treat all NumPy operators as regular Python binary operators, which would mis-model what the user dunder receives. Fixes #146951 Generated by my agent Test Plan: python test/dynamo/test_functions.py -v DefaultsTests.test_numpy_operator_user_dispatch_input_va...",https://github.com/pytorch/pytorch/pull/185815,21ee37582bc0621fd2ca82d90f8c7858f19a6b4b830231c4d671cc391d907648 closes,pr,184974,issue,170834,high,pr.body,"no user setup_context, keyword-only metadata, TensorList inputs, and a torch.compile fullgraph regression with a fake implementation. Fixes #170834 Generated by my agent Test Plan: python test/test_custom_ops.py -k test_register_autograd_with_torch_func_grad_compile python tes...",https://github.com/pytorch/pytorch/pull/184974,8e609f6e1d680a176005401a6c1d4c0ee69983b7dd76d033502b3763ef90867a review guidance,pr,184974,issue,170834,high,pr.reviews[0].body,"ping me again in ~5 days, this might conflict with some of the eager custom ops work",https://github.com/pytorch/pytorch/pull/184974,11cd34289cd4fe4f2cf2ecf59fb58aa2e37841a69cadc8b586cc687163417f1e review guidance,pr,184974,pr,184974,high,pr.reviews[0].body,"ping me again in ~5 days, this might conflict with some of the eager custom ops work",https://github.com/pytorch/pytorch/pull/184974,c2eb7e58b428f32defc6e2155aaa7dd84e8ef66110a67515f000781e069a34ec review guidance,pr,184974,issue,170834,high,pr.reviews[1].body,"ping me again in 7 days, this might conflict with some of the eager custom ops work",https://github.com/pytorch/pytorch/pull/184974,84b256ebedc06b02ae804d987b90994ba4a7d69f58d58aaefe971049bfbd87bb review guidance,pr,184974,pr,184974,high,pr.reviews[1].body,"ping me again in 7 days, this might conflict with some of the eager custom ops work",https://github.com/pytorch/pytorch/pull/184974,db34e342e8293e0d9c6d0040d78fdd6fd7a360b7e97e6ab483e2d2dec53b040c closes,pr,184975,issue,170781,high,pr.body,ers use overwrite semantics; optimizers or external parameter references should not be set up for these fake module conversion flows. Fixes #170781 Generated by my agent Test Plan: python test/test_fake_tensor.py FakeTensorOperatorInvariants.test_move_fake_module_under_fake Fa...,https://github.com/pytorch/pytorch/pull/184975,121b109d0d5c1d8547be6bab0b68acb32cda63a686868dcff811f3d71afebc7c closes,pr,184976,issue,153056,high,pr.body,the userland FakeTensorMode dtype conversion and the non-strict export Module.to(...) path that fakifies module state. Fixes #170770 Fixes #153056 Generated by my agent Test Plan: python - <<'PY' import torch import torch.nn as nn from torch._subclasses.fake_tensor import Fake...,https://github.com/pytorch/pytorch/pull/184976,5250def63b04253f9a8704267a5021a95276fc30ce89027196016874765ed0b2 closes,pr,184976,issue,170770,high,pr.body,akref failure: the userland FakeTensorMode dtype conversion and the non-strict export Module.to(...) path that fakifies module state. Fixes #170770 Fixes #153056 Generated by my agent Test Plan: python - <<'PY' import torch import torch.nn as nn from torch._subclasses.fake_ten...,https://github.com/pytorch/pytorch/pull/184976,cc00ddc5499b0626c00fd1133df3a3862445dc8e419421d2c225952f59f7cad8 closes,pr,184977,issue,170696,high,pr.body,"rate missing closure on internal wrappers, but that would only mask the symptom and still leave guards pointed at the wrong callable. Fixes #170696 Generated by my agent Test Plan: python - <<'PY' ... minimal nn.Module with @torch.compile(fullgraph=True, backend=""eager"") + @to...",https://github.com/pytorch/pytorch/pull/184977,b93e4159296022cb78df36b47d1ffe34d749b8df766adefa412c2df43c20c055 closes,pr,188641,issue,188355,high,pr.body,cause the unsupported behavior belongs to the operator's CUDA implementation regardless of the compile mode that tries to capture it. Fixes #188355 Generated by my agent Test Plan: Ran the issue repro before the fix and reproduced the max-autotune CUDA graph capture crash on i...,https://github.com/pytorch/pytorch/pull/188641,48cda560ad67b0e61c20db0b3a4376933dfd39495a88c856e25c0807c65236c7 closes,pr,184051,issue,161113,high,pr.body,"ck ops without eager_input_vals, while the default needs_exact_strides custom-op contract remains gated on explicit eager_input_vals. Fixes #161113 Generated by my agent Test Plan: python -m pytest test/inductor/test_needs_exact_strides.py -q lintrunner -r HEAD~ --skip PYREFLY...",https://github.com/pytorch/pytorch/pull/184051,4ebc9c30e7d7539d546b23ff17580499c6243b6d601a86f9b581a8d09c727bb8 review guidance,pr,184051,issue,161113,high,pr.reviews[5].body,"> Before this change, GraphLowering only applied exact-stride constraints for the lowering path when the FX node carried node.meta[""eager_input_vals""]. Dynamo/AOTAutograd adds that metadata for ops present during the original trace, but the vLLM selective_scan case inserts the custom op in a post...",https://github.com/pytorch/pytorch/pull/184051#pullrequestreview-4367190985,43ea2a1b77e6be3dc8848eaa490135f11b31efb9d4d0b8407d9836e251d3ed54 review guidance,pr,184051,pr,184051,high,pr.reviews[5].body,"> Before this change, GraphLowering only applied exact-stride constraints for the lowering path when the FX node carried node.meta[""eager_input_vals""]. Dynamo/AOTAutograd adds that metadata for ops present during the original trace, but the vLLM selective_scan case inserts the custom op in a post...",https://github.com/pytorch/pytorch/pull/184051#pullrequestreview-4367190985,91dac19bc60af316314170bac7b454b845f116bddb2bb0ffba682d85d5ef326c review guidance,pr,184051,issue,161113,high,pr.reviews[6].body,"are the new helper functions _normalize_args_kwargs, _layout_constraints_for_target, _fake_args_kwargs_for_layout_constraints being used more than once? If not I think I'd prefer them just inlined directly in the code, so that it makes it easier for me to review.",https://github.com/pytorch/pytorch/pull/184051#pullrequestreview-4404607733,97f84fc029108da2b5e468f3c2e45ab8f5102c52d666ad179e78bd60c266858c review guidance,pr,184051,pr,184051,high,pr.reviews[6].body,"are the new helper functions _normalize_args_kwargs, _layout_constraints_for_target, _fake_args_kwargs_for_layout_constraints being used more than once? If not I think I'd prefer them just inlined directly in the code, so that it makes it easier for me to review.",https://github.com/pytorch/pytorch/pull/184051#pullrequestreview-4404607733,cfd7631c088e0dab2938ec35ff8ddf455a00814b5a61cd622c6e0c2f5e888517 review guidance,pr,184051,issue,161113,high,pr.reviews[7].body,sorry I'm finding the refactor difficult to read. I'm not really sure what would make it better,https://github.com/pytorch/pytorch/pull/184051#pullrequestreview-4500404988,5a67e5108e55b75c14a11905ac178da76c92ff61418e20d54203b6c6fff236b7 review guidance,pr,184051,pr,184051,high,pr.reviews[7].body,sorry I'm finding the refactor difficult to read. I'm not really sure what would make it better,https://github.com/pytorch/pytorch/pull/184051#pullrequestreview-4500404988,d3d3692979b293c5750e300cbf45d7a0eb9cca5bb2cb2262dd8d9233ff3dae9f closes,pr,185953,issue,145529,high,pr.body,end graph-break compiles. This clears only the compile-time weakref caches at the point where execution is about to resume in Python. Fixes #145529 Generated by my agent Test Plan: python test/dynamo/test_repros.py ReproTests.test_swap_tensors_after_custom_backend_graph_break...,https://github.com/pytorch/pytorch/pull/185953,a05c7bd363c69eeac56a3d58fbef1eca8182fda9ae6724500886058c6ad8ae0b references,pr,185953,issue,145529,medium,pr.comments[2].body,swap_module_params_on_conversion_after_custom_backend_graph_break`: Tests the higher-level `Module.to()` path from the original bug report (#145529). Uses `finally` to restore the `swap_module_params_on_conversion` flag correctly. Both tests explicitly set `invalidate_compile_...,https://github.com/pytorch/pytorch/pull/185953,47ddabc954810c41a043bbcdb6a26728965c04d72d9336fa6109daad247a88b9 references,pr,185953,issue,145529,medium,pr.comments[7].body,test_swap_module_params_on_conversion_after_custom_backend_graph_break` — tests the higher-level `Module.to()` path from the original issue #145529 All three use `config.patch(invalidate_compile_context_weakrefs=None)` to test the default codepath explicitly. **No blocking iss...,https://github.com/pytorch/pytorch/pull/185953,dc0256d2f8580e5525f11196d5bd4a884b669438c23cfa19b7dc507d89dbde61 references,pr,185953,issue,145529,medium,pr.comments[14].body,"lds. No blocking issues. The fix is well-scoped and the test coverage (direct repro, free-threaded GC path, and the `Module.to()` path from #145529) maps cleanly onto the three behaviors being changed. • `gh/jansel/1274/head`",https://github.com/pytorch/pytorch/pull/185953,0d8904ee390b3dd94c32c1753bbae3fa4f8c41da69e5ddbdc923df4897ee6940 closes,pr,185288,issue,161111,high,pr.body,"ch shapes. Tracking only the symbolic length matches the data actually needed by the grid and avoids materializing host-side offsets. Fixes #161111 Generated by my agent Benchmark Results: Concrete list/range compile microbenchmark, 5 warmup iterations and 50 measured torch.co...",https://github.com/pytorch/pytorch/pull/185288,bb18d60568ff62f28a27ea53f52106bf5b2c91c1862b8c0a0a6611150a15ee9c closes,pr,184978,issue,170672,high,pr.body,"xport sees a bogus lifted fake tensor constant, it used to raise an internal diagnostic ending with ""Please file an issue on github."" Issue #170672 hits this path by reading parameters through .data before copying back into the parameter under no_grad. Direct no_grad parameter...",https://github.com/pytorch/pytorch/pull/184978,7c3f8c1eda513e858fb5857909a898b9c8bddeb3a130f5c22729c8567db4f315 closes,pr,184981,issue,170509,high,pr.body,e temporary Dynamo skip from test_geometric_kstest and add eager/Dynamo regressions for the underlying bool/numeric subtract pattern. Fixes #170509 Generated by my agent Test Plan: python - <<'PY' ... torch._numpy bool subtract probe ... PY pytest test/torch_np/test_ufuncs_bas...,https://github.com/pytorch/pytorch/pull/184981,7c2a3099d4f2faf903de5b8096e00f7b02670a712f7e71ef1e760f456a7b1600 closes,pr,184983,issue,170127,high,pr.body,whether a fused kernel is definitely legal. Treating unknown fused-kernel preconditions as not viable is the narrower framework fix. Fixes #170127 Generated by my agent Test Plan: PYTHONPATH=/data/users/jansel/pytorch-issue-fixer/pytorch ninja -C build Minimized CPU repro for...,https://github.com/pytorch/pytorch/pull/184983,9f8d94d4844f592cd8d52bdde80eba0466dfc59d4d023f6ee3822678647710a9 closes,pr,184987,issue,169779,high,pr.body,the existing refs decomposition fixes all backends that consume the decomposition and avoids duplicating index_select lowering logic. Fixes #169779 Generated by my agent Test Plan: python test/test_decomp.py TestDecompCPU.test_comprehensive_index_select_cpu_complex32 -v python...,https://github.com/pytorch/pytorch/pull/184987,b1d3f3346078d5866a107dd9263a17a263750822be95cb642e2e321d388e2e1a references,pr,184987,issue,187463,medium,pr.reviews[0].body,See - #187463 . im pretty sure these are just eager bugs. not sure whether we should go ahead with this or just fix eager.,https://github.com/pytorch/pytorch/pull/184987,41cd6c1e97250184173c3b9cf12548e560b2eae461c20593d1ae711cb8dc361a review guidance,pr,184987,issue,169779,high,pr.reviews[0].body,See - #187463 . im pretty sure these are just eager bugs. not sure whether we should go ahead with this or just fix eager.,https://github.com/pytorch/pytorch/pull/184987,1a4712b33d4dcedcdb1fa0129a29b77058edb813e20e9558975bb3fb499418c4 review guidance,pr,184987,pr,184987,high,pr.reviews[0].body,See - #187463 . im pretty sure these are just eager bugs. not sure whether we should go ahead with this or just fix eager.,https://github.com/pytorch/pytorch/pull/184987,64e3c0bc096671c0bbb85e9ac13f4d0b8e34cfa85a7c9a678e13d49ab8994ef1 closes,pr,186027,issue,143495,high,pr.body,"out inference has failed, so valid symbolic broadcasts can keep the precise inferred size instead of forcing the slow path too early. Fixes #143495 Generated by my agent Test Plan: python -m py_compile torch/_refs/init.py torch/_prims/init.py torch/_subclasses/fake_impls.py py...",https://github.com/pytorch/pytorch/pull/186027,2a670f9fbf9f17b400ecc0bd0f696b520742d6bfb518fffb574fa067da926beb closes,pr,188636,issue,188422,high,pr.body,trace and verifies FixedLayout.is_stride_ordered can evaluate it without falling back to optimization hints or raising. Fixes #188423 Fixes #188422 Generated by my agent Test Plan: python test/inductor/test_unbacked_symints.py -k stride_ordered_handles_reciprocal_in_divisibili...,https://github.com/pytorch/pytorch/pull/188636,e72a9d0793e7c8bf39eba9922fa10de9a19194ee014283b4204a1f3fc8bd6cb3 closes,pr,188636,issue,188423,high,pr.body,e issue stack trace and verifies FixedLayout.is_stride_ordered can evaluate it without falling back to optimization hints or raising. Fixes #188423 Fixes #188422 Generated by my agent Test Plan: python test/inductor/test_unbacked_symints.py -k stride_ordered_handles_reciprocal...,https://github.com/pytorch/pytorch/pull/188636,eed4bb51f63c5c3ab516ce7065a60bb13d3a218c1da1113f8d904cf4b1a7890c closes,pr,186893,issue,160271,high,pr.body,s/fake_impls.py torch/_prims_common/init.py test/dynamo/test_recompiles.py test/dynamo/test_logging.py git diff --check lintrunner -a Fixes #160271 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/186893,a663fb3b9b234b572d7eebf655ab4a2c9f7e886e74f518311b2ca41b127f5b17 references,pr,186893,issue,160271,medium,pr.comments[2].body,"and nonstandard-stride cases each pin a distinct behavior. `test_dynamic_true_does_not_specialize_size_one` is the core regression test for #160271. Consider also asserting no recompile for a `size 1 -> 0` transition does the *right* thing (it should recompile, which `test_dyn...",https://github.com/pytorch/pytorch/pull/186893,6fbd1a1fdc857112c04d2e2b3ec5c661f13084c2e99523d8f7c7e87b97a00764 references,pr,186893,issue,160271,medium,pr.comments[5].body,"-branch, and nonstandard-stride preservation each pin a separate behavior, and `test_dynamic_true_does_not_specialize_size_one` is the core #160271 regression. The expected-inline updates in `test_logging.py` (`2 <= ... -> 1 <= ...`, plus the new `or`-form broadcast guard) and...",https://github.com/pytorch/pytorch/pull/186893,60a55c2477203e4b5e9a661b8e0b733ceb1ba664298e2d8a24264dbbeb3b9e63 references,pr,186893,issue,160271,medium,pr.comments[8].body,"exercises the `ExpandView` path end-to-end through Inductor with a nonstandard-strided size-1 input, which is the real-world failure mode (#160271-adjacent). Good that it's a numeric `allclose` check, not just a recompile count. - `test_dynamic_true_singleton_nonstandard_strid...",https://github.com/pytorch/pytorch/pull/186893,97917f051a83c4cf56a48b6952be61a99c752de34f18109fd3582e45e1d99153 references,pr,186893,issue,160271,medium,pr.comments[14].body,"und (I traced both-equal-hint, A-hint-1, and fallback cases). ### Tests Coverage is strong and each test pins a distinct behavior: the core #160271 regression (`does_not_specialize_size_one`), zero-branch preservation/exclusion, singleton broadcast generalization + the correct...",https://github.com/pytorch/pytorch/pull/186893,13e90ed160dc79ecffbf542fc619dc3fa89dea387bf7a0aa0cab24d20de2b19b closes,pr,184988,issue,169769,high,pr.body,al semantics. The narrower placeholder passthrough check keeps the existing safety boundary while accepting pure passthrough outputs. Fixes #169769 Generated by my agent Test Plan: python - <<'PY' original while_loop repro before/after fix python test/dynamo/test_higher_order_...,https://github.com/pytorch/pytorch/pull/184988,63d54eae667c8ae9d15b56b807c9882332638d8a8cadba2591bd471802b56088 closes,pr,184990,issue,157217,high,pr.body,e an invalid in-place shape) was surfaced as a compile error instead of being caught the way eager would. Fixes #169538 Fixes #183887 Fixes #157217 Generated by my agent Test Plan: python test/dynamo/test_exceptions.py ExceptionTests.test_fake_tensor_runtime_error_in_try_excep...,https://github.com/pytorch/pytorch/pull/184990,9e2f501ff0817c57d4174f38b3340cf69a584cbaefa57f6381e7e9ae14bec59d closes,pr,184990,issue,169538,high,pr.body,user try/except (for example an invalid in-place shape) was surfaced as a compile error instead of being caught the way eager would. Fixes #169538 Fixes #183887 Fixes #157217 Generated by my agent Test Plan: python test/dynamo/test_exceptions.py ExceptionTests.test_fake_tensor...,https://github.com/pytorch/pytorch/pull/184990,7fe84df20091e6e02be1a01bf97bad04f8128ba6ae7ca76902fd85118882345c closes,pr,184990,issue,183887,high,pr.body,pt (for example an invalid in-place shape) was surfaced as a compile error instead of being caught the way eager would. Fixes #169538 Fixes #183887 Fixes #157217 Generated by my agent Test Plan: python test/dynamo/test_exceptions.py ExceptionTests.test_fake_tensor_runtime_erro...,https://github.com/pytorch/pytorch/pull/184990,7514c1e8764f5b9548f3c5a6838d98aa52df031c214487f2395e52c1120c759a references,pr,184990,issue,169538,medium,pr.comments[2].body,scenarios: - **Happy path**: shape mismatch and custom op errors route through the user's `except RuntimeError` handler (the original issue #169538) - **Type preservation**: a `ValueError` raised by a fake impl is not caught by `except RuntimeError` (type matching is preserved...,https://github.com/pytorch/pytorch/pull/184990,efb78cd538c1f28c738930b624ce26a0aaa75eaaec477f1a67aa7ac6fcbf6cee closes,pr,184995,issue,169477,high,pr.body,s avoided because ops such as contiguous can allocate inference tensors and must continue to preserve inference-mode mutation errors. Fixes #169477 Generated by my agent Test Plan: python test/dynamo/test_aot_autograd.py AotAutogradFallbackTests.test_inference_mode_decorator_g...,https://github.com/pytorch/pytorch/pull/184995,af9125c22fcd9e7738c417f627ce0d0428bb4b881958f8331acebe42026c1b22 closes,pr,186032,issue,143463,high,pr.body,w metadata returns None when viewability or stride selection is undecidable rather than manufacturing strides from the example shape. Fixes #143463 Generated by my agent Benchmark Results: Not run. The change affects compile/export metadata fast paths rather than eager tensor...,https://github.com/pytorch/pytorch/pull/186032,4122e6e43f3300e53297713fa8071d61e24659a8a7a3c5fce253f16c40967fd4 closes,pr,185008,issue,168975,high,pr.body,omparison tools. Test Plan: python test/distributed/tensor/debug/test_debug_mode.py TestDebugModeUtils git diff --check lintrunner -a Fixes #168975 Generated by my agent,https://github.com/pytorch/pytorch/pull/185008,0c5f7d0b3cca0430ed36e7e45f4ac1d80684ca166e43be4ac568fc0d3b7136f2 closes,pr,185009,issue,168974,high,pr.body,opt-in flag instead of weakening the existing behavior so current callers that rely on strict validation keep the same failure mode. Fixes #168974 Generated by my agent Test Plan: python -m py_compile torch/utils/_debug_mode/_mode.py python test/distributed/tensor/debug/test_d...,https://github.com/pytorch/pytorch/pull/185009,9bda031304d133f0fef846f4e92c704a95f23fdc293065c598fddbe66f1ad662 closes,pr,185017,issue,168136,high,pr.body,added more complexity to attribute resolution. Moving field initialization earlier fixes the construction-order root cause directly. Fixes #168136 Generated by my agent Test Plan: python test/test_per_overload_api.py TestPerOverloadAPI.test_opoverloadpacket_init_under_sys_sett...,https://github.com/pytorch/pytorch/pull/185017,be14c5d658755f7879da38ce99e2b9115de81d22f6a1aa3020d682387b5e7a53 review guidance,pr,185017,issue,168136,high,pr.reviews[0].body,"we are not going to make torch._ops more complicated. I would prefer more work in Dynamo to support this thing correctly than complicating torch._ops, or making torch._ops less complicated so that it traces better.",https://github.com/pytorch/pytorch/pull/185017,962e385ed378bea96baa625762cc72501521bd0fff8275917990e326602863a2 review guidance,pr,185017,pr,185017,high,pr.reviews[0].body,"we are not going to make torch._ops more complicated. I would prefer more work in Dynamo to support this thing correctly than complicating torch._ops, or making torch._ops less complicated so that it traces better.",https://github.com/pytorch/pytorch/pull/185017,b735809a8fb3fc9bb38c2f8ddf9fb28b0019f71ba62c7b681389cbad0ac92e16 closes,pr,184628,issue,182217,high,pr.body,tedTensor keepdim reductions do not create unguardable ephemeral guards. Leave mixed or unsupported SingletonInt expressions unknown. Fixes #182217 Fixes #183369 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzhe...,https://github.com/pytorch/pytorch/pull/184628,458928505aae5f4f9f518d2398abbb6814934324386d60baccf9e234ccf6af7e closes,pr,184628,issue,183369,high,pr.body,dim reductions do not create unguardable ephemeral guards. Leave mixed or unsupported SingletonInt expressions unknown. Fixes #182217 Fixes #183369 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/184628,da78876e681c420e6e56b55c5234af07a1c87b01d2fdab0dbe66249aaceade00 review guidance,pr,184628,issue,182217,high,pr.reviews[0].body,Minor test coverage comment,https://github.com/pytorch/pytorch/pull/184628,cade43a7b3688aa3b97c6b72af8d71a462c667fd1745cebd0888808051cae9d5 review guidance,pr,184628,issue,183369,high,pr.reviews[0].body,Minor test coverage comment,https://github.com/pytorch/pytorch/pull/184628,c9deffd6fa589c7cc0a2ff2019eb78c16e8b09749ac0b5567698742c1f922fdc review guidance,pr,184628,pr,184628,high,pr.reviews[0].body,Minor test coverage comment,https://github.com/pytorch/pytorch/pull/184628,1a2909335f83f33fa192253568e5b42042c5e5fbc606c7e1bad4f247655cea1c closes,pr,184481,issue,93617,high,pr.body,c-base storage groups so compiled InputBuffer view reconstruction is not reused after input alias topology or storage offsets change. Fixes #93617 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/184481,dcedc4334dfc01a794b335f77651c9f5966051287c41ad66862aa1e804906953 references,pr,184481,issue,93617,medium,pr.comments[2].body,"ion (autograd_cache.py, runtime_wrappers.py) - [x] Review tests - [x] Post review feedback --- ### Summary This PR fixes a correctness bug (#93617) where AOTAutograd's synthetic-base view reconstruction could be incorrectly reused after input alias topology or storage offsets...",https://github.com/pytorch/pytorch/pull/184481,80ac55d38a82c068e5735dbd7071ea8d2501b54f0bc751c99abb821e3620abaf references,pr,184481,issue,93617,medium,pr.comments[4].body,"ion (autograd_cache.py, runtime_wrappers.py) - [x] Review tests - [x] Post review feedback --- ### Summary This PR fixes a correctness bug (#93617) where AOTAutograd's synthetic-base view reconstruction could be incorrectly reused after input alias topology or storage offsets...",https://github.com/pytorch/pytorch/pull/184481,af5590f5fb5e845da15ba2da5dce569622e7eb5b56978dc535266d041bbd84aa references,pr,184481,pr,185891,medium,pr.comments[6].body,-base case where a same-storage group already exists at trace time and keeps the topology guard grouped to avoid the O(P^2) shape that hurt #185891. The different-storage-inputs-that-later-overlap case still belongs in the other guard work; I rebased this PR and kept the confl...,https://github.com/pytorch/pytorch/pull/184481,3810965d5b94ff3b1bc7670814be481f53524b479b4d184c7baad54ae9c16953 supersedes,pr,184481,pr,184694,medium,pr.reviews[6].body,"(Reviewed by me, assisted by AI) [question] This overlaps with two other open PRs reworking AOTAutograd input-overlap guards: #184694 (same author, separate stack -- guards different-storage/non-overlap groups plus cache-key replay) and the third-party #185891 (replaces St",https://github.com/pytorch/pytorch/pull/184481,0b55415a38af81a7021fcfea8b9996262655ed714e2f683e21f666bbf4dc637b supersedes,pr,184481,pr,185891,medium,pr.reviews[6].body,"rlap guards: #184694 (same author, separate stack -- guards different-storage/non-overlap groups plus cache-key replay) and the third-party #185891 (replaces StorageOverlap with a new partition guard). All three drop the same `and symbolic` gate in `compute_overlapping_inputs`...",https://github.com/pytorch/pytorch/pull/184481,04351a991e9b42dca5122132f41ef2e9a58bbcd7bef5baea76b97021e5ecbd55 references,pr,184481,issue,188133,medium,pr.review_comments[15].body,My agent says I created #188133 to track the pre-existing Pallas CPU synthetic-base shape mismatch and updated the skip comment to reference it directly. The skip remains,https://github.com/pytorch/pytorch/pull/184481,ed8762fd6656cbb9cf313e53067dd48865fd42d61ddf6740edc58bf26dfe4d4a review guidance,pr,184481,issue,93617,high,pr.reviews[8].body,"(Reviewed by me, assisted by AI) [question] This overlaps with two other open PRs reworking AOTAutograd input-overlap guards: #184694 (same author, separate stack -- guards different-storage/non-overlap groups plus cache-key replay) and the third-party #185891 (replaces StorageOverlap with a new...",https://github.com/pytorch/pytorch/pull/184481#pullrequestreview-4555420043,913f059a97f17a8419204de1e84ecfb90373c04ee37adbe402484b89c3b9c5bf review guidance,pr,184481,pr,184481,high,pr.reviews[8].body,"(Reviewed by me, assisted by AI) [question] This overlaps with two other open PRs reworking AOTAutograd input-overlap guards: #184694 (same author, separate stack -- guards different-storage/non-overlap groups plus cache-key replay) and the third-party #185891 (replaces StorageOverlap with a new...",https://github.com/pytorch/pytorch/pull/184481#pullrequestreview-4555420043,4bd82210d7f507739056c966cf00b637b44fdc2f19be6de78ebce59af5e65bf8 closes,pr,189122,issue,188900,high,pr.closingIssuesReferences,pr #189122 declares a closing reference to issue #188900.,https://github.com/pytorch/pytorch/pull/189122,c9aff83c72c14370a8a4d16301c44c5a1547ca2e88d745a20a33493e4f77ec12 closes,pr,189122,issue,188900,high,pr.body,Fixes #188900 Summary Multiplying a dense tensor by a sparse COO tensor whose size-1 sparse dimension must broadcast up to a larger size silently dropped,https://github.com/pytorch/pytorch/pull/189122,1332268b0a3914ec83083e2a8ce59de96cd94e9353ad12b329acd8dec479d70b review guidance,pr,189122,issue,188900,high,pr.reviews[0].body,Looks correct,https://github.com/pytorch/pytorch/pull/189122,3862839751a8d1eab6664ad49bedc4db2055b0bad337f7e8393d50cc16b9f9ba review guidance,pr,184079,pr,156103,high,pr.reviews[0].body,probably needs more thorough testing/benchmarking before enable,https://github.com/pytorch/pytorch/pull/184079,d46a73d148488a4734f963b37f54b56356dafd1920854ac7c1e2f0c6623a4c63 review guidance,pr,184079,pr,184079,high,pr.reviews[0].body,probably needs more thorough testing/benchmarking before enable,https://github.com/pytorch/pytorch/pull/184079,f740bece45b241ec247c514934bfc81b48aa6467b21e621c7c3485e82db8ab81 review guidance,pr,184079,pr,156103,high,pr.reviews[1].body,"will defer to @anijain2305 . this has a lot of interactions and consequence, nto sure we shoudl rush into it.",https://github.com/pytorch/pytorch/pull/184079,fea081377b850924cf3d4186b1ab3aad0bffb99604a7c663d7215a9fcdeaacf4 review guidance,pr,184079,pr,184079,high,pr.reviews[1].body,"will defer to @anijain2305 . this has a lot of interactions and consequence, nto sure we shoudl rush into it.",https://github.com/pytorch/pytorch/pull/184079,8f24c0a9de1bffecbcbe99e3d7cfc3acf8a35548d53bbca43628b357a2553b2b closes,pr,185025,issue,167872,high,pr.body,e before any artifact can call back into it is the narrower root-cause fix and matches the state ordering required by the graph path. Fixes #167872 Generated by my agent Test Plan: Reproduced the AttributeError with a compact serializer repro using a dynamic-shape ExportedProg...,https://github.com/pytorch/pytorch/pull/185025,921d25e51fb2d05ba43267dd1d40ad507e40d9973135e36ab201579c7c9157d2 closes,pr,185033,issue,167732,high,pr.body,ig patch scoped to the internal backward rewrite fixes the requested behavior without changing direct autograd.grad default behavior. Fixes #167732 Generated by my agent Test Plan: python test/dynamo/test_fwd_loss_bwd.py -k TestTensorBackwardDefaultConfig python test/dynamo/te...,https://github.com/pytorch/pytorch/pull/185033,2a599f783577237454390413a2a555929ca58a3e1806600f50493dc1afd5bb12 closes,pr,185034,issue,167729,high,pr.body,"sue. Existing compiled-autograd, GradientEdge, external grad_fn, consumed-grad_fn, and returned-output safety checks remain in place. Fixes #167729 Generated by my agent Test Plan: python -m py_compile torch/_dynamo/variables/torch.py test/dynamo/test_fwd_loss_bwd.py python te...",https://github.com/pytorch/pytorch/pull/185034,ccb4734c752a4480249e2cd921f3bdea02e24d5fc0425cc53d4b4b0f94f89cd1 references,pr,185034,issue,174929,medium,pr.comments[0].body,"possibly due to flakiness on trunk: inductor / inductor-cpu-test / test (cpu_inductor_torchbench, 1, 2, linux.2xlarge.amx, unstable) (gh) (#174929) RuntimeError: dets should have the same type as scores pull / linux-jammy-aarch64-py3.10 / test (default, 3, 5, linux.arm64.m8g.4...",https://github.com/pytorch/pytorch/pull/185034,af4c5371bf31d191536cec6e0d467da67f97e614c07f1ec7d0d6aefd8d58ad0d closes,pr,185035,issue,167719,high,pr.body,ixing formatting keeps the error consistent for other f-string/dict-key uses without making ModuleDict special-case symbolic strings. Fixes #167719 Generated by my agent Test Plan: python test/export/test_export.py -k test_data_dependent_fstring_moduledict_key python test/test...,https://github.com/pytorch/pytorch/pull/185035,cfc241dfb8b2cbf992d3e0cddceb340896cd4ad71e0072eaf38c84df0a56f0df review guidance,pr,185035,issue,167719,high,pr.reviews[0].body,"Nope. In non-strict, there's only so much you can do to avoid naughty situations. The repr choice for integers is one such thing where there is not a canonical answer. I value more giving a symbolic repr for unbacked symint than I value ""fidelity"", since full fidelity for non-strict is a fools ga...",https://github.com/pytorch/pytorch/pull/185035,ac399af7c01111e5c370dc239de7e2ba2d45ca9ffdc438002abbec1eb6831dd1 review guidance,pr,185035,pr,185035,high,pr.reviews[0].body,"Nope. In non-strict, there's only so much you can do to avoid naughty situations. The repr choice for integers is one such thing where there is not a canonical answer. I value more giving a symbolic repr for unbacked symint than I value ""fidelity"", since full fidelity for non-strict is a fools ga...",https://github.com/pytorch/pytorch/pull/185035,2452de8e5ebe8b73a58d31ed812bf32b8ea4f1979c35ea911e44a662c39d5dd5 review guidance,pr,185035,issue,167719,high,pr.reviews[1].body,"Nope. - In non-strict, there's only so much you can do to avoid naughty situations. The repr choice for integers is one such thing where there is not a canonical answer. I value more giving a symbolic repr for unbacked symint than I value ""fidelity"", since full fidelity for non-strict is a fools...",https://github.com/pytorch/pytorch/pull/185035#pullrequestreview-4415001128,3f6199abee321992d479c3cd3e7057746b175b8cfce0571b691f3c36de6ba839 review guidance,pr,185035,pr,185035,high,pr.reviews[1].body,"Nope. - In non-strict, there's only so much you can do to avoid naughty situations. The repr choice for integers is one such thing where there is not a canonical answer. I value more giving a symbolic repr for unbacked symint than I value ""fidelity"", since full fidelity for non-strict is a fools...",https://github.com/pytorch/pytorch/pull/185035#pullrequestreview-4415001128,76c143fa0044603130e87e9429bd0f4e697590a1b19a551fd5f1799098752a46 closes,pr,185042,issue,167596,high,pr.body,"reported module-name pattern because the root issue is residual bytecode value specialization, not nn.Module attributes specifically. Fixes #167596 Generated by my agent Test Plan: python inline issue repro with CompileCounter; after fix frame_count=1 and op_count=1 across mod...",https://github.com/pytorch/pytorch/pull/185042,2ebc7b901f5b1cf6584b810c671d269e1aeb991c23cf5aa7ad40fb0548f87ba0 references,pr,185042,issue,167596,medium,pr.comments[8].body,"produces plain `dict`s, this is fine. **6. Test coverage is excellent** The tests cover: - Basic dict key side-effect replay (the original #167596 issue pattern) - Tuple keys containing lazy constants - `dict.update` with lazy keys - Lazy constant as dict *value* (no unnecessa...",https://github.com/pytorch/pytorch/pull/185042,21934fc6ee5c20485478f6853c71824e36c195e20aea1e558db0fd113f48b9c0 closes,pr,186953,issue,186799,high,pr.body,"inary vectorized elementwise codegen intact while making the backward equality mask compare against the same forward-computed values. Fixes #186799 Generated by my agent Benchmark Results: Compiled CPU torch.atan2(x, y) on 1024x1024 float32, one thread, using torch.utils.bench...",https://github.com/pytorch/pytorch/pull/186953,15fbe3ef99535dad98a4669039b4a417ded895fbae6e0daa5a7b9d9a913802f1 closes,pr,184590,issue,183266,high,pr.body,r would create. This avoids specializing on data-dependent .item() results while preserving the returned NestedTensor metadata cache. Fixes #183266 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/184590,f24e51892b1f0bfbe8cb06ebc14c63aedb12fdb96ed2bcccb728526ad80151ad closes,pr,185043,issue,167568,high,pr.body,"ntains triton_kernel_wrapper_functional, no longer contains the mutating wrapper, and returns the functional wrapper's cloned output. Fixes #167568 Generated by my agent Test Plan: python test/higher_order_ops/test_local_map.py TestLocalMap.test_mutations_triton_deferred_inlin...",https://github.com/pytorch/pytorch/pull/185043,88cef3df5f0b9c8479f248967a5d5585f4f4c7fd771e6c871b648aab1a2d03e0 closes,pr,185050,issue,167191,high,pr.body,"mpile. A few Dynamo dynamic-shapes lazy module tests now pass with this fix, so their stale expected-failure annotations are removed. Fixes #167191 Fixes #173252 Generated by my agent Test Plan: python test/nn/test_lazy_modules.py TestLazyModules.test_lazy_linear_with_compile...",https://github.com/pytorch/pytorch/pull/185050,35664dc42b357d9654d84165dc0ccdc8dd6628a7eacff0b5c54ae549027ba332 closes,pr,185050,issue,173252,high,pr.body,"ynamo dynamic-shapes lazy module tests now pass with this fix, so their stale expected-failure annotations are removed. Fixes #167191 Fixes #173252 Generated by my agent Test Plan: python test/nn/test_lazy_modules.py TestLazyModules.test_lazy_linear_with_compile TestLazyModule...",https://github.com/pytorch/pytorch/pull/185050,32726d6e64637bae1ea11501a354ff142adaab3c455815d1db2ae3e28d466ef9 closes,pr,185051,issue,167117,high,pr.body,"ort-block formatting rather than changing type_repr so normal FX codegen behavior and existing annotation rendering remain unchanged. Fixes #167117 Generated by my agent Test Plan: Reproduced the original torch.save failure with a symbolic_trace function annotated as ""torch.Te...",https://github.com/pytorch/pytorch/pull/185051,79def2b598e5021dd11a3a1da300363386bd41ca9753588ac7c8c87098c192da closes,pr,185052,issue,167098,high,pr.body,"icated Blackwell fp4 conversion intrinsic could improve this further, but it does not address the kernel split root cause fixed here. Fixes #167098 Generated by my agent Test Plan: python test/inductor/test_quantization.py TestQuantization.test_bfloat16_to_float4_pack_fuses py...",https://github.com/pytorch/pytorch/pull/185052,d4e19d1ae8caf93d27ffad2b5a97d6fe3497a7987857ded619535e641e455071 closes,pr,184920,issue,173578,high,pr.body,llables. Keeping a raw-id fallback for non-weakrefable objects was also rejected because it preserves the same stale-id failure mode. Fixes #173578 Generated by my agent Test Plan: python -m pytest test/dynamo/test_trace_rules.py -q python -m pytest test/dynamo/test_decorators...,https://github.com/pytorch/pytorch/pull/184920,80bdba4f15da1bcb0c30793f333fc86ba37d47afdce654d7f3b6d8c3ea96abc5 closes,pr,185775,issue,148695,high,pr.body,check git diff --cached --check lintrunner -a (attempted; failed on unrelated existing XPU/mkldnn clang-tidy/generated-header issues) Fixes #148695 Generated by my agent,https://github.com/pytorch/pytorch/pull/185775,470590ba0ee3f46d11a3a94b64419fe48184930d7f36f9a85f6afd3df87993ab references,pr,185775,issue,148695,medium,pr.comments[2].body,"te a genuine `OpaqueTensorImpl` tensor with `PrivateUse1` dispatch key, which exactly matches the real-world scenario described in #148695. It validates both `MetaConverter` and `FakeTensorMode.from_tensor`, covering the full fakification path. #### Verdict LGTM. The...",https://github.com/pytorch/pytorch/pull/185775,f8f4f15e10f97bb12e5cc1243939ce7da04754c7ac7d3d4109f17059ffc36c18 closes,pr,186879,issue,164313,high,pr.body,"was closed unmerged. This patch keeps the same remapping direction and adds guard-source handling plus nested/collision regressions. Fixes #164313 Generated by my agent Benchmark Results: Issue-shaped export benchmark, CPU, strict=False, 3 warmups and 20 timed torch.export.exp...",https://github.com/pytorch/pytorch/pull/186879,559316d726aa1227b8d38e7e4845fb1d56266192133a2485f9eb496b264fa40e closes,pr,184830,issue,175267,high,pr.body,"riptor lookups can synthesize fresh objects, so the narrower guarded-source approach avoids immediate guard failures and stale folds. Fixes #175267 Generated by my agent Test Plan: python test/dynamo/test_repros.py ReproTestsDeviceCUDA.test_is_with_mutated_tensor_alias_source_...",https://github.com/pytorch/pytorch/pull/184830,f1b9016f5007f4555a53060883679a233cd3b1fd7a0f32fb538a28a21d4ba59d closes,pr,186886,issue,164124,high,pr.body,s intended keyed autotune semantics. Keeping a per-key runtime winner matches Triton's behavior while preserving dynamic compilation. Fixes #164124 Generated by my agent Test Plan: python -m py_compile torch/_inductor/runtime/triton_heuristics.py torch/_inductor/ir.py torch/_i...,https://github.com/pytorch/pytorch/pull/186886,139de9de9174e8281ab496d0485817a7b339763f8fc13aafbd33a00094d24a7b closes,pr,185100,issue,163713,high,pr.body,g reconstruction out of proxy tracing fixes the root cause without globally disabling the fake tensor cache or reducing cacheability. Fixes #163713 Generated by my agent Test Plan: python inline minimal cache-output synthesis repro python test/test_fake_tensor.py -k test_cache...,https://github.com/pytorch/pytorch/pull/185100,d12d10c9b6f5b011b42828e6fc9e42a8bd11a1d9f055a553bcf30a2135814eea closes,pr,185103,issue,163641,high,pr.body,"ked symbols, but that would reintroduce the binding time-travel hazard that the existing replacement restrictions are meant to avoid. Fixes #163641 Generated by my agent Test Plan: python test/test_dynamic_shapes.py -k 'test_div_unbacked_eq_item' python test/test_dynamic_shape...",https://github.com/pytorch/pytorch/pull/185103,a911250e84f16f2dd5cb9c6d7d660e7b93b0ad84fee2e545de9ec69e06012945 closes,pr,185105,issue,163323,high,pr.body,_device_pass is the supported graph-casting path. Keeping the fix in that pass avoids changing general FX GraphModule.to() semantics. Fixes #163323 Generated by my agent Test Plan: python test/export/test_passes.py TestPasses.test_move_to_device_pass_recompiles_graph_module Te...,https://github.com/pytorch/pytorch/pull/185105,e66d8c04374866b9140c43815c43c82b49f3b4f7800b3cac1cb639eca5149f85 closes,pr,185106,issue,163303,high,pr.body,into user-visible containers. Matching registered state during constant lifting keeps the fix localized to export graph construction. Fixes #163303 Generated by my agent Test Plan: python test/export/test_lift_unlift.py python test/export/test_export.py TestExport.test_export_...,https://github.com/pytorch/pytorch/pull/185106,0843c8592e54efdf5fc5fd909e78604acc5c15a070b67aa296ccbfb18c81fd74 closes,pr,185116,issue,163143,high,pr.body,Stack from ghstack (oldest at bottom): -> #185116 The failure in #163143 came from functionalization discovering an effectful op inside a torch.cond branch after the cond operands had already been fixed. That lef,https://github.com/pytorch/pytorch/pull/185116,9fc27c6250e6d562fe8dee8fdaa719795b5ec619aa0e912c910abd4bfa9b9abd closes,pr,185116,issue,165981,high,pr.body,"rately via the Inductor lowering path; threading the effect token through cond at the source fixes both manifestations. Fixes #163143 Fixes #165981 Generated by my agent Test Plan: python wrapper that predefines the missing local torchvision::nms schema, then runs test/higher_...",https://github.com/pytorch/pytorch/pull/185116,9576e4b0ac92f8228ef6b9619d6e7b7875899e2cf91438362edf0e8ba64920af references,pr,179550,pr,176688,medium,pr.body,e enabled with instantiate_device_type_tests and onlyAccelerator. Stack from ghstack (oldest at bottom): -> #179550 #178849 #179549 #176689 #176688 #178565,https://github.com/pytorch/pytorch/pull/179550,4ebd52dd16cb2d04261b3fe2ab9abb8d3239f8439b7cd6e2bc3b5499acfc5854 references,pr,179550,pr,176689,medium,pr.body,they are enabled with instantiate_device_type_tests and onlyAccelerator. Stack from ghstack (oldest at bottom): -> #179550 #178849 #179549 #176689 #176688 #178565,https://github.com/pytorch/pytorch/pull/179550,73cbb97bb0fc2fe42a9f4f1730c635eaae06818980b05fdad92b50f466c2f766 references,pr,179550,pr,178849,medium,pr.body,"n test_torch.py, they are enabled with instantiate_device_type_tests and onlyAccelerator. Stack from ghstack (oldest at bottom): -> #179550 #178849 #179549 #176689 #176688 #178565",https://github.com/pytorch/pytorch/pull/179550,e9d8685392c59ab3cd968fd29d8086680bb1c6ccbad0165bf87692379cf1404e references,pr,179550,pr,179549,medium,pr.body,"orch.py, they are enabled with instantiate_device_type_tests and onlyAccelerator. Stack from ghstack (oldest at bottom): -> #179550 #178849 #179549 #176689 #176688 #178565",https://github.com/pytorch/pytorch/pull/179550,ea43f256108072053bf3f965ca2373856c5806135913b4dbc84e0a4230e70ffb references,pr,179550,issue,188987,medium,pr.comments[0].body,"but was likely due to flakiness present on trunk: xpu / linux-noble-xpu-n-py3.10 / test (default, 12, 12, linux.idc.xpu) (gh) (disabled by #188987) test/dynamo/test_logging.py::LoggingTests::test_optimizer_non_static_param This comment was automatically generated by Dr. CI and...",https://github.com/pytorch/pytorch/pull/179550,ff2366d795b325dae45d1adacca8d9df71b274b5f36676e18e70bbf17bd1110e closes,pr,184944,issue,172869,high,pr.body,inline C++ extension repro verified OpaqueBase + OpaqueBaseMeta pybind class registers and constructs git diff --check lintrunner -a Fixes #172869 Generated by my agent,https://github.com/pytorch/pytorch/pull/184944,8d5522efdc2744fe9defffe9e463a42d714b4b66169adbd261bff025baa3f92e review guidance,pr,184944,issue,172869,high,pr.reviews[0].body,"The context of the original issue is that python classes that want to be opaque just need to subclass OpaqueBase, but for pybinded classes like ProcessGroup, we have to metaclass OpaqueBaseMeta. So the issue was tracking for a way to have the pybinded class just need to subclass OpaqueBase. But i...",https://github.com/pytorch/pytorch/pull/184944,b28f71d87ebd58b8c29a61f7ccc42c944616a2297c7902431ac43ff2c8c83611 review guidance,pr,184944,pr,184944,high,pr.reviews[0].body,"The context of the original issue is that python classes that want to be opaque just need to subclass OpaqueBase, but for pybinded classes like ProcessGroup, we have to metaclass OpaqueBaseMeta. So the issue was tracking for a way to have the pybinded class just need to subclass OpaqueBase. But i...",https://github.com/pytorch/pytorch/pull/184944,04c1a1611767539709ed9d654d5b6a2dd812da09cabcf92b11220754e9cbd87d closes,pr,184007,issue,167725,high,pr.body,ve stale cache directories in child processes. This also keeps Triton's libdevice path knob synchronized with the worker environment. Fixes #167725 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/184007,9cd44e79f02357334b40bd8a01ba83745a0e6fbe94a249f4a9c7fc637a27edab closes,pr,185121,issue,163040,high,pr.body,"r. This revives the direction from abandoned PR #163285 while preserving structural file-like support and closed-file error behavior. Fixes #163040 Generated by my agent Test Plan: python - <<'PY' import runpy import sys import types sys.modules[""torch._native""] = types.Module...",https://github.com/pytorch/pytorch/pull/185121,df4f158b76662db5488ca940e13827ab1c3ea9a92c82f251d21b33730dd37308 references,pr,188004,pr,157149,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 #188639 #188638 -> #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188004,9702b4435eff0c13a0515fa08604f7ef956ad47037cae98ddb9fcb4077bf75d6 references,pr,188004,pr,187690,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 #188639 #188638 -> #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188004,31e2d4095de2402c2ea8f0fbed8cf6eb05a052beb9f410f6a2d8a9b8edaac787 references,pr,188004,pr,187744,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 #188639 #188638 -> #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188004,696e8cc230cd2ad5008c6cb256d626d81b934443770f31ba70e8f604455b79ec references,pr,188004,pr,188638,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 #188639 #188638 -> #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188004,72eae962ac40d0a43f8a99b60ac563a12215c29e7a1e7232710521984e197c50 references,pr,188004,pr,188639,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 #188639 #188638 -> #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188004,303c59da9afeeb109ea392b8645e3d25a8d24eb750a9aeb7484dc4560daf739a references,pr,188004,pr,188824,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 #188639 #188638 -> #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188004,5a486ac9a2de2bdf26a677ffef00368fe6d1169574593462d8995995f70aebbf references,pr,188004,pr,188825,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 #188639 #188638 -> #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188004,9574ad451420480d90808c0fcb66f6db587ac84a5516780d955028071a4ff3e2 references,pr,188004,pr,188834,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 #188639 #188638 -> #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188004,f7d5992e356a2daa0b762fce8f206af2cb7192733a8c85f06ae3000b42ab1fa8 references,pr,188004,pr,189024,medium,pr.body,Stack from ghstack (oldest at bottom): #189024 #188825 #188834 #188824 #157149 #188639 #188638 -> #188004 #187744 #187690,https://github.com/pytorch/pytorch/pull/188004,369f83ecb535242d115265d165ef93cbfcd2ed27e96127499123b369622fa020 review guidance,pr,188004,pr,157149,high,pr.reviews[1].body,"My claude flagged this concern: Exception-stack imbalance in the pre-3.12 PEP 479 conversion (functions.py, gen_send_ex2) The new except ObservedUserStopIteration handler pushes the escaping StopIteration onto the shared exception stack and never pops it back off: prev = tracer.exn_vt_stack.get_r...",https://github.com/pytorch/pytorch/pull/188004,7908959c8ee124e041d8a5b321a97f40a020619499ab1f50786ded3026d94e73 review guidance,pr,188004,pr,187690,high,pr.reviews[1].body,"My claude flagged this concern: Exception-stack imbalance in the pre-3.12 PEP 479 conversion (functions.py, gen_send_ex2) The new except ObservedUserStopIteration handler pushes the escaping StopIteration onto the shared exception stack and never pops it back off: prev = tracer.exn_vt_stack.get_r...",https://github.com/pytorch/pytorch/pull/188004,8a7de150a18e041e9840b380d55226352b324bc2a45ea871447fd28779f11a3b review guidance,pr,188004,pr,187744,high,pr.reviews[1].body,"My claude flagged this concern: Exception-stack imbalance in the pre-3.12 PEP 479 conversion (functions.py, gen_send_ex2) The new except ObservedUserStopIteration handler pushes the escaping StopIteration onto the shared exception stack and never pops it back off: prev = tracer.exn_vt_stack.get_r...",https://github.com/pytorch/pytorch/pull/188004,598e1841110b0d263c31f00b7e4148a78e717ee43788c1279e78d9c051ddbb59 review guidance,pr,188004,pr,188004,high,pr.reviews[1].body,"My claude flagged this concern: Exception-stack imbalance in the pre-3.12 PEP 479 conversion (functions.py, gen_send_ex2) The new except ObservedUserStopIteration handler pushes the escaping StopIteration onto the shared exception stack and never pops it back off: prev = tracer.exn_vt_stack.get_r...",https://github.com/pytorch/pytorch/pull/188004,5c54d0c0005ffbf6b40b33091d1c9de0bdc976d5c86d4a0c8f8622fbfd4d5e43 review guidance,pr,188004,pr,188638,high,pr.reviews[1].body,"My claude flagged this concern: Exception-stack imbalance in the pre-3.12 PEP 479 conversion (functions.py, gen_send_ex2) The new except ObservedUserStopIteration handler pushes the escaping StopIteration onto the shared exception stack and never pops it back off: prev = tracer.exn_vt_stack.get_r...",https://github.com/pytorch/pytorch/pull/188004,a7f56aed5f3264ce9d2d04bb6785d44ea30bd76870a51f3c4c49ad8ef366c652 review guidance,pr,188004,pr,188639,high,pr.reviews[1].body,"My claude flagged this concern: Exception-stack imbalance in the pre-3.12 PEP 479 conversion (functions.py, gen_send_ex2) The new except ObservedUserStopIteration handler pushes the escaping StopIteration onto the shared exception stack and never pops it back off: prev = tracer.exn_vt_stack.get_r...",https://github.com/pytorch/pytorch/pull/188004,0d1e52a49b874385909fd3c088b3c66d22ae250c9390b854bd27baf2cec3c2b1 review guidance,pr,188004,pr,188824,high,pr.reviews[1].body,"My claude flagged this concern: Exception-stack imbalance in the pre-3.12 PEP 479 conversion (functions.py, gen_send_ex2) The new except ObservedUserStopIteration handler pushes the escaping StopIteration onto the shared exception stack and never pops it back off: prev = tracer.exn_vt_stack.get_r...",https://github.com/pytorch/pytorch/pull/188004,d60d0adade3b84e5f89510fcf47a39e8372c6c127c9744657af261448ed3bd07 review guidance,pr,188004,pr,188825,high,pr.reviews[1].body,"My claude flagged this concern: Exception-stack imbalance in the pre-3.12 PEP 479 conversion (functions.py, gen_send_ex2) The new except ObservedUserStopIteration handler pushes the escaping StopIteration onto the shared exception stack and never pops it back off: prev = tracer.exn_vt_stack.get_r...",https://github.com/pytorch/pytorch/pull/188004,c0f03872340e6a0af3856d96c358a0b9891cf7c1bd0c33a3d762ce87958984ad review guidance,pr,188004,pr,188834,high,pr.reviews[1].body,"My claude flagged this concern: Exception-stack imbalance in the pre-3.12 PEP 479 conversion (functions.py, gen_send_ex2) The new except ObservedUserStopIteration handler pushes the escaping StopIteration onto the shared exception stack and never pops it back off: prev = tracer.exn_vt_stack.get_r...",https://github.com/pytorch/pytorch/pull/188004,07c675410e27fada706082368e754b1a8f9d61c947e2ebf071ea1aa787694efc review guidance,pr,188004,pr,189024,high,pr.reviews[1].body,"My claude flagged this concern: Exception-stack imbalance in the pre-3.12 PEP 479 conversion (functions.py, gen_send_ex2) The new except ObservedUserStopIteration handler pushes the escaping StopIteration onto the shared exception stack and never pops it back off: prev = tracer.exn_vt_stack.get_r...",https://github.com/pytorch/pytorch/pull/188004,632aa41cf407900215ac373a02199668cbd1ff1477bf6c0fae811c3c6e478dfb closes,pr,185124,issue,157610,high,pr.body,"s run_decompositions during export — all of which previously failed with the same alias-annotation error. Fixes #163037 Fixes #162734 Fixes #157610 Generated by my agent Test Plan: python setup.py develop python - <<'PY' ... torch.compile(ReproBroadcastInDim(), backend=""induct...",https://github.com/pytorch/pytorch/pull/185124,328e72f26bdb92fc6e06fe9cdaa2628d350ba350494f91528dc39f8dbd8e6a0a closes,pr,185124,issue,162734,high,pr.body,"path as well as run_decompositions during export — all of which previously failed with the same alias-annotation error. Fixes #163037 Fixes #162734 Fixes #157610 Generated by my agent Test Plan: python setup.py develop python - <<'PY' ... torch.compile(ReproBroadcastInDim(), b...",https://github.com/pytorch/pytorch/pull/185124,aae9edb00a09f413ac3cd9f60f685a1499bfd9f0bcea63056a5c1f491dd62f7c closes,pr,185124,issue,163037,high,pr.body,ograd compile path as well as run_decompositions during export — all of which previously failed with the same alias-annotation error. Fixes #163037 Fixes #162734 Fixes #157610 Generated by my agent Test Plan: python setup.py develop python - <<'PY' ... torch.compile(ReproBroad...,https://github.com/pytorch/pytorch/pull/185124,8c017a43a35a8d154b10ad02bd72f117ac69d8b19581c343a24ea6401e9048d6 review guidance,pr,185124,issue,157610,high,pr.reviews[0].body,"sorry, don't want to deal with this right now. I don't think a user can actually run into this problem using torch.compile. For the torch.export path: i don't understand how this actually gets run into. If we get a better sense of that then I might be convinced that this is worth fixing",https://github.com/pytorch/pytorch/pull/185124,520c95f8bd1994575c27285a0e1c4bd1d000af9d2200e003cf4434d1efcfd2a4 review guidance,pr,185124,issue,162734,high,pr.reviews[0].body,"sorry, don't want to deal with this right now. I don't think a user can actually run into this problem using torch.compile. For the torch.export path: i don't understand how this actually gets run into. If we get a better sense of that then I might be convinced that this is worth fixing",https://github.com/pytorch/pytorch/pull/185124,279a87971f50ec0186725488c405509d212eb677ca2cf90d4576cb45efff312a review guidance,pr,185124,issue,163037,high,pr.reviews[0].body,"sorry, don't want to deal with this right now. I don't think a user can actually run into this problem using torch.compile. For the torch.export path: i don't understand how this actually gets run into. If we get a better sense of that then I might be convinced that this is worth fixing",https://github.com/pytorch/pytorch/pull/185124,1b94b2dcc00c887735de25d7bdda7a4950434f859c62138beca76ef1fbca0302 review guidance,pr,185124,pr,185124,high,pr.reviews[0].body,"sorry, don't want to deal with this right now. I don't think a user can actually run into this problem using torch.compile. For the torch.export path: i don't understand how this actually gets run into. If we get a better sense of that then I might be convinced that this is worth fixing",https://github.com/pytorch/pytorch/pull/185124,4a762b77cbf9304e72055db90bf760c1f3fc493190f3f82af58d566f96949f98 review guidance,pr,185124,issue,157610,high,pr.reviews[1].body,"My agent says I updated the PR to address the reachability concern: the regression test now starts from ordinary torch.broadcast_to(x, (2, 3)).clone(), runs export decompositions to produce prims.broadcast_in_dim, and verifies a follow-up run_decompositions({}) no longer fails functionalization....",https://github.com/pytorch/pytorch/pull/185124,14b0bb147e0e483e8078c5715d79e3b4249ca8b51a07beeebae6e3b82acf9150 review guidance,pr,185124,issue,162734,high,pr.reviews[1].body,"My agent says I updated the PR to address the reachability concern: the regression test now starts from ordinary torch.broadcast_to(x, (2, 3)).clone(), runs export decompositions to produce prims.broadcast_in_dim, and verifies a follow-up run_decompositions({}) no longer fails functionalization....",https://github.com/pytorch/pytorch/pull/185124,b7dbe9237bba0344a0d5236f8137a578d17adfa04cf2d441122f9894f2ae1884 review guidance,pr,185124,issue,163037,high,pr.reviews[1].body,"My agent says I updated the PR to address the reachability concern: the regression test now starts from ordinary torch.broadcast_to(x, (2, 3)).clone(), runs export decompositions to produce prims.broadcast_in_dim, and verifies a follow-up run_decompositions({}) no longer fails functionalization....",https://github.com/pytorch/pytorch/pull/185124,677b12f10b68e6e5e3484a6b0446ed2f38fee8b6dbf4ae89d065093ed96e5d49 review guidance,pr,185124,pr,185124,high,pr.reviews[1].body,"My agent says I updated the PR to address the reachability concern: the regression test now starts from ordinary torch.broadcast_to(x, (2, 3)).clone(), runs export decompositions to produce prims.broadcast_in_dim, and verifies a follow-up run_decompositions({}) no longer fails functionalization....",https://github.com/pytorch/pytorch/pull/185124,898cfc4e0fec96bb505b2158faa498cacd4c1caa02ae4dcb71d853e0567863c9 references,pr,179549,pr,176688,medium,pr.body,ype list in DecoratorInfo as xpu has same xfail behavior as cuda. Stack from ghstack (oldest at bottom): #179550 #178849 -> #179549 #176689 #176688 #178565,https://github.com/pytorch/pytorch/pull/179549,daa4df5c91706372e1433e745b6e52d7eeb6db94668ae45c978dcac11c2be2aa references,pr,179549,pr,176689,medium,pr.body,device_type list in DecoratorInfo as xpu has same xfail behavior as cuda. Stack from ghstack (oldest at bottom): #179550 #178849 -> #179549 #176689 #176688 #178565,https://github.com/pytorch/pytorch/pull/179549,21e2e2481e13f4dfa314fe611cab2925a42b3fda7ea2689cd6d47639e687c683 references,pr,179549,pr,178849,medium,pr.body,test_ops.py using device_type list in DecoratorInfo as xpu has same xfail behavior as cuda. Stack from ghstack (oldest at bottom): #179550 #178849 -> #179549 #176689 #176688 #178565,https://github.com/pytorch/pytorch/pull/179549,552c21ab13eb6cf9a0999d997bda8c158a92c57c6c0471f8431eb0089500a096 references,pr,179549,pr,179550,medium,pr.body,p_db for test_ops.py using device_type list in DecoratorInfo as xpu has same xfail behavior as cuda. Stack from ghstack (oldest at bottom): #179550 #178849 -> #179549 #176689 #176688 #178565,https://github.com/pytorch/pytorch/pull/179549,9bd147501fb55bb7ed03ed5e2ab7e3c0e834f5581e8f4582128fdd0eba1c356f closes,pr,182224,issue,178522,high,pr.body,Stack from ghstack (oldest at bottom): -> #182224 Fixes #178522,https://github.com/pytorch/pytorch/pull/182224,e874175db879484cbd895bd9e904e64c8bb51f455b6386c9ce4ff6ec1fde5b5f references,pr,176830,issue,116396,medium,pr.body,#116396 Add support for operator.setitem and operator.delitem by handling them in BuiltinVariable(same as the existing getitem). Add tests covering,https://github.com/pytorch/pytorch/pull/176830,c53f78c2abdbc9d3ebee4c2c42222174fcff3f356886e32095a1d7f97ad7187f review guidance,pr,176830,issue,116396,high,pr.reviews[0].body,"Instead of polyfills, let's use the builtins approach in torch/_dynamo/variables/builtin.py",https://github.com/pytorch/pytorch/pull/176830,06742c98f61bfbf835243616d81bacb259961f76312552c62ddfeb1bb653a4b6 closes,pr,185126,issue,162973,high,pr.body,"pped from about 3.34s cold / 2.91s warm to about 0.91s, and GB_REGISTRY all-files lintrunner dropped from about 3.50s to about 1.27s. Fixes #162973 Generated by my agent Test Plan: python3 -m py_compile tools/linter/adapters/gb_registry_linter.py tools/dynamo/gb_id_mapping.py...",https://github.com/pytorch/pytorch/pull/185126,5aa47a53d7027822f8bc79d0540ed5f444e8fb252e36f5aed2b8a2b342366895 references,pr,176689,pr,176688,medium,pr.body,low_xpu=True augument skip cases in op db if xpu has limitations. Stack from ghstack (oldest at bottom): #179550 #178849 #179549 -> #176689 #176688 #178565,https://github.com/pytorch/pytorch/pull/176689,51457d84ee0177ae21b6b6a8d5c367492b35426dc779c5f3dc01d62d85920db2 references,pr,176689,pr,178849,medium,pr.body,iate_device_type_tests() allow_xpu=True augument skip cases in op db if xpu has limitations. Stack from ghstack (oldest at bottom): #179550 #178849 #179549 -> #176689 #176688 #178565,https://github.com/pytorch/pytorch/pull/176689,214ebd35aa261f046d4eb249aece3fd255740c240a2280167b79eac1b9c13c97 references,pr,176689,pr,179549,medium,pr.body,ice_type_tests() allow_xpu=True augument skip cases in op db if xpu has limitations. Stack from ghstack (oldest at bottom): #179550 #178849 #179549 -> #176689 #176688 #178565,https://github.com/pytorch/pytorch/pull/176689,665bb37574fe0ccca5266b539f412a8bc7cda2411c33fff67f1f12e891ef3f0d references,pr,176689,pr,179550,medium,pr.body,instantiate_device_type_tests() allow_xpu=True augument skip cases in op db if xpu has limitations. Stack from ghstack (oldest at bottom): #179550 #178849 #179549 -> #176689 #176688 #178565,https://github.com/pytorch/pytorch/pull/176689,acd0a0a3ed6340bf520f98886cc8f61613fc134a9bf545faed8db4574259ecdf references,pr,176688,pr,176689,medium,pr.body,() of TestCommon for xpu Update the op_db to align XPU op dyptes with OpInfo Stack from ghstack (oldest at bottom): #179550 #178849 #179549 #176689 -> #176688 #178565 disable torch.half and torch.complex32 on XPU for some fft ops before they are implemented update dot and vdot...,https://github.com/pytorch/pytorch/pull/176688,65471be78faf83056d098071e40a1c3511240cbee88751a37dd5ace9c3784439 references,pr,176688,pr,178849,medium,pr.body,able test_dtypes() of TestCommon for xpu Update the op_db to align XPU op dyptes with OpInfo Stack from ghstack (oldest at bottom): #179550 #178849 #179549 #176689 -> #176688 #178565 disable torch.half and torch.complex32 on XPU for some fft ops before they are implemented upd...,https://github.com/pytorch/pytorch/pull/176688,89b8adafb547f8c195e30d3e46f48f51fc314b26f3ab66c2607ce0c7226db75f references,pr,176688,pr,179549,medium,pr.body,t_dtypes() of TestCommon for xpu Update the op_db to align XPU op dyptes with OpInfo Stack from ghstack (oldest at bottom): #179550 #178849 #179549 #176689 -> #176688 #178565 disable torch.half and torch.complex32 on XPU for some fft ops before they are implemented update dot...,https://github.com/pytorch/pytorch/pull/176688,fc460fa958ff93c09c777e8145e4588fe20a1d848266c9eab53418b070c845a2 references,pr,176688,pr,179550,medium,pr.body,Enable test_dtypes() of TestCommon for xpu Update the op_db to align XPU op dyptes with OpInfo Stack from ghstack (oldest at bottom): #179550 #178849 #179549 #176689 -> #176688 #178565 disable torch.half and torch.complex32 on XPU for some fft ops before they are implemented u...,https://github.com/pytorch/pytorch/pull/176688,296429523c08b06ffa4bfb13a3f9887d8d9da202070143a6b002f163d374fd32 closes,pr,185132,issue,162687,high,pr.body,ients preserves the dispatcher elision behavior and matches the existing acceptance of trailing None gradients for non-tensor inputs. Fixes #162687 Generated by my agent Test Plan: python - <<'PY' ... PY # issue-style torch.compile(fullgraph=True) repro with default-valued cus...,https://github.com/pytorch/pytorch/pull/185132,3a37065f8578868eee6dbc08546b30a3a3b978e2a4f1a2e8a3bbc2fb51982507 review guidance,pr,185132,issue,162687,high,pr.reviews[0].body,can you measure how much overhead this adds?,https://github.com/pytorch/pytorch/pull/185132,d5a2db89e53435806d9d595c4c9793516c1ebf1c72c3d20a04c543885d6c83d2 review guidance,pr,185132,pr,185132,high,pr.reviews[0].body,can you measure how much overhead this adds?,https://github.com/pytorch/pytorch/pull/185132,e2e777af02d90a4e5b480cecd80818bd99a6a99c8e6e2b7c2cdfc3e5833f7894 closes,pr,185138,issue,162415,high,pr.body,d-frame count at the wrapper boundary matches the existing zero-frame enforcement point and keeps the fullgraph contract centralized. Fixes #162415 Generated by my agent Test Plan: ninja -C build torch_python python test/dynamo/test_compile.py -v FullgraphTests python test/dyn...,https://github.com/pytorch/pytorch/pull/185138,32b9b26ba48ebef9390b875d14e8df7580c289fd679036692f6dd0780d926972 closes,pr,185140,issue,162386,high,pr.body,ython test/export/test_passes.py TestPasses.test_predispatch_set_grad TestPasses.test_predispatch_autocast_and_set_grad lintrunner -a Fixes #162386 Generated by my agent,https://github.com/pytorch/pytorch/pull/185140,e8d0955964492af4a9fad6e3d4952c958f21d1dd2ca929b510b63ccc8875a2fe closes,pr,185147,issue,162287,high,pr.body,he public symbolic helper. Keeping the mapping local to z3op matches the existing validator handling for other torch.sym_* functions. Fixes #162287 Generated by my agent Test Plan: python test/dynamo/test_exc.py ExcTests.test_z3op_sym_not -v python test/dynamo/test_misc.py Mis...,https://github.com/pytorch/pytorch/pull/185147,a81864f2f76b668301fec8a03b79f937637dcad529335efc6078555ac1bbef6d closes,pr,185149,issue,162151,high,pr.body,Stack from ghstack (oldest at bottom): -> #185149 Fixes #162151 The issue repro assigns x[0] = x[0].t() after a permute. That lowers to a copy where the source and destination are aliases with the same f,https://github.com/pytorch/pytorch/pull/185149,de7c6628b12899a2239f427ee9fa0eadd948dff3e7dc25561bb3c2386212cddc references,pr,185149,issue,174929,medium,pr.comments[0].body,"possibly due to flakiness on trunk: inductor / inductor-cpu-test / test (cpu_inductor_torchbench, 1, 2, linux.2xlarge.amx, unstable) (gh) (#174929) detectron2_maskrcnn_r_50_fpn inductor / inductor-cpu-test / test-osdc (cpu_inductor_torchbench, 1, 2, mt-l-x86iamx-8-64, unstable...",https://github.com/pytorch/pytorch/pull/185149,a14523eb67159f64114a87376c1157269702210cadfb2c658ec00ecdfaa43abd references,pr,163503,pr,171826,medium,pr.body,"Stack from ghstack (oldest at bottom): #171826 -> #163503 #185524 #188173 #186100 #186657 This will prevent recompilations (i.e. logs to TORCH_LOGS=""recompiles"") due to wrap_inline's gua",https://github.com/pytorch/pytorch/pull/163503,269e2eb6c9e07ea9de3baade74d35bd0666c84290eab87ca9ed6f58d6f4d0373 references,pr,163503,pr,185524,medium,pr.body,"Stack from ghstack (oldest at bottom): #171826 -> #163503 #185524 #188173 #186100 #186657 This will prevent recompilations (i.e. logs to TORCH_LOGS=""recompiles"") due to wrap_inline's guard on fn. If we cal",https://github.com/pytorch/pytorch/pull/163503,26e4d0fdbd2f240b09907bf5a33be8d66959d9a3e57521a3bad85802096757c7 references,pr,163503,pr,186100,medium,pr.body,"Stack from ghstack (oldest at bottom): #171826 -> #163503 #185524 #188173 #186100 #186657 This will prevent recompilations (i.e. logs to TORCH_LOGS=""recompiles"") due to wrap_inline's guard on fn. If we call wrap_inline on",https://github.com/pytorch/pytorch/pull/163503,123b20b0ef6da89a963cb120e48ee04664c98ebc0eaec440171dbff7b1cdc8cb references,pr,163503,pr,188173,medium,pr.body,"Stack from ghstack (oldest at bottom): #171826 -> #163503 #185524 #188173 #186100 #186657 This will prevent recompilations (i.e. logs to TORCH_LOGS=""recompiles"") due to wrap_inline's guard on fn. If we call wrap_i",https://github.com/pytorch/pytorch/pull/163503,7c69c68ca090f6fb4cb09ef2b79b3259c26c7388cd6eb7bb5833a24e9a203df3 references,pr,163503,issue,174929,medium,pr.comments[0].body,"possibly due to flakiness on trunk: inductor / inductor-cpu-test / test (cpu_inductor_torchbench, 1, 2, mt-l-x86iamx-8-64, unstable) (gh) (#174929) fastNLP_Bert inductor / inductor-cpu-test / test (cpu_inductor_torchbench, 2, 2, mt-l-x86iamx-8-64, unstable) (gh) (#174929) visi...",https://github.com/pytorch/pytorch/pull/163503,09ced63063d4ee439f9aa56b6ce8f2c1a1912c14f5abf3f0d5a04227c6ab2160 references,pr,186100,pr,163503,medium,pr.body,Stack from ghstack (oldest at bottom): #171826 #163503 #185524 #188173 -> #186100 #186657 cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-,https://github.com/pytorch/pytorch/pull/186100,ecdec5b5ef2f85b713dbeed680b968bbb6576b44cab2ff53342285721b2e7197 references,pr,186100,pr,171826,medium,pr.body,Stack from ghstack (oldest at bottom): #171826 #163503 #185524 #188173 -> #186100 #186657 cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng,https://github.com/pytorch/pytorch/pull/186100,cad7cf77b8908066228ceea2274b36b5965cdb168ae7e7706defb34b2ce0a3e3 references,pr,186100,pr,185524,medium,pr.body,Stack from ghstack (oldest at bottom): #171826 #163503 #185524 #188173 -> #186100 #186657 cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv @jia,https://github.com/pytorch/pytorch/pull/186100,7ecaa53000ff0dda164ddb4bd94757431d7b8509214b7bff48428f79bb593f94 references,pr,186100,pr,188173,medium,pr.body,Stack from ghstack (oldest at bottom): #171826 #163503 #185524 #188173 -> #186100 #186657 cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv @jiayisunx @,https://github.com/pytorch/pytorch/pull/186100,3e644fddea53d77fb89fd0db9dc5262706200b080ca73a54ef085c86d96e0e75 references,pr,186100,issue,174929,medium,pr.comments[0].body,"possibly due to flakiness on trunk: inductor / inductor-cpu-test / test (cpu_inductor_torchbench, 1, 2, mt-l-x86iamx-8-64, unstable) (gh) (#174929) fastNLP_Bert inductor / inductor-cpu-test / test (cpu_inductor_torchbench, 2, 2, mt-l-x86iamx-8-64, unstable) (gh) (#174929) visi...",https://github.com/pytorch/pytorch/pull/186100,b5271444074519037ffb77ab73ccfc4520e53bc91f4ff0d575383dba746a380d references,pr,185524,pr,163503,medium,pr.body,Stack from ghstack (oldest at bottom): #171826 #163503 -> #185524 #188173 #186100 #186657 Add NestedGraphBreaksStrong test variant that forces graph breaks at every leaf function return via debu,https://github.com/pytorch/pytorch/pull/185524,80fa2bcd4899dc35d756b8c17059f0b2989e48881914b324b6187c4a2838ff05 references,pr,185524,pr,171826,medium,pr.body,Stack from ghstack (oldest at bottom): #171826 #163503 -> #185524 #188173 #186100 #186657 Add NestedGraphBreaksStrong test variant that forces graph breaks at every leaf function return,https://github.com/pytorch/pytorch/pull/185524,c199117c9faf3e2641add32502a17bb2519d0b18ecd2244385cdac56fa441135 references,pr,185524,pr,186100,medium,pr.body,Stack from ghstack (oldest at bottom): #171826 #163503 -> #185524 #188173 #186100 #186657 Add NestedGraphBreaksStrong test variant that forces graph breaks at every leaf function return via debug_force_graph_break_on_leaf,https://github.com/pytorch/pytorch/pull/185524,416920a9186bccf214f98b343cc8d245bf354564b6efebf1d1d3084498a2428a references,pr,185524,pr,188173,medium,pr.body,Stack from ghstack (oldest at bottom): #171826 #163503 -> #185524 #188173 #186100 #186657 Add NestedGraphBreaksStrong test variant that forces graph breaks at every leaf function return via debug_force_graph_break,https://github.com/pytorch/pytorch/pull/185524,aba4a4758b227ba0a4097fdf4682b94c6626c8a6039ff8155faaa26e2f5d190a references,pr,185524,issue,174929,medium,pr.comments[0].body,"possibly due to flakiness on trunk: inductor / inductor-cpu-test / test (cpu_inductor_torchbench, 1, 2, mt-l-x86iamx-8-64, unstable) (gh) (#174929) fastNLP_Bert inductor / inductor-cpu-test / test (cpu_inductor_torchbench, 2, 2, mt-l-x86iamx-8-64, unstable) (gh) (#174929) visi...",https://github.com/pytorch/pytorch/pull/185524,f2be0e0f212c0aae9f9cb25fc706f0166740579d4d8342336f037a255bbbe8a8 closes,pr,185155,issue,161796,high,pr.body,"break message is surfaced directly, replace the generic backend-repo hint with guidance to allow eager fallback via fullgraph=False. Fixes #161796 Generated by my agent Test Plan: python test/dynamo/test_error_messages.py ErrorMessagesTest.test_backend_fake_tensor_exc ErrorMes...",https://github.com/pytorch/pytorch/pull/185155,29d9474187b464385a5f512a53b287bfba36087b834cb919d33a1a47f4854a67 closes,pr,185159,issue,161764,high,pr.body,ad behavior is the local rewrite choosing a GEMM form without a usable GEMM template; other max-autotune behavior can still be valid. Fixes #161764 Generated by my agent Test Plan: CUDA_VISIBLE_DEVICES=1 TORCHINDUCTOR_CACHE_DIR=/tmp/torchinductor_161764_manager_tests python te...,https://github.com/pytorch/pytorch/pull/185159,b0812891122df2451c5655fb34f95317346696b1ab93039a7ce1d6bcf34ce53b competes with,pr,185159,issue,161764,medium,pr.comments[2].body,"age in `test/inductor/test_max_autotune.py` - [x] Post review feedback --- This is a well-structured fix for a real performance regression (#161764). The approach is sound: gate the conv→GEMM rewrite on the availability of a usable non-ATen template, rather than letting it fal...",https://github.com/pytorch/pytorch/pull/185159,8dadcf5dcae6666b7ddee44945daecb4d9f1320f0c5c1f9aa270e9abfc0a8dfd references,pr,185159,issue,161764,medium,pr.comments[8].body,"coverage in `test/inductor/test_max_autotune.py` - [x] Post review feedback --- This is a well-targeted fix for the regression described in #161764. The core problem is clear: on small GPUs where GEMM templates aren't viable, the 1x1 conv→GEMM rewrite was still firing and repl...",https://github.com/pytorch/pytorch/pull/185159,97b3a664abb7631dcd7637e11af8f4e5a9d1514c9c50966088089b1e897494c0 review guidance,pr,181027,pr,131043,high,pr.reviews[0].body,This now enables int input with float opt_dtype. Is that expected?,https://github.com/pytorch/pytorch/pull/181027,826f07f58c42641b161caaa0fb97b86f73c98c4349a1a92939d47b03d3ce0914 review guidance,pr,181027,pr,172809,high,pr.reviews[0].body,This now enables int input with float opt_dtype. Is that expected?,https://github.com/pytorch/pytorch/pull/181027,74fa10d18fca6ad552000ccc9cd9444d01c50928814888489af78b69ba6f6168 closes,pr,185166,issue,161705,high,pr.body,ression when possible preserves existing symbolic relationships and only avoids the guards that encode eager slice clamping branches. Fixes #161705 Generated by my agent Test Plan: python test/export/test_export.py -k test_export_slice_static_bound_crosses_dynamic_dim python t...,https://github.com/pytorch/pytorch/pull/185166,091d75c4831d116daa2104a6650cd748c2045d902992d969c81348f8ff513ad6 closes,pr,181062,issue,146790,high,pr.closingIssuesReferences,pr #181062 declares a closing reference to issue #146790.,https://github.com/pytorch/pytorch/pull/181062,a42c0f8472789c013a8de90a54e77c1c1a9425bcb1f1204d34e12515ca0d59b6 closes,pr,181062,issue,146790,high,pr.body,Fixes #146790. Summary torch.ops.aten.multi_margin_loss_backward segfaults when reduction='none' and grad_output doesn't have shape [nframe]. Both CPU an,https://github.com/pytorch/pytorch/pull/181062,349b856837c33cfb37cce57b5f6f5f4c906834428417aee0807147ddedb4fafd closes,pr,181062,issue,146790,high,pr.comments[1].body,"kernel is still launched with reduce=false, and line 127 of the kernel dereferences gradOutput[0] on empty storage — the original segfault #146790 still reproduces. Either drop the target_.dim() > 0 gate, or gate on something that matches the kernel's actual indexing behavior...",https://github.com/pytorch/pytorch/pull/181062,c2d411416dd6f237f24fe78230116c42aa1f53286a8c04439786f559008f0e09 closes,pr,185170,issue,161671,high,pr.body,"izing the entire GraphModule as an opaque nested byte blob, which would hide its tensor storages from the outer torch.save machinery. Fixes #161671 Generated by my agent Test Plan: python test/export/test_serialize.py TestSaveLoad.test_torch_save_exported_program_with_call_mod...",https://github.com/pytorch/pytorch/pull/185170,66b75c3b8e02b24439fda6ed01c940be86585d054d177baa41126595e2dbc9a5 closes,pr,181036,issue,178089,high,pr.closingIssuesReferences,pr #181036 declares a closing reference to issue #178089.,https://github.com/pytorch/pytorch/pull/181036,24972d9d96a709bc3696b765afecbcfc5fa0eac163f761050c5f5cc98c56f95d closes,pr,181036,issue,178089,high,pr.body,"Fixes #178089. Summary torch.sparse.spdiags computes per-diagonal nnz as nnz_per_diag = at::where( offsets_1d.le(0), offsets_1d.add(shape[0]).clamp_max_(",https://github.com/pytorch/pytorch/pull/181036,51968fa05005b696f29f753bab2f4ddf5b9591d889a34bf9e52c6e21d7ce0134 review guidance,pr,181036,issue,178089,high,pr.reviews[0].body,Confirmed the behavior here is correct referenced against scipy's implementation. Thanks!,https://github.com/pytorch/pytorch/pull/181036,ff57746abe9ad1317197831a20fc4478ba0bc2a5a26035ca46d120b022c8f931 review guidance,pr,181036,issue,178089,high,pr.reviews[1].body,To unblock. Trusting @amjames .,https://github.com/pytorch/pytorch/pull/181036,9359c62137de230eba589cfe37d4a91974763faeed05862fff7f8aa5f9abd1cb closes,pr,185171,issue,161650,high,pr.body,on knobs. This belongs on docs/source/checkpoint.md because the setting controls compile-time activation rematerialization tradeoffs. Fixes #161650 Generated by my agent Test Plan: Confirmed torch._functorch.config.activation_memory_budget exists with default 1.0 and torch._dy...,https://github.com/pytorch/pytorch/pull/185171,0d05a66c5b23a2fb8b1e9538167c3691606d0120c18b658f1196da99a15e6ab6 review guidance,pr,185171,issue,161650,high,pr.reviews[1].body,"I don't know where the right file is, deferring to the agent to find this.",https://github.com/pytorch/pytorch/pull/185171,a82cf19fd225ecad67384f4fea66d289eed063db6d76498036be8c18d60f9289 review guidance,pr,185171,pr,185171,high,pr.reviews[1].body,"I don't know where the right file is, deferring to the agent to find this.",https://github.com/pytorch/pytorch/pull/185171,b292832163d1136a7089ca0b006c6bb5964b98298d9a0679a4e56f923c576809 closes,pr,188039,issue,188034,high,pr.closingIssuesReferences,pr #188039 declares a closing reference to issue #188034.,https://github.com/pytorch/pytorch/pull/188039,21d78a85bd54a3c65c6df716ff27430c28a88019ab9116764456f7297dc519f4 closes,pr,188039,issue,188034,high,pr.body,st that returns a sentinel object from setup_context and asserts sys.getrefcount is unchanged after repeated forward+backward passes. Fixes #188034 Test plan python test/test_autograd.py TestAutograd.test_custom_function_setup_context_no_refcount_leak Existing test_custom_func...,https://github.com/pytorch/pytorch/pull/188039,b8bac777b58593cf9b1b9cf88681f4f7ac91f8b88ff770c96f31de5d8a6098a8 references,pr,188173,pr,163503,medium,pr.body,"Stack from ghstack (oldest at bottom): #171826 #163503 #185524 -> #188173 #186100 #186657 When a graph break originates from a function on NGB_SUPPRESS_INLINELIST (e.g. torch.distributed), the p",https://github.com/pytorch/pytorch/pull/188173,aedccc4037a9713f578c2be02fc01e872c52e56008e332a1d7410bf152b5988f references,pr,188173,pr,171826,medium,pr.body,Stack from ghstack (oldest at bottom): #171826 #163503 #185524 -> #188173 #186100 #186657 When a graph break originates from a function on NGB_SUPPRESS_INLINELIST (e.g. torch.distributed,https://github.com/pytorch/pytorch/pull/188173,f4d4c29038242ac92b246545b8442c8b7fc7ee33edf5124a0f45b047d6374131 references,pr,188173,pr,185524,medium,pr.body,"Stack from ghstack (oldest at bottom): #171826 #163503 #185524 -> #188173 #186100 #186657 When a graph break originates from a function on NGB_SUPPRESS_INLINELIST (e.g. torch.distributed), the prior fix",https://github.com/pytorch/pytorch/pull/188173,787f9d078bfc917ac7a146a82e182cf539e696b052d8a171b55913af119a67e8 references,pr,188173,pr,186100,medium,pr.body,"Stack from ghstack (oldest at bottom): #171826 #163503 #185524 -> #188173 #186100 #186657 When a graph break originates from a function on NGB_SUPPRESS_INLINELIST (e.g. torch.distributed), the prior fix only suppressed NG",https://github.com/pytorch/pytorch/pull/188173,a46c0801ec4b3f80bf126c3fe4838e1227848853948dd2403418ade3013b6061 references,pr,188173,issue,174929,medium,pr.comments[0].body,"possibly due to flakiness on trunk: inductor / inductor-cpu-test / test (cpu_inductor_torchbench, 1, 2, mt-l-x86iamx-8-64, unstable) (gh) (#174929) fastNLP_Bert inductor / inductor-cpu-test / test (cpu_inductor_torchbench, 2, 2, mt-l-x86iamx-8-64, unstable) (gh) (#174929) visi...",https://github.com/pytorch/pytorch/pull/188173,eaa80717817602895c5ecb468f7793ef6d9346e67e2fa51ea5e38b336d1189a4 closes,pr,185191,issue,161456,high,pr.body,t every fake tensor leaf returned for a node. This keeps the safety check at the storage layer instead of special-casing any one HOP. Fixes #161456 Generated by my agent Test Plan: python test/dynamo/test_higher_order_ops.py -k test_hop_aliasing_checks_all_subclass_inner_tenso...,https://github.com/pytorch/pytorch/pull/185191,111b35a29c6da9d2cf65ff321af9cf5c10baf8335349c55c71d1b82951f10465 closes,pr,185274,issue,161289,high,pr.body,"improve the missing-credential error, but that preserves the private-credential requirement and does not fix the documented workflow. Fixes #161289 Generated by my agent Test Plan: python test/dynamo/test_ci_expected_accuracy.py env -u CH_KEY_ID -u CH_KEY_SECRET python - <<'PY...",https://github.com/pytorch/pytorch/pull/185274,48886e35bff91ef45c8d2f5ad8f47ee690fb6ae537c9bb5b6b2c78d6bfa77e09 review guidance,pr,185274,issue,161289,high,pr.reviews[0].body,Test plan insufficient. You have to actually inspect the source checkout after those commands,https://github.com/pytorch/pytorch/pull/185274,13cc53568af33f4b005b06b6222e616f701be4fd4ff451e3a3ddc2a0f3fb67bd review guidance,pr,185274,pr,185274,high,pr.reviews[0].body,Test plan insufficient. You have to actually inspect the source checkout after those commands,https://github.com/pytorch/pytorch/pull/185274,6d2cbb75c7f91dcdcbfe03c3570f96aff028ec7eb04a38663fcb9bfdcc97aa24 closes,pr,184047,issue,161138,high,pr.body,tic by default so large scalar reductions can avoid split-reduction follow-up kernels when cross-block synchronization is profitable. Fixes #161138 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/184047,62b315a8a0edf00d3499fd10e1c285e470b131a6de15d159321bf3044cd28ec7 review guidance,pr,184047,issue,161138,high,pr.reviews[0].body,Needs more global metrics/testing,https://github.com/pytorch/pytorch/pull/184047,98e6c784b92dfc0863fe688d272501d46ec0ca8c2c8b434b3f6fdf0556a5f8c6 review guidance,pr,184047,pr,184047,high,pr.reviews[0].body,Needs more global metrics/testing,https://github.com/pytorch/pytorch/pull/184047,5507b8b840ddd2366d5ca562276bd9aca14830211e4e9756f87a49d484a0380f review guidance,pr,184047,issue,161138,high,pr.reviews[1].body,Needs more perf validation.,https://github.com/pytorch/pytorch/pull/184047,a3b3ed78cefea48f1ab88f1e990fd5c3ea67929d99a7b5a9565f11e84ea2c8eb review guidance,pr,184047,pr,184047,high,pr.reviews[1].body,Needs more perf validation.,https://github.com/pytorch/pytorch/pull/184047,0f139737bdce80fbec19336978d5cc22d64bd47978d03ec05a2c83033e37763a closes,pr,148674,issue,113490,high,pr.body,"Fixes: #113490 When using the Microsoft Visual C++ Compiler with Intel® OpenMP, it's needed to avoid linking the Microsoft OpenMP runtime library (vcomp)",https://github.com/pytorch/pytorch/pull/148674,46638109c324a000b5fe9cab6f3a202191d410963381c6d1ca7716b5f2d170aa closes,pr,185293,issue,161088,high,pr.body,"he ONNX test that exercises Where with a model-input condition now uses a bool condition, since uint8 predicates are no longer valid. Fixes #161088 Generated by my agent Test Plan: cmake --build build --target install -- -j 64 python test/test_torch.py -k test_where_condition_...",https://github.com/pytorch/pytorch/pull/185293,f73301170cf597193be222e89f121a7a84760c2eae5d90cb209a5612add0edb6 closes,pr,186859,issue,181304,high,pr.body,"Stack from ghstack (oldest at bottom): -> #186859 Fixes #181304 Generated by my agent When torch.compile traces torch.func.grad over a model with an N-D linear input, AOTAutograd's joint tracing can run",https://github.com/pytorch/pytorch/pull/186859,763ff9b5e7d1451928212f8bd116e063df9378bd6edcd2670f66bf6dbb2e0f5e references,pr,184779,issue,77764,medium,pr.body,"e.perf_counter()-t0)*1000/5:.2f} ms/call"") # ~2.1 ms on M4 Max (was ~18,000 ms) Related issues #149325 — large-tensor MPS coverage tracker. #77764 — MPS op-coverage tracker. Authored by Jacob Kaplan. (Also — I'm a former Meta employee; my unixname there was jtkaplan!)",https://github.com/pytorch/pytorch/pull/184779,e1ff0c1d20345cc406f583bb45a417937ff311bffc1b15063fef5d07a93a6f0c references,pr,184779,issue,149325,medium,pr.body,"N + 1); torch.mps.synchronize() print(f""{(time.perf_counter()-t0)*1000/5:.2f} ms/call"") # ~2.1 ms on M4 Max (was ~18,000 ms) Related issues #149325 — large-tensor MPS coverage tracker. #77764 — MPS op-coverage tracker. Authored by Jacob Kaplan. (Also — I'm a former Meta employ...",https://github.com/pytorch/pytorch/pull/184779,7aca6d7a10bf380ddfd981e1bcb6bb387104e2e44451dfae38ef6498f7791b32 review guidance,pr,184779,issue,77764,high,pr.reviews[1].body,Please address the comments and benchmark for more numels/shapes/dtypes/layouts,https://github.com/pytorch/pytorch/pull/184779,6c557d64ca78bb85800bce73928197cd2ea721b058a353230c3661d6834cf068 review guidance,pr,184779,issue,149325,high,pr.reviews[1].body,Please address the comments and benchmark for more numels/shapes/dtypes/layouts,https://github.com/pytorch/pytorch/pull/184779,8b71263128c37a855901f9b1938a5b45a6e494c4010548b243cf5af64f2ce72d closes,pr,184780,issue,97310,high,pr.closingIssuesReferences,pr #184780 declares a closing reference to issue #97310.,https://github.com/pytorch/pytorch/pull/184780,4d7cc2bec5fab7a9c0e3f8c58e3f331ef3fee95be9d315673e20873965b570a2 references,pr,184780,issue,77764,medium,pr.body,"3). #111173 — closes (silent correctness bug from ScatterModeAdd congestion, open since October 2023). #149325 — large-tensor MPS coverage. #77764 — MPS op-coverage tracker. Authored by Jacob Kaplan. (Also — I'm a former Meta employee; my unixname there was jtkaplan!)",https://github.com/pytorch/pytorch/pull/184780,1ca9b6cdacf6b5f1d5bfc7030f11fd164ffaffbe45aa779d2f2d145cf9b2fde4 closes,pr,184780,issue,97310,high,pr.body,"Title: [MPS] Fast atomic-free path for flat torch.unique (~19 s → 1.7 ms, fixes #97310) TL;DR Closes #97310 (perf, open since March 2023) and #111173 (silent correctness bug, open since October 2023). On the workload-shape inp",https://github.com/pytorch/pytorch/pull/184780,bd21d3f2f6950ac16915a2be4e7612e21f0ff7f55b2f6244a4a1cd9536dbaa1e references,pr,184780,issue,111173,medium,pr.body,"tle: [MPS] Fast atomic-free path for flat torch.unique (~19 s → 1.7 ms, fixes #97310) TL;DR Closes #97310 (perf, open since March 2023) and #111173 (silent correctness bug, open since October 2023). On the workload-shape input (2M int64 elements, single long duplicate run), to...",https://github.com/pytorch/pytorch/pull/184780,fa64b5da76831087be9e159b3529cd92180921fee0ebeceddb33d578a107a4dd references,pr,184780,issue,149325,medium,pr.body,"— closes (perf, open since March 2023). #111173 — closes (silent correctness bug from ScatterModeAdd congestion, open since October 2023). #149325 — large-tensor MPS coverage. #77764 — MPS op-coverage tracker. Authored by Jacob Kaplan. (Also — I'm a former Meta employee; my un...",https://github.com/pytorch/pytorch/pull/184780,b6f87e8903e3cb1fdba2bbd4f0d686496e0c6a02c6dfb1ef4627270acbfa7701 closes,pr,184787,issue,176428,high,pr.body,p_serdes python test/export/test_schema.py TestSchema.test_schema_compatibility TestSchema.test_thrift_schema_unchanged lintrunner -a Fixes #176428 Generated by my agent,https://github.com/pytorch/pytorch/pull/184787,00dc4780e49d823570c56223f59dfaa6f74729fa6ac6d9ae734494f87e6ece9f closes,pr,185054,issue,184842,high,pr.closingIssuesReferences,pr #185054 declares a closing reference to issue #184842.,https://github.com/pytorch/pytorch/pull/185054,56ebd8e82a6b441a922e821bdd8687287792d38fc0ef8625da9b1715bbb841d2 closes,pr,185054,issue,184842,high,pr.body,"Fixes #184842 torch.autograd.functional.jacobian(strategy=""forward-mode"", vectorize=False) is an unsupported combination, but the API used to accept it a",https://github.com/pytorch/pytorch/pull/185054,21cdfc5348b4d769bf85331020b8acd5f97d096e03c142a29bc2d698415d6181 closes,pr,187161,issue,187155,high,pr.closingIssuesReferences,pr #187161 declares a closing reference to issue #187155.,https://github.com/pytorch/pytorch/pull/187161,9e3874ead2eeaf267f0004ae7d31f8e3b1efec67f81430c12376d9a26fcc2b9f closes,pr,187161,issue,187155,high,pr.body,"Summary Fixes #187155 When fully_shard is called with reshard_after_forward as an int (e.g., 2), each layer creates its own DeviceMesh and NCCL communicator in _",https://github.com/pytorch/pytorch/pull/187161,97db8566324f0bd7ce591cd640cbfed82d833e64ba2a89fe80102f1e46aa313b review guidance,pr,187161,issue,187155,high,pr.reviews[0].body,"the problem is valid - just implementation details needs to be polished: support fully_shard(list[Module]) because it's used in torchtitan as a way to support [lm_head, norm] add unit test, eg #179402",https://github.com/pytorch/pytorch/pull/187161,187941d874a94fe2f21198e668804aebd5fe1ca09b2695873791d2594705a2fd review guidance,pr,187161,issue,187155,high,pr.reviews[1].body,CI error looks real,https://github.com/pytorch/pytorch/pull/187161,54b329f0d2da39955cb0144f34c33d173b13bdae590c50f0f2b4ff5fedc3af1f references,pr,171826,pr,163503,medium,pr.body,Stack from ghstack (oldest at bottom): -> #171826 #163503 #185524 #188173 #186100 #186657 cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv,https://github.com/pytorch/pytorch/pull/171826,ca0d63db58325b729a95760fd8194c3f7d57f3725448716589834ae9c9223624 references,pr,171826,pr,185524,medium,pr.body,Stack from ghstack (oldest at bottom): -> #171826 #163503 #185524 #188173 #186100 #186657 cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv @jiayis,https://github.com/pytorch/pytorch/pull/171826,f6177e2bcef89193a7ac30daa3292de8bdc7f1fd133dabe9024dad9351b8029a references,pr,171826,pr,186100,medium,pr.body,Stack from ghstack (oldest at bottom): -> #171826 #163503 #185524 #188173 #186100 #186657 cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv @jiayisunx @kadeng @cha,https://github.com/pytorch/pytorch/pull/171826,c19e0e8eb6bf61c43c71bab8b360f86f4e0b26f5e07a2d04904a0181caf841c0 references,pr,171826,pr,188173,medium,pr.body,Stack from ghstack (oldest at bottom): -> #171826 #163503 #185524 #188173 #186100 #186657 cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv @jiayisunx @kad,https://github.com/pytorch/pytorch/pull/171826,d38e613d164183e720760cc2a593c9ddd93db029d27e4167bb2186f0d597bc89 references,pr,171826,issue,174929,medium,pr.comments[0].body,"possibly due to flakiness on trunk: inductor / inductor-cpu-test / test (cpu_inductor_torchbench, 1, 2, mt-l-x86iamx-8-64, unstable) (gh) (#174929) fastNLP_Bert inductor / inductor-cpu-test / test (cpu_inductor_torchbench, 2, 2, mt-l-x86iamx-8-64, unstable) (gh) (#174929) Runt...",https://github.com/pytorch/pytorch/pull/171826,a81a55f67ce60ddfb80241af6e25175bd87f7367bb015bb34df0f834743de596 references,pr,188979,pr,188978,medium,pr.body,"my changes to see original python impl, added 2 tests to test the register_fake mechanism Stack from ghstack (oldest at bottom): -> #188979 #188978",https://github.com/pytorch/pytorch/pull/188979,f1c6ffc7d36f93fbe91df8de7aaf4f8cb338cc61bf705c0764b97ecb2732999b review guidance,pr,188979,pr,188978,high,pr.reviews[1].body,can you add some tests to exercise this? without tests this is dead code,https://github.com/pytorch/pytorch/pull/188979,364dd59ec3387186daa58b31a1784799be53b9a4f05ff1e6efe64cf616dbc8d5 review guidance,pr,188979,pr,188979,high,pr.reviews[1].body,can you add some tests to exercise this? without tests this is dead code,https://github.com/pytorch/pytorch/pull/188979,0a1671c608f27e04f7a40560be2f2f7ce27301a2f02231e2d4d1e0c99183f0bf references,pr,188999,issue,180975,medium,pr.body,ths still compile # and run (verified natively on an aarch64 host): clang harness.c -o t && ./t # rc=0 Part of the RISC-V enablement effort #180975 (Phase 1.3: RISC-V CI / arch enablement). cc @jgong5 @mingfeima @XiaobingSuper @sanchitintel @ashokei @jingxu10 @aditew01,https://github.com/pytorch/pytorch/pull/188999,80f9c17158de4ba0a964d07041a44fed47c5b2fa54a9d6f22b133a33ca3551c4 references,pr,188999,issue,180975,medium,pr.comments[1].body,cc @fernchen — small RISC-V enablement fix from #180975. cpu_atomic_add_float's spin loop had no relax hint on RISC-V (it fell through to an empty _mm_pause() macro); this wires up the Zihintpaus,https://github.com/pytorch/pytorch/pull/188999,3bd0df602955f98efcabb38f6c2e98e87f61fee7b1559814699e1f94eb8c309a review guidance,pr,188042,pr,188042,high,pr.reviews[1].body,I don't like this change. it's overly specific and wordy,https://github.com/pytorch/pytorch/pull/188042,00500e99fa2c06cc44e0585dd6e0dd78f066603a9216de290a6941f1ad89beb0 closes,pr,185073,issue,165641,high,pr.body,ribing free state as original module state. Reclassifying the inputs at the cleanup boundary fixes the underlying signature mismatch. Fixes #165641 Generated by my agent Test Plan: python issue-shaped strict export script with torch._dynamo.config.install_free_tensors=True and...,https://github.com/pytorch/pytorch/pull/185073,4f19779f0a20c0d9480b421794d06511ffc8b968ecde501968b2085061883ccc closes,pr,185075,issue,165547,high,pr.body,isting maybe_get_fake_mode helper so actual FakeTensors and fake-containing wrapper subclasses are not misclassified as real tensors. Fixes #165547 Generated by my agent Test Plan: python test/test_fake_tensor.py -k test_deepcopy_real_tensor_in_fake_mode_error python test/test...,https://github.com/pytorch/pytorch/pull/185075,fe378219c7a00adfadb2b56b8afb070c7a35ad135895ad88afa2fa226e5ff1a6 review guidance,pr,185075,issue,165547,high,pr.reviews[0].body,"At minimum this needs to be split into two PRs: one for the error message, one for the real metadata fix. The diff seems overly long for just improving an error message. I would not want to commit to so much code for just an error message fix. It is possible that we should make deepcopy work on f...",https://github.com/pytorch/pytorch/pull/185075,38a3c0f2a81e8380d827dd37451211554350283b0cd090c46c6bb993cf481645 review guidance,pr,185075,pr,185075,high,pr.reviews[0].body,"At minimum this needs to be split into two PRs: one for the error message, one for the real metadata fix. The diff seems overly long for just improving an error message. I would not want to commit to so much code for just an error message fix. It is possible that we should make deepcopy work on f...",https://github.com/pytorch/pytorch/pull/185075,9eac3eb743b898706305e6db3a404ea6e12da3a23aaac796c795e018af890e46 review guidance,pr,182145,pr,180430,high,pr.reviews[0].body,Thank you for the fix!,https://github.com/pytorch/pytorch/pull/182145,bccb8c8cc9e9b7894d313b4ff089d2fd14f94e264f54c077096d0f0f5ee13aeb review guidance,pr,182145,pr,181955,high,pr.reviews[0].body,Thank you for the fix!,https://github.com/pytorch/pytorch/pull/182145,b06e83c2c9f25cb7d3d0e4cf4ddc66c7be7933a973ec6cae0c6610a564ca01ee references,pr,176678,issue,170753,medium,pr.body,ovided by Nvidia with RTX GPUs. Scope: This is purely an infrastructure addition; no core PyTorch source code is modified. Issue Reference: #170753 Release Notes: This is a CI infrastructure change and not user-facing. cc @peterjc123 @mszhanyi @skyline75489 @nbcsm @iremyux @Bl...,https://github.com/pytorch/pytorch/pull/176678,3d5196f7ae6314997fb7ef22dfbd98f149d38fb889b7c41c6eaea99fc5ecf08c review guidance,pr,176678,issue,170753,high,pr.reviews[0].body,"Why do you need separate _win-rtx-build.yml and _win-rtx-test.yml workflows? What is wrong with existing _win-build one? (Which is closer to binary pipeline workflows that still produces those binaries) I see this PR targets sm_89 architecture, but we already have plenty of tests for those on Lin...",https://github.com/pytorch/pytorch/pull/176678,9661ee08afbff5f9603d404ee82c18bd1a2356e7f501a312c89429dfd9163625 review guidance,pr,176678,issue,170753,high,pr.reviews[1].body,"On Workflow Separation (_win-rtx-build.yml vs. _win-build) Let me be very explicit: I'm fine with if/else path in existing _win-build.yml, but I don't want separate .yml file. If they are fundamentally different, explain why this is the case? To the best of my understanding, we can totally use ex...",https://github.com/pytorch/pytorch/pull/176678,cefd364e03cf638da90e694feb85f91fa27935081ea23dd976d7e492db7dd9ef closes,pr,187481,issue,187480,high,pr.closingIssuesReferences,pr #187481 declares a closing reference to issue #187480.,https://github.com/pytorch/pytorch/pull/187481,0b8dcf73c1e1eafdfa9bd6559fb9593dbd968a793770f6de87b3ce1afbe78f02 closes,pr,187481,issue,187480,high,pr.body,"but the FBGEMM version does not. This PR adds the overflow check to the FBGEMM copy, fixing some quantization test failures on GB200. Fixes #187480 Authored with Claude cc @eqy",https://github.com/pytorch/pytorch/pull/187481,69a4e24405ee363e0933c778b7067d22970a4bf582b03494ce632ef5c210d7b1 references,pr,186245,issue,188602,medium,pr.body,tructed on a best-effort basis by BOLT. I am not fully aware if this causes any degradation in the debugger experience or stacktraces. RFC: #188602 Assisted-by: Claude Opus 4.8,https://github.com/pytorch/pytorch/pull/186245,995bf75ac4ae524ceda7e954f68ac76dff4dc26d7d0f07232eb67a99576bae4a review guidance,pr,186245,issue,188602,high,pr.reviews[0].body,"Some Open Questions Set up PyTorch CI with llvm-bolt How often do we refresh the profiles? Script for reporting profile staleness? How do we make the profiles easily auditable? Submodule? Do we need to BOLT optimize all libraries, or can we limit it to the large libraries only? Does this impact s...",https://github.com/pytorch/pytorch/pull/186245,4418b6b6a32d8ecce8c137354fbe8a39c1e1bd94a801c2b53c722f8a32c76d20 closes,pr,182972,issue,175608,high,pr.closingIssuesReferences,pr #182972 declares a closing reference to issue #175608.,https://github.com/pytorch/pytorch/pull/182972,7dff9ccea598c9c105d60e2ae405ef207a1ac2eb0338ff226cd6db60d946928e closes,pr,182972,issue,175608,high,pr.body,"fixes #175608 I realize this does not fix every try/except behavior error, but it does fix the repro with a pattern that can be resued. I can extend it o",https://github.com/pytorch/pytorch/pull/182972,68f3ee3df1f13b45366de3147f770cc6a51a231cc266baaa6f922c2fca77ae81 review guidance,pr,182972,issue,175608,high,pr.reviews[0].body,"What about an approach that takes advantage of Dynamo's observed exception handling, erroring out if a FakeTensor error bubbles outside of the compiled region? On another note, a graph break in a try block skips the entire frame.",https://github.com/pytorch/pytorch/pull/182972,6e42c5795cee00c0768eee9b81efd59de4cc15c93372cb02e8222c4cff5a87e1 closes,pr,188605,issue,188547,high,pr.closingIssuesReferences,pr #188605 declares a closing reference to issue #188547.,https://github.com/pytorch/pytorch/pull/188605,9c910e26bc603e2263d6172fca9d12ddb97a39f8e226dfc2aabf14d907027c53 closes,pr,188605,issue,188547,high,pr.body,"Fixes #188547 Fixes infinite recursion when calling int( ), float( ), or indexing pybind11 enums in compiled functions. Basically detects GetAttrVariable",https://github.com/pytorch/pytorch/pull/188605,2b7eae6d1db2c48aa97f709610bb00d8f85a09cb9c9abbdbc0978e50b22ad8ff review guidance,pr,188605,issue,188547,high,pr.reviews[0].body,"Overall looks good, I think we can simplify the test, thanks!",https://github.com/pytorch/pytorch/pull/188605,01e33dd4048d808290ff1e4f136d759d275ae2ec8f992dfa9dd7555f7812953e closes,pr,184276,issue,183988,high,pr.closingIssuesReferences,pr #184276 declares a closing reference to issue #183988.,https://github.com/pytorch/pytorch/pull/184276,01cf1e30fbc9c5ce40883e052072e0d4f7867a4eecf582bde862fed61581d99e closes,pr,184276,issue,183988,high,pr.body,"Fixes: #183988 In [1]: import torch ...: ...: def fn(x): ...: idx = torch.randperm(x.shape[0], device=x.device) ...: return x[idx] ...: ...: x = torch.ran",https://github.com/pytorch/pytorch/pull/184276,7adcfe09585146a2827f8729acc0d11105564df36e7854cad19d8b271698066e references,pr,184276,issue,188721,medium,pr.comments[0].body,"akiness on trunk: B200 Smoke Tests / linux-jammy-cuda13.0-py3.12-gcc11-sm100 / test-osdc (smoke_b200, 1, 1, mt-l-x86iamx-22-225-b200) (gh) (#188721) test/inductor/test_flex_flash.py::TestFlexFlashCUDA::test_captured_table_int64_index_cuda vllm-test / vllm-x-pytorch-test / test...",https://github.com/pytorch/pytorch/pull/184276,a8446bf3c78811ce46f5e06a530b9c2cd9581d83a59078ba4b675b30b5d21c5e review guidance,pr,184276,issue,183988,high,pr.reviews[0].body,fix failure,https://github.com/pytorch/pytorch/pull/184276,496b023a773dd4995000e15c0ec20ff139e34ff83f998cc9d2752a7ac6709dac closes,pr,185084,issue,165032,high,pr.body,r boolean behavior for normal contiguity checks while keeping or_false metadata queries conservative when singleton-ness is symbolic. Fixes #165032 Generated by my agent Test Plan: python test/test_dynamic_shapes.py TestPySymInt.test_contiguous_metadata_does_not_guard_on_symbo...,https://github.com/pytorch/pytorch/pull/185084,590b9b6dfe6eca6376dcb92abc86eea176be4fcfcb068f72ef5b638318617837 closes,pr,185460,issue,156511,high,pr.body,improvement. The checker also accepts abs_latency as a metric so the CPU path can move to latency baselines once those targets exist. Fixes #156511 Generated by my agent Test Plan: python test/dynamo/test_check_perf_csv.py python benchmarks/dynamo/check_perf_csv.py --help lint...,https://github.com/pytorch/pytorch/pull/185460,f9a5237104c54d3d48e366717d385e29cc7339db4c7e0ed7e375a8dc3a6b7a4a closes,pr,185946,issue,185240,high,pr.closingIssuesReferences,pr #185946 declares a closing reference to issue #185240.,https://github.com/pytorch/pytorch/pull/185946,a0c515681d05a1af4efc50bef0652f19dfdfc407fdd1a3b476cda76c86b52aaf closes,pr,185946,issue,185240,high,pr.body,"ns for successful NVML calls with softer checks and warnings to prevent hard crashes on Tegra, which does not have full NVML support. Fixes #185240",https://github.com/pytorch/pytorch/pull/185946,382b0e27f6ed8618c88c78136c7dbe4fb11580bbf40d2e82f26f0d58e2d4f090 review guidance,pr,185946,issue,185240,high,pr.reviews[0].body,Does it make sense to diverge the behavior here depending on compute capability or would that be too complex?,https://github.com/pytorch/pytorch/pull/185946,81b2e5b84d62af421c0ab7ac0465a4118ff95685eb613cd2ccd3bff7b360840c closes,pr,185299,issue,161067,high,pr.body,"ch.test_fx_graph_cse_preserves_impure_barriers, TestAOTDispatch.test_aot_dispatch_output_requires_grad_in_no_grad_views lintrunner -a Fixes #161067 Generated by my agent",https://github.com/pytorch/pytorch/pull/185299,35c87bda72c560f48a140184de007a38056f4cd17859e676f38f1f0f66a15f35 review guidance,pr,185299,issue,161067,high,pr.reviews[0].body,"Yeah, I think the Richard problem has to be reasoned through more carefully. Also, we are also planning on sending full train-step with forward-backward into the compiler, and we need to not pessimize in this case @IvanKobzarev",https://github.com/pytorch/pytorch/pull/185299,1e258ad6213ec7cc11b2d607e7d28fd389b5fcdba0c2acb1ab6b40309fc77708 review guidance,pr,185299,pr,185299,high,pr.reviews[0].body,"Yeah, I think the Richard problem has to be reasoned through more carefully. Also, we are also planning on sending full train-step with forward-backward into the compiler, and we need to not pessimize in this case @IvanKobzarev",https://github.com/pytorch/pytorch/pull/185299,f284fc8085a07dcf3647ea258e5aa2b864da0f1c39c33b1721801e4d9ec375bb review guidance,pr,185299,issue,161067,high,pr.reviews[1].body,"There are a few reasons to be careful about CSE: increased liveness, increasing memory not being able to fuse into uses. interactions with side affects/mutations. i think we want to be more careful before rolling this out.",https://github.com/pytorch/pytorch/pull/185299,192db0723d1a52694d74a81d30c060f761cc83f3f1ed5b5602bb6bde5db13865 review guidance,pr,185299,pr,185299,high,pr.reviews[1].body,"There are a few reasons to be careful about CSE: increased liveness, increasing memory not being able to fuse into uses. interactions with side affects/mutations. i think we want to be more careful before rolling this out.",https://github.com/pytorch/pytorch/pull/185299,8981dbb4aacf3ffe4504a426d981b40768d0c2af7cf7529fe7de953fca837f51 closes,pr,185300,issue,161053,high,pr.body,"ate AttrProxy compile microbenchmark over 50 iterations: main median 15.008 ms, mean 15.084 ms; fix median 15.467 ms, mean 15.574 ms. Fixes #161053 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...",https://github.com/pytorch/pytorch/pull/185300,6af54978d864c469c4685a818533ddc94dcabd99906cea14812faf162f6a69e1 closes,pr,181860,issue,163836,high,pr.closingIssuesReferences,pr #181860 declares a closing reference to issue #163836.,https://github.com/pytorch/pytorch/pull/181860,dafc37dd83b8d43ee3eaf753a14ae1b2d925f725ebd77505d0e20849c6d4e72a closes,pr,181860,issue,163836,high,pr.body,"Backends can migrate to the new signatures incrementally; once all backends have migrated, the legacy void* methods can be removed. Fixes: #163836",https://github.com/pytorch/pytorch/pull/181860,6d68064bb37b801f6d9a6e2d2cd8a1e1d1eafc013af9a1c2859d59d8ae87531f closes,pr,185320,issue,160544,high,pr.body,"direct torch.ops.higher_order.scan calls, but that bypasses the frontend flattening and metadata contract the Dynamo path relies on. Fixes #160544 Generated by my agent Benchmark Results: Micro-benchmark of the added scan partial validation in isolation: before median: 0.01840...",https://github.com/pytorch/pytorch/pull/185320,17a8a6c61fbe0254b6d5ca1cf2eeda0f31d51e3bed2cf6ed17f2cfb5f916ad41 closes,pr,184052,issue,160388,high,pr.body,backend AOT graphs agree with export on preserved functional CIA ops. Keep direct aot_export_module decomposition behavior unchanged. Fixes #160388 Generated by my agent,https://github.com/pytorch/pytorch/pull/184052,b6e2f10a617eed197046574235d9596bba729f206a77b6133ee21fb173f82154 review guidance,pr,184052,issue,160388,high,pr.reviews[0].body,"change has code motion which makes it difficult for human to review in github ui, please amend agent code authoring instructions",https://github.com/pytorch/pytorch/pull/184052,3458483113833af36f12fbcbcbf3f89716ba1bbdc2b7005ab7fcbdcdd0d29dbf review guidance,pr,184052,pr,184052,high,pr.reviews[0].body,"change has code motion which makes it difficult for human to review in github ui, please amend agent code authoring instructions",https://github.com/pytorch/pytorch/pull/184052,179ce8866674d911fac2a7c0b071daf42a906c17d37ac13441c2ba90df12cfe2 review guidance,pr,184052,issue,160388,high,pr.reviews[1].body,Needs a test for the decompositions=None behavior... Approving w/ comments,https://github.com/pytorch/pytorch/pull/184052,006dc26efcbd790146a7aa6d5988fc033483ca626399ad56555c7ea2b21555ce review guidance,pr,184052,pr,184052,high,pr.reviews[1].body,Needs a test for the decompositions=None behavior... Approving w/ comments,https://github.com/pytorch/pytorch/pull/184052,386737ba7e60c084a9dc8cb06ebb66f0360c1042d7b0db1912d7015d35225341 closes,pr,184059,issue,160124,high,pr.body,"m scheduler-local IO so generated metadata reflects each kernel's actual inputs and outputs, including fused producers and epilogues. Fixes #160124 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...",https://github.com/pytorch/pytorch/pull/184059,f2307f471aae83adb8e15d64f159ad31c88b9e0932b46075bef4f215138a44c8 closes,pr,187012,issue,160094,high,pr.body,"lt-in, typing.List, typing.Sequence, and optional-element spellings, plus mutation/torch.compile coverage for Optional[List[Tensor]]. Fixes #160094 Generated by my agent Test Plan: python test/test_custom_ops.py TestCustomOp.test_infer_schema_supported TestCustomOpAPI.test_mut...",https://github.com/pytorch/pytorch/pull/187012,0e00fcfce17bf2341f163c21d24f1ca383b5e3d4b9403c7820c2658bc252b94f closes,pr,185336,issue,160083,high,pr.body,"ring a later forward pre-hook with prepend=True mutates the module hook OrderedDict order with move_to_end(last=False). In the failure from #160083, Dynamo traced hooks in OrderedDict iteration order but used guard sources derived from dict.keys(...) positions. When those orde...",https://github.com/pytorch/pytorch/pull/185336,7ca35acca0f28ce3258f1dbe6226f7f09d6916b879131fd2ddcbe9e504f6310e closes,pr,185347,issue,136586,high,pr.body,because the compiled graph only needs to distinguish whether runtime dispatch should preserve FakeTensorMode semantics. Fixes #160057 Fixes #136586 Generated by my agent Test Plan: python test/dynamo/test_misc.py MiscTests.test_compile_with_userland_fake_tensor_mode python tes...,https://github.com/pytorch/pytorch/pull/185347,cb854e3c682bae1321f376e35143b4940a61a8c65ab0733321361be9eb2013d5 closes,pr,185347,issue,160057,high,pr.body,is sufficient because the compiled graph only needs to distinguish whether runtime dispatch should preserve FakeTensorMode semantics. Fixes #160057 Fixes #136586 Generated by my agent Test Plan: python test/dynamo/test_misc.py MiscTests.test_compile_with_userland_fake_tensor_m...,https://github.com/pytorch/pytorch/pull/185347,3957d3252ef788c18fcfd06f4bd4d3102ca375c87d50ec2fa5333d73884efd50 review guidance,pr,185347,issue,136586,high,pr.reviews[0].body,"What are we trying to support exactly ? It's not clear to me what parts of the stack are meant to be recreated, and under what fidelity. Could we get context from the issue raiser what they want?",https://github.com/pytorch/pytorch/pull/185347,d03e9f0c9589a62745afac5177a9459e484dd96869932b4ab86617e29df2ed70 review guidance,pr,185347,issue,160057,high,pr.reviews[0].body,"What are we trying to support exactly ? It's not clear to me what parts of the stack are meant to be recreated, and under what fidelity. Could we get context from the issue raiser what they want?",https://github.com/pytorch/pytorch/pull/185347,c313acf60e6e01492335c0e251341821d2d2dd06bf35a75b6bf71bef86380f45 review guidance,pr,185347,pr,185347,high,pr.reviews[0].body,"What are we trying to support exactly ? It's not clear to me what parts of the stack are meant to be recreated, and under what fidelity. Could we get context from the issue raiser what they want?",https://github.com/pytorch/pytorch/pull/185347,be5cc185fb96c88617328d37e4526113deb51047bda317ecfbe2e3574575f66e review guidance,pr,185347,issue,136586,high,pr.reviews[1].body,"What are we trying to support exactly ? It's not clear to me what parts of the stack are meant to be recreated, and under what fidelity. Could we get context from the issue raiser what they want?",https://github.com/pytorch/pytorch/pull/185347#pullrequestreview-4421071588,a235ec6d3054bedec522ca03e9a1ba020733274c2dff2c8f75c42c61f7e6333d review guidance,pr,185347,issue,160057,high,pr.reviews[1].body,"What are we trying to support exactly ? It's not clear to me what parts of the stack are meant to be recreated, and under what fidelity. Could we get context from the issue raiser what they want?",https://github.com/pytorch/pytorch/pull/185347#pullrequestreview-4421071588,9bfc5b964e318c1f2f1951fce39c43802aa0a986593b891274b2ffbf9f566f5e review guidance,pr,185347,pr,185347,high,pr.reviews[1].body,"What are we trying to support exactly ? It's not clear to me what parts of the stack are meant to be recreated, and under what fidelity. Could we get context from the issue raiser what they want?",https://github.com/pytorch/pytorch/pull/185347#pullrequestreview-4421071588,d04fae509b26628698f7a6296641c93b1859b1c0d93d9bf75787b9f8765ac180 closes,pr,185349,issue,160018,high,pr.body,n.py -k test_backend_triton_decode -v lintrunner -a torch/_inductor/kernel/flex/flex_decoding.py test/inductor/test_flex_attention.py Fixes #160018 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/185349,e284482c7c6ceec240f0ff189d68e81493ab11293bd15bb5c4a5e96aaae327e1 closes,pr,185353,issue,159918,high,pr.body,"original1 naming globally, since those names are part of parametrization internals and changing them would be a much broader BC risk. Fixes #159918 Generated by my agent Test Plan: python test/export/test_export.py TestExport.test_subclasses_parameterization TestExport.test_su...",https://github.com/pytorch/pytorch/pull/185353,488e91e7d2701e5bed7c6b44e934d07dc1b8bdfebda2d0bfcba3dc58fcc39b8d closes,pr,185356,issue,159843,high,pr.body,"he lazy conversion as a compatibility path for compiled reductions that still pair a FunctionalTensor mask with a plain local tensor. Fixes #159843 Generated by my agent Benchmark Results: CPU fake-PG microbenchmark, world size 4, measured with torch.utils.benchmark. Command u...",https://github.com/pytorch/pytorch/pull/185356,e6e49db7e2e193c0154d00b5c726b7d4b7fba7ccac3a528231e0d6330e6cbbaa closes,pr,186003,issue,144820,high,pr.body,"repeatedly re-enters Dynamo for the same skipped branch and does not implement the guard-based fallback model requested by the issue. Fixes #144820 Generated by my agent Benchmark Results: Repeated skipped-call micro-benchmark, 5,000 calls per repeat, 5 repeats, backend='eager...",https://github.com/pytorch/pytorch/pull/186003,c5cd06c6c2fd4a0cb527e2c298a3159977bd92acee5e19e764edb42701d13167 closes,pr,185357,issue,159831,high,pr.body,r membership fall back through getitem(0). Non-container modules continue to raise the underlying TypeError from the original module. Fixes #159831 Generated by my agent Test Plan: python - <<'PY' manual repro from #159831 before and after the fix python test/dynamo/test_modul...,https://github.com/pytorch/pytorch/pull/185357,0aefbb3ded09be45a4b68f0da3243bf274951d655609cd67a39d026d94691ed7 review guidance,pr,185357,issue,159831,high,pr.reviews[0].body,"deferring to others on if this in scope, im not sure it should be.",https://github.com/pytorch/pytorch/pull/185357,c674c67a1f0812692c8514c794690d685673cb613e8d6eeb10ad1a4e2979d1c1 review guidance,pr,185357,pr,185357,high,pr.reviews[0].body,"deferring to others on if this in scope, im not sure it should be.",https://github.com/pytorch/pytorch/pull/185357,649fc24a1b2013175fb45913f61f84d5a9adace6350ff460b8c2ea097a7b2411 closes,pr,188698,issue,188772,high,pr.closingIssuesReferences,pr #188698 declares a closing reference to issue #188772.,https://github.com/pytorch/pytorch/pull/188698,ae03321070aa4b31ad5fcdf86f09ea5beb02c32a8d04f5c795bda96675b1f44a closes,pr,188698,issue,188772,high,pr.body,Fixes #188772 This PR makes FsspecReader and FsspecWriter public by renaming _fsspec_filesystem.py to fsspec_filesystem.py and exposing them in the torch,https://github.com/pytorch/pytorch/pull/188698,5ce7280544f0d1c6464c24b5cf160fce756bc7d58f77e4073621f40b9c0502c0 review guidance,pr,188698,issue,188772,high,pr.reviews[0].body,Can we keep deprecated aliases with the underscore to avoid breaking people's downstream code?,https://github.com/pytorch/pytorch/pull/188698,977241effd5ca54dcdfca3c40a3052396b2df4bc33f9bf00da2dcdc62cf8ee3c review guidance,pr,188698,issue,188772,high,pr.reviews[1].body,overall LGTM and agree with the problem described in the issue :) thank you for fixing this!,https://github.com/pytorch/pytorch/pull/188698,c95c71da8d72776c4612cba867985d10fc1cb1afab34764e591a4daad7ac4e9d review guidance,pr,188946,pr,111774,high,pr.reviews[0].body,Pull request overview Fixes a correctness issue in legacy FSDP optimizer-state reconstruction for TP+FSDP by capturing the out-of-place return value from DTensor.redistribute() during _unflatten_orig_param_states(). Changes: Assign the return value of DTensor.redistribute() before subsequent resh...,https://github.com/pytorch/pytorch/pull/188946,0fa495885d8d0c155d183be80a7b891098d49511904b31ccc3e7aba331bbaf5b review guidance,pr,188946,pr,111774,high,pr.reviews[1].body,"Addressed all 3 Copilot review comments: Hardcoded line number (Medium): Removed (line 1463) from docstring, now references torch/distributed/fsdp/_optim_utils.py by path only. Test scope (High): Acknowledged, the docstring already explicitly states this is an API-contract test, not an integratio...",https://github.com/pytorch/pytorch/pull/188946,de53a4eec08fad0c2c08e1280135d234b67046b336e6a3d04395c6b1924f8f0d references,pr,188816,issue,149325,medium,pr.body,"Summary Addresses the nonzero op case of #149325 (the umbrella tracking MPS 32-bit index limits). MPS nonzero rejected any tensor with numel >= INT_MAX, and the internal flat accumulator i",https://github.com/pytorch/pytorch/pull/188816,980eea7cff46e6400e74e0af0b9f70fe45d3cae485bf7bcb1ed598cfa39d69a4 review guidance,pr,188816,issue,149325,high,pr.reviews[1].body,"Pull request overview This PR updates the MPS implementation of nonzero to handle tensors with numel >= INT_MAX by removing the hard guard, widening host-side sizing math to 64-bit, and widening the Metal scatter index decomposition accumulator to 64-bit; it also adds new MPS tests for nonzero, i...",https://github.com/pytorch/pytorch/pull/188816,85205bea027b493dcedc078940b13a8ce4c9d0fb8fe992c18dced0539cdcb126 review guidance,pr,177540,pr,177540,high,pr.reviews[0].body,address automated feedback from claude,https://github.com/pytorch/pytorch/pull/177540,f7496e0d6d59ffc3e58d36fd3da2c434bdfbbc3995a9df8d3c9fbe27a4a24574 review guidance,pr,188396,pr,188396,high,pr.reviews[0].body,"SGTM, thank you for contributing this one.",https://github.com/pytorch/pytorch/pull/188396,14c802e684c9e676cb7798a74912d4029b476b73624b26ffcde3464924ae694c review guidance,pr,188396,pr,188396,high,pr.reviews[1].body,solid stuff,https://github.com/pytorch/pytorch/pull/188396,3857cc62e8ea6532a1a6d65c5de4aed6f73ce5791820a831390acd81d74f1e64 closes,pr,186622,issue,168126,high,pr.body,mposed/fused path. aten.var_mean is also excluded from the Inductor decomposition table so it reaches the coherent var_mean lowering. Fixes #168126 Generated by my agent Test Plan: python -m compileall -q torch/_inductor/autocast_utils.py torch/_inductor/codecache.py torch/_in...,https://github.com/pytorch/pytorch/pull/186622,96f2d9b32b901575d708e350aae73f0391a6390c981f312791ea186af1d202e7 closes,pr,185365,issue,159613,high,pr.body,"ase lists or cat arguments, but the bug is in generic placeholder remapping and can affect any supported FX aggregate argument shape. Fixes #159613 Generated by my agent Test Plan: git diff --cached --check pytest -q test/fx/test_subgraph_rewriter.py::TestSubgraphRewriter::tes...",https://github.com/pytorch/pytorch/pull/185365,b8b706b29ab6fb82e74fe98480d9fa47765120097fc1f719520957d676973c07 review guidance,pr,185365,issue,159613,high,pr.reviews[0].body,this pattern matcher is not in use in compile - deferring to @SherlockNoMad on this one,https://github.com/pytorch/pytorch/pull/185365,c0dbad58814186edbaa592088399d61021499d060fefa91ec49511df73bc3b15 review guidance,pr,185365,pr,185365,high,pr.reviews[0].body,this pattern matcher is not in use in compile - deferring to @SherlockNoMad on this one,https://github.com/pytorch/pytorch/pull/185365,113a2473c2db54b42031f5476c4de8080c81ba80d6cc4179f238528a427729e2 review guidance,pr,188331,pr,188331,high,pr.reviews[0].body,"Hi thanks for the contribution! If the intent here is to make TestProfilerTree more device-agnostic, I'll suggest instead following the typical pytorch test framework pattern where we generate device-specific test classes using instantiate_device_type_tests() + appropriate decorator-based skips a...",https://github.com/pytorch/pytorch/pull/188331,a27cfdde65afbec6490b036fa168cd6d03c7e3fa8f6aee80412e3ac37fd1e830 closes,pr,188828,issue,188757,high,pr.closingIssuesReferences,pr #188828 declares a closing reference to issue #188757.,https://github.com/pytorch/pytorch/pull/188828,9c51d2efcbb37d0cb380e2a05e42e52930747fcf72603c60bcdf86cd72a31cd3 closes,pr,188828,issue,188757,high,pr.body,"as Linux. This issue disabled the test, but it was closed #182869 and moved to in source skip here. That skip does not cover windows Fixes #188757 Test plan python -m pytest .\pytorch\test\test_overrides.py Authored with help of an AI assistant cc @nWEIdia @ptrblck",https://github.com/pytorch/pytorch/pull/188828,cda059c46aa18ea566189440efe4cbe2b4e6b4d54febf2aed36696150b5b139a closes,pr,185369,issue,159550,high,pr.body,"he first set reuses one compiled frame, exercises changing input shapes for correctness, and covers the large-window fallback branch. Fixes #159550 Fixes #185575 Generated by my agent Test Plan: python test/inductor/test_cpu_repro.py -k test_adaptive_avg_pool2d_dynamic_input_o...",https://github.com/pytorch/pytorch/pull/185369,e058a79b99a93b97f9dda5cd4339f23547c856b3f336f94f830c467ad8aaa6b7 closes,pr,185369,issue,185575,high,pr.body,"euses one compiled frame, exercises changing input shapes for correctness, and covers the large-window fallback branch. Fixes #159550 Fixes #185575 Generated by my agent Test Plan: python test/inductor/test_cpu_repro.py -k test_adaptive_avg_pool2d_dynamic_input_output_sizes No...",https://github.com/pytorch/pytorch/pull/185369,5be8597fa12ab5b36f0834d6a17fbdf0efdc0cca3a0a46bfb01f855c7a6f638d closes,pr,185371,issue,159468,high,pr.body,"before median 15.885 us, after median 16.225 us, +2.15% on this small issue-sized input. Test Plan: Inline ONNX/ORT reproduction for issue #159468 now returns output spatial shapes 16x10, 16x5, and rounded 16x9. python test/test_nn.py TestNN.test_interpolate_bilinear_antialias...",https://github.com/pytorch/pytorch/pull/185371,0ac9b816a0dc0b93fdb2af91c232c42087b4734147ea6e0ff06ba4f8beaf9470 review guidance,pr,185371,issue,159468,high,pr.reviews[1].body,"Pull request overview This PR fixes a Dynamo export / ONNX export shape specialization bug when F.interpolate(..., antialias=True) is driven by symbolic scale_factor values (e.g., 16 / x.shape[2]). The key goal is to prevent symbolic scale factors from being coerced through float[] (which can for...",https://github.com/pytorch/pytorch/pull/185371,b191c2b8d951503448e4925d8b4672c284177aa765eea0836d27f02d3de1c53f review guidance,pr,185371,pr,185371,high,pr.reviews[1].body,"Pull request overview This PR fixes a Dynamo export / ONNX export shape specialization bug when F.interpolate(..., antialias=True) is driven by symbolic scale_factor values (e.g., 16 / x.shape[2]). The key goal is to prevent symbolic scale factors from being coerced through float[] (which can for...",https://github.com/pytorch/pytorch/pull/185371,d84c827199002d8998fcd456976bff370ccc79564a3efec39163bf873653025d review guidance,pr,185371,issue,159468,high,pr.reviews[3].body,"## Pull request overview This PR fixes a Dynamo export / ONNX export shape specialization bug when `F.interpolate(..., antialias=True)` is driven by **symbolic** `scale_factor` values (e.g., `16 / x.shape[2]`). The key goal is to prevent symbolic scale factors from being coerced through `float[]`...",https://github.com/pytorch/pytorch/pull/185371#pullrequestreview-4401919356,da810479dc3b9126fea48c1233848346a3e44ea6cef1c819a3f5fbfc60281e4d review guidance,pr,185371,pr,185371,high,pr.reviews[3].body,"## Pull request overview This PR fixes a Dynamo export / ONNX export shape specialization bug when `F.interpolate(..., antialias=True)` is driven by **symbolic** `scale_factor` values (e.g., `16 / x.shape[2]`). The key goal is to prevent symbolic scale factors from being coerced through `float[]`...",https://github.com/pytorch/pytorch/pull/185371#pullrequestreview-4401919356,d7dfb35646fd61e5fcc47543ea9efc8c950d691c8001f53d48fe14f16ea9dc78 review guidance,pr,184648,pr,184648,high,pr.reviews[0].body,"Hi, @orangeH25. Thanks for your contribution! I left some comments, PTAL.",https://github.com/pytorch/pytorch/pull/184648,60e347f44dbc7ed2f65dffb0f3d0372716df9ca9227ceac69d6f92069c5b4a92 closes,pr,185373,issue,159457,high,pr.body,"neric Tensor registration path, but limiting this to assume_constant_result avoids changing unrelated ConstantSource Tensor handling. Fixes #159457 Generated by my agent Test Plan: python inline script from #159457 using torch.compile(model, backend=""eager"") python test/dynamo...",https://github.com/pytorch/pytorch/pull/185373,ef2d31898eab246f5cbd5513585e8f1b8c2e0fb86960b8e57d28c2a0c72d752f closes,pr,185892,issue,184841,high,pr.body,tches. Keeping the translation limited to direct FakeTensorDeviceMismatchError preserves other export errors on their existing paths. Fixes #184841 Generated by my agent Test Plan: python test/export/test_export.py TestExport.test_export_fake_tensor_device_mismatch_error TestE...,https://github.com/pytorch/pytorch/pull/185892,5c863867dc29db2519cb9ecfb71d709fbf7fd3f6037f3c46ba5199b250a27ce9 references,pr,185892,issue,184841,medium,pr.comments[2].body,"`isinstance(exc, ...)` check won't match. This is acceptable per the stated design (""limited to direct""), but worth noting that issues like #184841 could recur if wrapping is introduced upstream. --- **Summary:** The implementation is clean, well-scoped, and the tests cover bo...",https://github.com/pytorch/pytorch/pull/185892,67bbf5529f7c0abe6eec98d6fce7c59394956d0b3a3c1e66e960739ed6a517c9 review guidance,pr,185892,issue,184841,high,pr.reviews[0].body,"Basically, I'm not sure if the change is needed since the existing behavior already matches the eager behavior, and the current implementation of parsing the error message seems a little fragile",https://github.com/pytorch/pytorch/pull/185892,dc49d42572736662fa5cbc31e3903480959f65760e7d5ec8e3a331977e388e61 review guidance,pr,185892,pr,185892,high,pr.reviews[0].body,"Basically, I'm not sure if the change is needed since the existing behavior already matches the eager behavior, and the current implementation of parsing the error message seems a little fragile",https://github.com/pytorch/pytorch/pull/185892,c646abed18203b7c1ced9aa0267d3920596639cba2e1379ab698847f9b910557 review guidance,pr,185892,issue,184841,high,pr.reviews[1].body,"My agent says I addressed the fragile parsing feedback by carrying structured inner exceptions instead of parsing formatted strings, and this update adds direct coverage for the _prims_common same-device path plus stronger assertions for the forward-created tensor case. The design intent is still...",https://github.com/pytorch/pytorch/pull/185892,783448882cce2089f7a9bc76c3e35ab796436e669ed672fb1efc519b2a1f1679 review guidance,pr,185892,pr,185892,high,pr.reviews[1].body,"My agent says I addressed the fragile parsing feedback by carrying structured inner exceptions instead of parsing formatted strings, and this update adds direct coverage for the _prims_common same-device path plus stronger assertions for the forward-created tensor case. The design intent is still...",https://github.com/pytorch/pytorch/pull/185892,7afd8864f1b8caf234e8ab8be3cecbdb61e71edf9e4ace5db30948772fd99a95 review guidance,pr,185892,issue,184841,high,pr.reviews[2].body,"Basically, I'm not sure if the change is needed since the existing behavior already matches the eager behavior, and the current implementation of parsing the error message seems a little fragile",https://github.com/pytorch/pytorch/pull/185892#pullrequestreview-4501234595,e72273df083415c51822a7e6a77ad698e15174b41de2e486912fdd88a3a8b1eb review guidance,pr,185892,pr,185892,high,pr.reviews[2].body,"Basically, I'm not sure if the change is needed since the existing behavior already matches the eager behavior, and the current implementation of parsing the error message seems a little fragile",https://github.com/pytorch/pytorch/pull/185892#pullrequestreview-4501234595,91b01252d812ca16d6bff7496bf681c2e0285393f8193159e62f7a81f9136537 review guidance,pr,185892,issue,184841,high,pr.reviews[3].body,"My agent says I addressed the fragile parsing feedback by carrying structured inner exceptions instead of parsing formatted strings, and this update adds direct coverage for the _prims_common same-device path plus stronger assertions for the forward-created tensor case. The design intent is still...",https://github.com/pytorch/pytorch/pull/185892#pullrequestreview-4539118084,af6df9731709e5d222c14fca76e88cb04b61cfbdd85884f08471aa32911c9f15 review guidance,pr,185892,pr,185892,high,pr.reviews[3].body,"My agent says I addressed the fragile parsing feedback by carrying structured inner exceptions instead of parsing formatted strings, and this update adds direct coverage for the _prims_common same-device path plus stronger assertions for the forward-created tensor case. The design intent is still...",https://github.com/pytorch/pytorch/pull/185892#pullrequestreview-4539118084,57109a935cb6bb56a94c521d16235ecf8782bebeee37acdaa9af9550d0c46b03 closes,pr,184103,issue,152548,high,pr.body,en they are lowerable. Seed subgraph updater work from placeholder metadata updates so dtype changes propagate through nested graphs. Fixes #152548 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/184103,36ef0bb3ef20abc5f4485db819d32c078fc47557e4d0d15b156a94da5527ec94 references,pr,184103,issue,152548,medium,pr.review_comments[0].body,"ed. The updater should not cache invalid/unprocessed nodes, or the hash needs to encode fake-arg validity. ### Original Issue Summary Issue #152548 reported that FakeTensorUpdater could leave stale fake metadata after manual FX graph edits. Two root causes were called out: HOP...",https://github.com/pytorch/pytorch/pull/184103,4c4e8bf97ff5176fe3e0c00933b2644cef282eedf4dbbc747587cbbb75fc6849 review guidance,pr,184103,issue,152548,high,pr.reviews[4].body,surprised this is that big of a pr.. deferring to @benjaminglass1,https://github.com/pytorch/pytorch/pull/184103#pullrequestreview-4479182599,e42d9cfc417173adad1ebca03acd9d645d83593dcdbee1bc27631d9bbad10246 review guidance,pr,184103,pr,184103,high,pr.reviews[4].body,surprised this is that big of a pr.. deferring to @benjaminglass1,https://github.com/pytorch/pytorch/pull/184103#pullrequestreview-4479182599,8ebdc94a56132ca87ff9b142f461129d57b0b29199a82cc8d8adc95dfd3bfb6d review guidance,pr,184103,issue,152548,high,pr.reviews[5].body,"Rebasing, then I can review.",https://github.com/pytorch/pytorch/pull/184103#pullrequestreview-4486429657,d2aaa8e92f92d4991fb503ce3df3feec20fb3c07cde4aa91828fd79ee9ec8ab5 review guidance,pr,184103,pr,184103,high,pr.reviews[5].body,"Rebasing, then I can review.",https://github.com/pytorch/pytorch/pull/184103#pullrequestreview-4486429657,02974df2700ef52a3cdafab99ac9c6c483b44f88a99c25ab3a00596c83cedf90 review guidance,pr,184103,issue,152548,high,pr.reviews[6].body,"@jansel I haven't closely reviewed all the new tests, but those seemed reasonably comprehensive with a quick pass. There are at least one or two bugs that got reintroduced by this that I've called out, but this should hopefully be good to go soon.",https://github.com/pytorch/pytorch/pull/184103#pullrequestreview-4499715083,2502e758ff3b917cbbea58ec7dc087b190f09f28945d828b3f6bd484f9777adb review guidance,pr,184103,pr,184103,high,pr.reviews[6].body,"@jansel I haven't closely reviewed all the new tests, but those seemed reasonably comprehensive with a quick pass. There are at least one or two bugs that got reintroduced by this that I've called out, but this should hopefully be good to go soon.",https://github.com/pytorch/pytorch/pull/184103#pullrequestreview-4499715083,bc17bb76752fb4ff54d6e159d14ceebf440c120ad4892f24681b6d420f62ac66 review guidance,pr,184103,issue,152548,high,pr.reviews[15].body,"Test failures seem unrelated, and review comments are fixed.",https://github.com/pytorch/pytorch/pull/184103#pullrequestreview-4545390040,6951b38be8f2702fae9e04a83effd18c39c6fd53340c9bfdcfc15c92ac7a2312 review guidance,pr,184103,pr,184103,high,pr.reviews[15].body,"Test failures seem unrelated, and review comments are fixed.",https://github.com/pytorch/pytorch/pull/184103#pullrequestreview-4545390040,0e9b1316b8e2dbd25f27d877c8d5b4dd66150f9f708362a7acf64a9e690f6552 closes,pr,185404,issue,159346,high,pr.body,cher.py -k test_bmm_to_mm lintrunner -a torch/_inductor/fx_passes/post_grad.py test/inductor/test_pattern_matcher.py git diff --check Fixes #159346 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/185404,0d1486ab116d0d86016172bb407e4b3ae661a0f25d4aa739192374bf65b58cf3 review guidance,pr,185404,issue,159346,high,pr.reviews[0].body,"Looks good, but, can we do a bit more benchmarking",https://github.com/pytorch/pytorch/pull/185404,4e06e0a630941d9f7b297d3446730b419b427a050571fea30fb2ce53367476d4 review guidance,pr,185404,pr,185404,high,pr.reviews[0].body,"Looks good, but, can we do a bit more benchmarking",https://github.com/pytorch/pytorch/pull/185404,001d7f873e89dcfdb385732e1da263817bb1bcb0234112a9d5db7fbffd3d1193 closes,pr,184061,issue,159192,high,pr.body,r resize side effects in the meta kernel so the generated functional variant returns accurate metadata for per-row fake quantization. Fixes #159192 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/184061,ce9e6af015430ed9e63d02a0a803a93d93b5f2fcac802bc368d50fac02c58ab4 review guidance,pr,184061,issue,159192,high,pr.reviews[0].body,"I don't really like this operator. Do people actually use it? if they don't, then let's close the issue as a wontfix",https://github.com/pytorch/pytorch/pull/184061,c5adc5704db2eb48b8a4c01004c59a1b818732318100d517b55258da51b2567a review guidance,pr,184061,pr,184061,high,pr.reviews[0].body,"I don't really like this operator. Do people actually use it? if they don't, then let's close the issue as a wontfix",https://github.com/pytorch/pytorch/pull/184061,efeaf6bfcfe9ed55cf51f24c4384a63151a3fcfa83a2abeba101e5710b20e9bb review guidance,pr,184061,issue,159192,high,pr.reviews[1].body,my understandign is torch.ao is deprecated in favor of torchao? if you agree then close this as wontfix pytorch/ao#2259,https://github.com/pytorch/pytorch/pull/184061,c03a3ec9ae8eb09f583a87dfeb4cd8b9184ada840bcb9fff85272526f3e6908f review guidance,pr,184061,pr,184061,high,pr.reviews[1].body,my understandign is torch.ao is deprecated in favor of torchao? if you agree then close this as wontfix pytorch/ao#2259,https://github.com/pytorch/pytorch/pull/184061,840361e47fb5308514baba287f15631eaaa551ed80a526096573d49e71ba0d72 closes,pr,186894,issue,158540,high,pr.body,"o aten.index.Tensor and matches the strict export behavior, while preserving the existing rewrite for ordinary scalar tensor indices. Fixes #158540 Generated by my agent Test Plan: python test/export/test_export.py -k test_non_strict_vmap_tensor_indexing python test/export/tes...",https://github.com/pytorch/pytorch/pull/186894,2f8b400c4f6cc960b342621a15ef0ec20c1ecafa0b02999ffc004c7943e84129 closes,pr,184065,issue,158521,high,pr.body,d fast CUDA launchers pack them as 16-bit values instead of 32-bit floats. Add regression coverage for fp16 and bf16 scalar launches. Fixes #158521 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/184065,dc6f04134dfd57f1dd5ef8639f5d72f7232fdd2e0f6be45d0e5404846adc3f42 closes,pr,185419,issue,158384,high,pr.body,the ShortenTraceback re-raise and add a regression test that keeps the backend stack visible while ensuring the hint token is absent. Fixes #158384 Generated by my agent Test Plan: python test/dynamo/test_error_messages.py ErrorMessagesTest.test_backend_compiler_failed_no_verb...,https://github.com/pytorch/pytorch/pull/185419,f4dd463d7d263d33a9d5b56911a58b50fe573ddcbf92b359523b9149ff568a93 closes,pr,185436,issue,158087,high,pr.body,res are Dynamo bugs. Matching only the known wrapper cause in the assertion helper is narrower and keeps unexpected failures wrapped. Fixes #158087 Generated by my agent Test Plan: python test/test_testing.py TestFrameworkUtils.test_plain_assert_raises_does_not_import_dynamo -...,https://github.com/pytorch/pytorch/pull/185436,1dd67bb3a166390f28bcd8a19e59049b0ee6ed7350835c11ff53072bd58fc6f5 closes,pr,185438,issue,157926,high,pr.body,"ent, which Dynamo intentionally skips. Keeping the disable scoped preserves that behavior without mutating user-visible global state. Fixes #157926 Generated by my agent Benchmark Results: Microbenchmark command measured a cached torch.compile(..., backend=""eager"") call for a...",https://github.com/pytorch/pytorch/pull/185438,8ab736596cb8fb2341fbc85738c6f96a2b4a32eaa237d65a3f8311184fc65e27 review guidance,pr,184718,pr,184718,high,pr.reviews[0].body,"Hi @RiyaP2508, nice done! I left a comment, PTAL.",https://github.com/pytorch/pytorch/pull/184718,6fb5cd53b850a7e106b6b020fbb0da7209b5fe9cba1a156ee7a3a0ed6ca23966 closes,pr,185450,issue,156786,high,pr.body,"t.default and aten._unsafe_index_put.default, and it accounts for with_effects positional argument offsets. An older abandoned draft PR for #156786 only checked the last positional argument, which misses schemas with an accumulate argument and does not cover indices. This vers...",https://github.com/pytorch/pytorch/pull/185450,2032e60aba8813f24e5571b0eb73d46946ced023b42933e0470b4a16e958bb53 review guidance,pr,185813,pr,180237,high,pr.reviews[0].body,Review Summary: This replaces the recursive license-files globs with an explicit 59-path list and adds an audit test that re-derives that list by globbing the working tree. The core problem is structural: it freezes a single static license set while PyTorch builds CPU/CUDA/ROCm wheels from differ...,https://github.com/pytorch/pytorch/pull/185813,ea15af137b02997acd4809f2964c6fce0f117e8ef60d4143df73f75d8a8b1d07 review guidance,pr,185813,pr,180237,high,pr.reviews[1].body,"The following review was created together with Claude Code. Thanks for reworking this, @tvukovic-amd - the structure is now exactly what we discussed: explicit license-files, an explicit excluded list plus a per-path SPDX map in the manifest, and one shared audit driving both the test and the lin...",https://github.com/pytorch/pytorch/pull/185813,0dcc71542af3e4db9a5e128f14b5c674f5920a50e62429c2c1107a71200f267e review guidance,pr,187165,pr,187165,high,pr.reviews[0].body,address automated feedback from claude,https://github.com/pytorch/pytorch/pull/187165,c50a58b69d2a02e2a3eb65a3c04434ca2e29bd83f13e3e6584234316b198bb5a review guidance,pr,185816,pr,185816,high,pr.reviews[0].body,Nice! Why does this need to be specific for SVE256?,https://github.com/pytorch/pytorch/pull/185816,c61f75469aa1d37e0673e045d1209145a578fb830010c69b3579f63356861413 review guidance,pr,184946,pr,184946,high,pr.reviews[0].body,"Looks reasonable overall, but I left some comments about edge cases, test cases to try, as well as some small style nits. I would like to take another look at this one.",https://github.com/pytorch/pytorch/pull/184946,a2998dabf189426bd9b05c0827a926050d8296e935c8585267be296c895a4b53 closes,pr,185451,issue,156640,high,pr.body,ws the conversion to the actual user callback invocation so internal comptime plumbing assertions are still wrapped as compiler bugs. Fixes #156640 Generated by my agent Test Plan: python test/dynamo/test_exceptions.py ExceptionTests.test_internal_assertion_error_wrapped Excep...,https://github.com/pytorch/pytorch/pull/185451,b18e006a8df2d9a411e294234965b7aba4fd55597860fce23439957743a91ac2 closes,pr,184985,issue,169991,high,pr.body,".compile wrapper, or called torch.compile inside forward, Dynamo still tried to compile while AOT export was tracing fake tensors. In issue #169991 that made FlexAttention/create_block_mask trip over nested compile state during joint export. Run aot_export_joint_with_descripto...",https://github.com/pytorch/pytorch/pull/184985,2785efee88c581c1560463e01a1dd87321933f97d4863326990a61c3d41a40e2 competes with,pr,184985,issue,169991,medium,pr.review_comments[0].body,"is HOP-dependent -- the safe fix is to mirror the export branch rather than rely on which HOPs survive inlining.) `create_block_mask` from #169991 is a plain `torch.compile`, not a HOP-internal compile, so the guard preserves this PR's fix. ```python if ( torch.compiler._is_no...",https://github.com/pytorch/pytorch/pull/184985,6dc22fcbf2b40cf469bec4790d0751d8d53a44d305d3b7faee0f84e40b92fb11 review guidance,pr,184985,issue,169991,high,pr.reviews[0].body,Breaks tests,https://github.com/pytorch/pytorch/pull/184985,ce980b2b789be4192d0647d75875f9096a45a7ae50726be9e7a379191b1ef0a9 review guidance,pr,184985,pr,184985,high,pr.reviews[0].body,Breaks tests,https://github.com/pytorch/pytorch/pull/184985,ad80ebaa262e03165b9fd10a0a350edc9a6cd343cdc1fbc400fbbe09e7be34dc review guidance,pr,184985,issue,169991,high,pr.reviews[1].body,"@claude No claude, this change sure doesn't look minimal lmao. Why did the original author think it was simple? What made this complicated",https://github.com/pytorch/pytorch/pull/184985,4a99e99005d0be46eb2e68b0ec2b963e564932a8b2370ee1e9be94ec2f386880 review guidance,pr,184985,pr,184985,high,pr.reviews[1].body,"@claude No claude, this change sure doesn't look minimal lmao. Why did the original author think it was simple? What made this complicated",https://github.com/pytorch/pytorch/pull/184985,e915766d7947960e532209e2e0a499e279b8e5a5018fe2ff3afaef4a5f9d25b3 review guidance,pr,184985,issue,169991,high,pr.reviews[2].body,Breaks tests,https://github.com/pytorch/pytorch/pull/184985#pullrequestreview-4367782794,8984d5999118138ed12385f4640f6b5ff1724bdd3709e40b96111f80dc2e206d review guidance,pr,184985,pr,184985,high,pr.reviews[2].body,Breaks tests,https://github.com/pytorch/pytorch/pull/184985#pullrequestreview-4367782794,d2ccf5c6438bc39b6de3068bec939cbba0c83e1e2e90c5fdb842ade754983208 review guidance,pr,184985,issue,169991,high,pr.reviews[3].body,"@claude No claude, this change sure doesn't look minimal lmao. Why did the original author think it was simple? What made this complicated",https://github.com/pytorch/pytorch/pull/184985#pullrequestreview-4432977079,3b328ae0922647f1377a55ed396672974ea26f410c00cb412840fc9803e177c3 review guidance,pr,184985,pr,184985,high,pr.reviews[3].body,"@claude No claude, this change sure doesn't look minimal lmao. Why did the original author think it was simple? What made this complicated",https://github.com/pytorch/pytorch/pull/184985#pullrequestreview-4432977079,1ab933116627513e409ef7b16cffa928498af215eef5bf04551fe756ef5e3b71 review guidance,pr,184985,issue,169991,high,pr.reviews[4].body,"(Reviewed by me, assisted by AI) [blocker] The two new tests only exercise a plain `torch.compile(...)` callable, so they don't cover the HOP-internal-compile path -- which is exactly where the inline `eval_frame.py` finding shows a regression (a module calling `associative_scan`, exported via `a...",https://github.com/pytorch/pytorch/pull/184985#pullrequestreview-4566352722,55de1f9f651817ef5cbf4ff21a38801b4930116aaa416b56e93f1132af26eaa1 review guidance,pr,184985,pr,184985,high,pr.reviews[4].body,"(Reviewed by me, assisted by AI) [blocker] The two new tests only exercise a plain `torch.compile(...)` callable, so they don't cover the HOP-internal-compile path -- which is exactly where the inline `eval_frame.py` finding shows a regression (a module calling `associative_scan`, exported via `a...",https://github.com/pytorch/pytorch/pull/184985#pullrequestreview-4566352722,33d0558e7a68cdb9b48c5bc7e7bec29b075490ed9aa9e985ec8634c8a4728a30 closes,pr,188994,issue,188723,high,pr.closingIssuesReferences,pr #188994 declares a closing reference to issue #188723.,https://github.com/pytorch/pytorch/pull/188994,8ee661ebe18772be0932093117b158bba948c444b8767045e2757dea3d87c8ba closes,pr,188994,issue,188723,high,pr.body,Fixes #188723 ShapeEnv._transfer_foreign_expr_as_unbacked had an order-dependent bug: when a composite foreign expression (e.g. u0 + u1) was transferred,https://github.com/pytorch/pytorch/pull/188994,654d8377e30923de036505d6613ddb394e86b5892ffd5ad4e6ed52ca13610298 review guidance,pr,188994,issue,188723,high,pr.reviews[0].body,.,https://github.com/pytorch/pytorch/pull/188994,b3e4cdcf3402e7f886ca7187d282b8f3e67901bbdc1d8655ddb7e81f16062e79 closes,pr,185057,issue,167026,high,pr.body,"om the propagated metadata is the narrower root-cause fix because it covers all generated code that can allocate for the tagged node. Fixes #167026 Generated by my agent Test Plan: python test/dynamo/test_ctx_manager.py -k ""cuda_use_mem_pool"" -v python test/dynamo/test_ctx_man...",https://github.com/pytorch/pytorch/pull/185057,942231d320b92708c28ec283644a91580b4bb053a5a54d7625eb98374455f0c9 review guidance,pr,185057,issue,167026,high,pr.reviews[4].body,"This looks great! Thank you for doing this. I'd love to see some tests that use `mode='reduce-overhead'`, and I worry there are places that we might trip up `torch._inductor.config.triton.slow_path_cudagraph_asserts`, but if desired those could be follow-ups.",https://github.com/pytorch/pytorch/pull/185057#pullrequestreview-4479376455,68ff99d2e717ea3d55e758ceb703441e7c9898d0e80eab1624065ea8fd55f0f8 review guidance,pr,185057,pr,185057,high,pr.reviews[4].body,"This looks great! Thank you for doing this. I'd love to see some tests that use `mode='reduce-overhead'`, and I worry there are places that we might trip up `torch._inductor.config.triton.slow_path_cudagraph_asserts`, but if desired those could be follow-ups.",https://github.com/pytorch/pytorch/pull/185057#pullrequestreview-4479376455,ddfdb04f1c0fa9936971d3d57ebec01beda7edf8e8a726db49d944d2be2a13a1 review guidance,pr,185057,issue,167026,high,pr.reviews[5].body,"Getting MemPool accurate is required for things like symmetric memory or you will have a hard error. How are we validating we are keeping this correct across pattern matcher, and graph transformations, such as CSE ? What is preventing CSE of two tensors computed across separate mempools such that...",https://github.com/pytorch/pytorch/pull/185057#pullrequestreview-4487453107,5de0edb87c7cd01f64c96eb676fe984e6a8ccab861a8a2436142099d2c5b8369 review guidance,pr,185057,pr,185057,high,pr.reviews[5].body,"Getting MemPool accurate is required for things like symmetric memory or you will have a hard error. How are we validating we are keeping this correct across pattern matcher, and graph transformations, such as CSE ? What is preventing CSE of two tensors computed across separate mempools such that...",https://github.com/pytorch/pytorch/pull/185057#pullrequestreview-4487453107,235c63ec4e9b981655b4026fae01188f462e905d0ce5b70f3121d76fe58b977d review guidance,pr,185057,issue,167026,high,pr.reviews[11].body,"Looks like I don't have permissions to resolve threads, so I 👍 'd my comments that looked done to me. Thanks for the follow-ups! ",https://github.com/pytorch/pytorch/pull/185057#pullrequestreview-4500232211,8e78d7b57f41178bb9b35bb327cd1cf34b2f1d623cb334d4f2b769f097237bab review guidance,pr,185057,pr,185057,high,pr.reviews[11].body,"Looks like I don't have permissions to resolve threads, so I 👍 'd my comments that looked done to me. Thanks for the follow-ups! ",https://github.com/pytorch/pytorch/pull/185057#pullrequestreview-4500232211,64e578d302c2b20548e44b515bc8149a97071f3b82df96441c3b1c36dcb947a8 references,pr,188766,pr,186282,medium,pr.body,"Stack from ghstack (oldest at bottom): #186282 -> #188766 Without the fix, TorchTitan graph_trainer + FlexCP with load balancing will trigger the following errors: Without squeeze(0) fix",https://github.com/pytorch/pytorch/pull/188766,1e66e8f44c71f1ad6bda1e9e1f5b0d8579064e6a68215ae428c653eb41003ee1 closes,pr,185461,issue,156191,high,pr.body,ing.info. A prior stale PR (#158439) returned None for len on the logger object; this patch handles the actual handlers list instead. Fixes #156191 Generated by my agent Test Plan: python test/dynamo/test_reorder_logs.py IgnoreLogsTests.test_ignore_module_level_logger_ignore_m...,https://github.com/pytorch/pytorch/pull/185461,46e00ddc947c3633779798467e5798785395f15487870f7011c1a1d42dbbf589 closes,pr,188948,issue,153410,high,pr.closingIssuesReferences,pr #188948 declares a closing reference to issue #153410.,https://github.com/pytorch/pytorch/pull/188948,77a890cae4f7d0cb0a4dcb1cfa622daef7e79850b114d3ee3b78d4ca8e51a32b closes,pr,188948,issue,153410,high,pr.body,"Fixes #153410 This change provides a reference implementation for triangular_solve on sparse CPU tensors when MKL is not available, by materializing diag",https://github.com/pytorch/pytorch/pull/188948,5436a961d9bebd0782f621762e47ef904c9cd305c49bfb493b3e223293ca0969 closes,pr,185462,issue,140884,high,pr.body,"ing torch-function dispatch for every ndarray call, since real torch_function overrides should continue to participate. Fixes #156162 Fixes #140884 Generated by my agent Test Plan: python test/dynamo/test_misc.py -k test_numpy_operator_with_default_device_context python test/d...",https://github.com/pytorch/pytorch/pull/185462,5dfde6fec2f2343595c8360389d7ce12079d72fa2c87e376c0b490f4c2cb0b9a closes,pr,185462,issue,156162,high,pr.body,"avoided bypassing torch-function dispatch for every ndarray call, since real torch_function overrides should continue to participate. Fixes #156162 Fixes #140884 Generated by my agent Test Plan: python test/dynamo/test_misc.py -k test_numpy_operator_with_default_device_context...",https://github.com/pytorch/pytorch/pull/185462,9f81a58fae9ebed0ecd9e3499bbe44d11b5237289d1ce68e8bd107db04444a69 closes,pr,185466,issue,156127,high,pr.body,"ead of special-casing Tensor.item() or individual graph-break sites, because the bug is in how captured stack positions are rendered. Fixes #156127 Generated by my agent Test Plan: python test/dynamo/test_error_messages.py ErrorMessagesTest.test_from_user_code_traceback_carets...",https://github.com/pytorch/pytorch/pull/185466,db25b1dc946880c2d9ef5fbc6202a2ff687ea22a31af6992641b659784e28113 closes,pr,185467,issue,156059,high,pr.body,"Apply the same keepalive reconstruction to DataPtrVariable, which represents the same class of raw pointer value across graph breaks. Fixes #156059 Generated by my agent Test Plan: python test/dynamo/test_misc.py MiscTests.test_storage_cdata_use_count_graph_break_lifetime lint...",https://github.com/pytorch/pytorch/pull/185467,7091a2cff1feca6cd3f274a27063b2a03978da3bbba156b8ea79b477c9dec6a3 closes,pr,185471,issue,155800,high,pr.body,"h.broadcast_shapes because that API rejects negative dimensions that _infer_size can propagate when paired with singleton dimensions. Fixes #155800 Generated by my agent Benchmark Results: Before this fix, the issue repro failed during Dynamo fake execution of torch._C._infer_...",https://github.com/pytorch/pytorch/pull/185471,8d733cfe07ae6b37edd379672fbab3410fd5bac2b1fddeedaba32c3dd0ddf19d closes,pr,184083,issue,155584,high,pr.body,g Inductor Triton launches. Add focused tests for preserving custom allocators and installing the fallback for the default allocator. Fixes #155584 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/184083,6d8412d0cea64f759c79f421b03d3fc8f0ff2bae362bf24484d295945e6c9099 closes,pr,185474,issue,155421,high,pr.body,opping only the optional min/max sequence length cache from tangents; offsets remain available to recompute those values when needed. Fixes #155421 Generated by my agent Test Plan: python test/test_nestedtensor.py -k test_compile_sdpa_backward_uncached_metadata_cache PASS. Ran...,https://github.com/pytorch/pytorch/pull/185474,8a2ff30215392ea544025a889b6c9f74b370c82649bb2acb6995a79fa0d6a8ff closes,pr,185483,issue,155238,high,pr.body,"at would only avoid this call site. The operator registration is the root cause, so the fix belongs at the aten/fake-tensor boundary. Fixes #155238 Generated by my agent Test Plan: ninja -C build torch_python ninja -C build install python test/dynamo/test_repros.py -k pad_pack...",https://github.com/pytorch/pytorch/pull/185483,0661b3b25fdcbffedcc1d09736de8c60b04fd0b0612bd52bddb7620f727f0e25 review guidance,pr,185207,pr,185207,high,pr.reviews[0].body,"Thanks! one Q for you: should we switch to ""accelerator"" APIs throughout instead of ""GPU"" (i.e. HAS_GPU / GPU_TYPE)? I realize that there are certain accelerators not expected to work today (e.g. MPS) but in the interest of making these fully available to out-of-tree backends, I think this would...",https://github.com/pytorch/pytorch/pull/185207,8c156e5b68a89e97107c0bc13f2aad0e22e5eadf8141621270bd7a919a96f0dd review guidance,pr,185207,pr,185207,high,pr.reviews[1].body,"(Claude Review) [cleanup] The PR description's ""Changes"" section still describes the original GPU_TYPE/HAS_GPU approach that was abandoned after review feedback: it lists device=""cuda"" -> device=GPU_TYPE, torch.cuda.is_available() gates -> HAS_GPU, torch.device(""cuda:0"") -> torch.device(f""{GPU_TY...",https://github.com/pytorch/pytorch/pull/185207,7b4aabd4abe9bcbfc941abbe116f320063e3570df846a6eaf8a60676bfe3348a closes,pr,184016,issue,166009,high,pr.body,016 Preserve real mutations into fallback alias outputs through lowering and scheduler DCE so returned base buffers observe the copy. Fixes #166009 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/184016,13633652eb640b54c6c2a8c24166941b771c5c484a54275209ad6ccba2ffa437 references,pr,184016,issue,166009,medium,pr.comments[5].body,priate for true isolation testing of DCE and matches the file's existing style. - `test_dce_fallback_alias_mutation` is a faithful repro of #166009. Good that it covers both the return-the-base and consume-later cases. ### Verdict LGTM. The core fix is correct and well-targete...,https://github.com/pytorch/pytorch/pull/184016,2c08199c4b3b3ea972458e27d8df6d4101c42da7f06e05a8a8c787ac5a4e285c review guidance,pr,184016,issue,166009,high,pr.reviews[7].body,CI is all red. I think this needs a rebase,https://github.com/pytorch/pytorch/pull/184016#pullrequestreview-4516429208,19cbeab96243f8207d6c1eeec9f7e3d0ec80ec66b4a8a4abf8bc4c1e433098a5 review guidance,pr,184016,pr,184016,high,pr.reviews[7].body,CI is all red. I think this needs a rebase,https://github.com/pytorch/pytorch/pull/184016#pullrequestreview-4516429208,86161ecb02fefbe5a1b7dc5838ea8221ecb106c883c6dd2671f72f39f89b8941 closes,pr,184086,issue,154807,high,pr.body,Stack from ghstack (oldest at bottom): -> #184086 Fixes #154807 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv @jiayisun,https://github.com/pytorch/pytorch/pull/184086,93c5dc88094485f35466f54177829fb4a99a83cf418beec5700cf1c75d5892c0 review guidance,pr,184086,issue,154807,high,pr.reviews[0].body,"cc @drisspg - do we still generate scaled_mm, or is that deprecated?",https://github.com/pytorch/pytorch/pull/184086,9c842310159fac8a33544e21470e32eda866c1533c33638c4955ea62fb92f0f8 review guidance,pr,184086,pr,184086,high,pr.reviews[0].body,"cc @drisspg - do we still generate scaled_mm, or is that deprecated?",https://github.com/pytorch/pytorch/pull/184086,5ebc8d40ccda9b3a653dee56bb84b6275ede5c91ac82d0dead9ba22655aacf7c closes,pr,185499,issue,154647,high,pr.body,d not remove nonzero_memo or skip the assertion because those would hide the missing fresh-binding invariant rather than maintain it. Fixes #154647 Generated by my agent Test Plan: Issue repro before fix: failed in ep.run_decompositions(...) with AssertionError: u2 -> u0 from...,https://github.com/pytorch/pytorch/pull/185499,481344a9b1cff13b7a0ca9f6ff817250143afdbbd35074080502c92aba613614 review guidance,pr,185499,issue,154647,high,pr.reviews[0].body,This is a misuse of the original intent of the epoch system. I'm not sure what an appropriate alternate fix is.,https://github.com/pytorch/pytorch/pull/185499,00bb2e7c3d72f5e19d9eb0296d0aff767b04320cb9c2f4a8ac31b73ae93905b3 review guidance,pr,185499,pr,185499,high,pr.reviews[0].body,This is a misuse of the original intent of the epoch system. I'm not sure what an appropriate alternate fix is.,https://github.com/pytorch/pytorch/pull/185499,4875df6a9c242a27fcc5f06fbd8b5f079fb428239227dbdf816e50e7f99e4866 closes,pr,184090,issue,154592,high,pr.body,"o avoid persisting user data. The mode bypasses FX graph cache, skips overlapping tensors, and compacts serialized CPU view payloads. Fixes #154592 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...",https://github.com/pytorch/pytorch/pull/184090,18427a4bc5c730f90f4750eab75cbf4d522258cb7ff470773d754f77ab4f3a75 review guidance,pr,184090,issue,154592,high,pr.reviews[0].body,"how is this different from TORCH_COMPILE_DEBUG_SAVE_REAL , can we just reuse ? cc @xmfan original issue author",https://github.com/pytorch/pytorch/pull/184090,c4ca3ae7308993f27bb920b9e87767874769c28fe6aa09e0c2a287093bb8bb02 review guidance,pr,184090,pr,184090,high,pr.reviews[0].body,"how is this different from TORCH_COMPILE_DEBUG_SAVE_REAL , can we just reuse ? cc @xmfan original issue author",https://github.com/pytorch/pytorch/pull/184090,7754963a73acc8169edd3cfc1ecb2eed5698463e2d1524d84186dc6af75be8b0 closes,pr,185504,issue,154559,high,pr.body,"tensors returned by both torch.cond branches, and removes the expected-failure marker from the existing unbacked SymInt closure test. Fixes #154559 Generated by my agent Test Plan: python test/export/test_export.py TestExport.test_cond_unbacked_symint_size_closure_non_strict -...",https://github.com/pytorch/pytorch/pull/185504,5cb3095072483c9885c610ae5e07d11a4528cf4f3df7a01cd4533f41203b9285 closes,pr,185508,issue,154454,high,pr.body,"his keeps reraises attributed to the original observed operation, so duplicate graph-break suppression sees the same source location. Fixes #154454 Generated by my agent Test Plan: python test/dynamo/test_exc.py -k test_reraised_observed_exception_graph_break_log python test/d...",https://github.com/pytorch/pytorch/pull/185508,fd8f6e40b49f7220be462df6fcca76ef2e29adbfa86263b50f7d97095a423cce closes,pr,184091,issue,154306,high,pr.body,fter-graph liveness mismatches as a re-recordable invariant failure and restore cleared managed inputs before recording the new path. Fixes #154306 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/184091,8a11482c1e6733b0f639e61917ddf74b767b7d069e40ecfe5752b45ecc0a1856 review guidance,pr,184091,issue,154306,high,pr.reviews[1].body,needs to be done in a way without regressing memory,https://github.com/pytorch/pytorch/pull/184091,65b2df5999ef3ac45914b2f359cd4655dfa116dce4185597f589dab3c7a5ef7f review guidance,pr,184091,pr,184091,high,pr.reviews[1].body,needs to be done in a way without regressing memory,https://github.com/pytorch/pytorch/pull/184091,b030c7d5ac3c158d3b1b6d49502123f659d75161534b5d2d9fead30363a184b5 closes,pr,184092,issue,154301,high,pr.body,ax-autotune so single-kernel pointwise graphs avoid replay input-copy overhead while still enabling cudagraphs for larger partitions. Fixes #154301 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/184092,63f8921520a62285bd14249d6c0a3c6a31bbce0dd662bb9e4f43b1fbf0274eec review guidance,pr,184092,issue,154301,high,pr.reviews[1].body,"why patch the config, instead of update the config value, so now it will never be read ?",https://github.com/pytorch/pytorch/pull/184092,a937128a57d9aaaad649048c5eae798f1e1345ee7fa45af650e98cd6e132c9b5 review guidance,pr,184092,pr,184092,high,pr.reviews[1].body,"why patch the config, instead of update the config value, so now it will never be read ?",https://github.com/pytorch/pytorch/pull/184092,9982cbf380a1e5085ef2ef505071b9846154be83fcaf97be2ec5fa41d1ad98b3 review guidance,pr,184092,issue,154301,high,pr.reviews[3].body,"why patch the config, instead of update the config value, so now it will never be read ?",https://github.com/pytorch/pytorch/pull/184092#pullrequestreview-4349027308,92bd775b057a1492ee2b9fb2377562e926363402360cad18bf548b5c3b668833 review guidance,pr,184092,pr,184092,high,pr.reviews[3].body,"why patch the config, instead of update the config value, so now it will never be read ?",https://github.com/pytorch/pytorch/pull/184092#pullrequestreview-4349027308,ac9473d132b3a683889511153073d2b8b325ec7092dae54238937a6387dc2287 closes,pr,185528,issue,154259,high,pr.body,og-graph-breaks --output /home/jansel/local/pytorch-issue-fixer/state/154259/vision_maskrcnn_after.csv git diff --check lintrunner -a Fixes #154259 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/185528,67a08beb214a1fa6b56d8543dac28cfc277b77a1dc376d428859c0fbef065d0a closes,pr,185536,issue,154161,high,pr.body,"Stack from ghstack (oldest at bottom): -> #185536 Issue #154161 exposed a PGO invariant mismatch. record_automatic_dynamic stored tensor sizes and strides for automatic dynamic shape decisions, but spars",https://github.com/pytorch/pytorch/pull/185536,ed233bce0e2f646f4874831ca2cf3254f6113bfbde35b8342a4a15ec1c1b6cae closes,pr,185545,issue,154137,high,pr.body,ed_slice python test/export/test_export.py TestExport.test_unbacked_slice_forward TestExport.test_unbacked_slice_simple lintrunner -a Fixes #154137 Generated by my agent,https://github.com/pytorch/pytorch/pull/185545,ee517d90d01e7fecaa7e6ad6e037c0feb0d32560d4aca09714551da9b2f5a8bd closes,pr,185564,issue,153387,high,pr.body,sted_subclass_constructed_in_forward_pre_dispatch TestExport.test_subclass_nested_attr_access lintrunner -a git diff --cached --check Fixes #153387 Generated by my agent,https://github.com/pytorch/pytorch/pull/185564,b91505ed71c7d2662361c9973c31d7fae41fe6a671bfef38844de87866057b79 closes,pr,185568,issue,153247,high,pr.body,port/fake/runtime-assert tracing behavior. A broader redesign of every data-dependent arange output shape form is outside this issue. Fixes #153247 Generated by my agent Test Plan: python test/export/test_export.py TestExport.test_export_arange_data_dependent_float_step TestEx...,https://github.com/pytorch/pytorch/pull/185568,1b101851a6383bc6707a1645b94d8da4634c511e13443f601d989caf022f0a6b closes,pr,187956,issue,187945,high,pr.body,"s fixes the metadata gap without blindly copying old output metadata, which would be wrong for shape- or dtype-changing replacements. Fixes #187945 Generated by my agent Benchmark Results: Microbenchmark: export a graph with 50 aten.div.Tensor matches and call replace_pattern...",https://github.com/pytorch/pytorch/pull/187956,d63f2ce6f5a4da45a4df3db0b2cad8bef69b0ebc458817c8e67c3f00a56f8cd0 closes,pr,185569,issue,153227,high,pr.body,"or_size_hint_short_circuits_unbacked_scalar TestUnbacked.test_deferred_sym_or_assert_backend_eager Original default-backend reproducer from #153227 returned tensor([100, 200, 300, 400, 500]). lintrunner -a torch/fx/experimental/symbolic_shapes.py test/test_dynamic_shapes.py Fi...",https://github.com/pytorch/pytorch/pull/185569,64e532fe4e7d6ecc3a758f88a83f8cda71503e1e241b4315bcdaf94d8516377c closes,pr,185588,issue,153175,high,pr.body,"cked_symint_compile (Ran 1 test, OK) git diff --cached --check lintrunner -a torch/_inductor/graph.py test/inductor/test_cpu_repro.py Fixes #153175 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...",https://github.com/pytorch/pytorch/pull/185588,1dbeeddf8e4890784a3ccd0fa8463785985c01c81efdeb9d558a66fd31e4ff8b closes,pr,185591,issue,153149,high,pr.body,"arguments, so runtime behavior is unchanged. Earlier PR #154353 documented the eager behavior; this patch fixes the exported IR side. Fixes #153149 Generated by my agent Benchmark Results Command: microbenchmark repeated torch.export.export calls for a 100-op relu chain and a...",https://github.com/pytorch/pytorch/pull/185591,9682dcd4341695052bd0934a4bdb3cbd659f5001f82fa36b5869ecb19fd027e2 closes,pr,185604,issue,152656,high,pr.body,"ng all torch._check lambdas to appear in runtime asserts, which would change many graph strings and broader user-visible diagnostics. Fixes #152656 Generated by my agent Test Plan: python test/export/test_export.py -k test_unbacked_infer_size python test/test_dynamic_shapes.py...",https://github.com/pytorch/pytorch/pull/185604,f1234d1da0e0ba170a194b8d2e03667b26c6a75c04ab51fde9eec9bc72a355e8 closes,pr,185690,issue,152346,high,pr.body,Stack from ghstack (oldest at bottom): -> #185690 Fixes #152346 AOTAutograd functionalization represents input view mutations by building an updated full input and then copying that full tensor back into,https://github.com/pytorch/pytorch/pull/185690,9afaf0b430d7e7e5131805e7b31346a2a6e5087f86779066ae9f315a14f32af6 review guidance,pr,185690,issue,152346,high,pr.reviews[0].body,I actually couldn't verify the original issue. In inductor generated code we just write to 1024 elements instead of full tensor. Would you investigate what fixed it ?,https://github.com/pytorch/pytorch/pull/185690,f1992b1d4643487789e9f747a443da408b711822639f5154ad6e1ddadbf7da28 review guidance,pr,185690,pr,185690,high,pr.reviews[0].body,I actually couldn't verify the original issue. In inductor generated code we just write to 1024 elements instead of full tensor. Would you investigate what fixed it ?,https://github.com/pytorch/pytorch/pull/185690,e6b3e6bf6125a1aef5a5b2794ad3666b7618b496074578289c228554d196224e closes,pr,185624,issue,152343,high,pr.body,ic_with_subclass_desugaring python test/dynamo/test_subclasses.py -k test_subclass_parameters_are_static_under_training lintrunner -a Fixes #152343 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/185624,cbcd5a2bc7d2d76be4abe5f8082504a1cf2990ce454895cd1e5149341fa1e28e review guidance,pr,185624,issue,152343,high,pr.reviews[0].body,deferring to @aorenste and @laithsakka,https://github.com/pytorch/pytorch/pull/185624,7c206a0d375eb23beab054e6a3984523168bdc87c6ea8f9ad791f2d53714cbea review guidance,pr,185624,pr,185624,high,pr.reviews[0].body,deferring to @aorenste and @laithsakka,https://github.com/pytorch/pytorch/pull/185624,7c9dd719f9dbdae4e1bf3d016fd8c108ab29c4a39e0f942f1aa68f9021310c0c closes,pr,185632,issue,152183,high,pr.body,gging path. SubclassMeta is covered because it is logged separately from ViewAndMutationMeta via dataclass_repr(maybe_subclass_meta). Fixes #152183 Generated by my agent Test Plan: python test/functorch/test_subclass_codegen.py TestSubclassCodegen.test_view_and_mutation_meta_r...,https://github.com/pytorch/pytorch/pull/185632,1c2e7133ff0e101573346086a56d5ccc8a21e632fd6fce42c1e9c8e988749643 closes,pr,184109,issue,151649,high,pr.body,"ent patterns traced through view to match post-grad graphs rewritten to reshape, while preserving exact matching for manual patterns. Fixes #151649 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...",https://github.com/pytorch/pytorch/pull/184109,f4e45efde7082c706eab02273cedd686e0500854845549a315ec990d5008fa59 review guidance,pr,184109,issue,151649,high,pr.reviews[0].body,"Hmm, when do we even insert the reshapes ? I remember it had something to do with channels last conversion. Ideally the post grad IR is supposed to be exact aliasing. Can you investigate why we insert them in the first place ? An alternative would be to use clone + view as needed. Assuming we're...",https://github.com/pytorch/pytorch/pull/184109,3f28ceb29edbc1228154fe412b2f84fe4d44a9ccc58a0e4db52db7389fd5177b review guidance,pr,184109,pr,184109,high,pr.reviews[0].body,"Hmm, when do we even insert the reshapes ? I remember it had something to do with channels last conversion. Ideally the post grad IR is supposed to be exact aliasing. Can you investigate why we insert them in the first place ? An alternative would be to use clone + view as needed. Assuming we're...",https://github.com/pytorch/pytorch/pull/184109,ed8acb4ef31ceec512c36511c87141241b01595089fee981c0fc81121b44a375 review guidance,pr,184109,issue,151649,high,pr.reviews[1].body,"I think we should avoid converting view to reshape. There should be some invariant that all ops in the inductor graph are just post-grad aten, and reshape is not post-grad-aten.",https://github.com/pytorch/pytorch/pull/184109,d269642b164f5efe45e85ae345b21ad80f140d60a48511a7029ca185dcedb02a review guidance,pr,184109,pr,184109,high,pr.reviews[1].body,"I think we should avoid converting view to reshape. There should be some invariant that all ops in the inductor graph are just post-grad aten, and reshape is not post-grad-aten.",https://github.com/pytorch/pytorch/pull/184109,942ed254984847eacc9b4df4537b96dc9c3a95794d17fbe04d89c5272030bec2 closes,pr,185718,issue,151196,high,pr.body,a GradTrackingTensor around a FakeTensor; discovering the inner fake tensor through the functorch wrapper resolves that case as well. Fixes #151196 Fixes #172428 Generated by my agent Test Plan: python -m py_compile torch/_subclasses/fake_tensor.py ninja -C build torch_cpu man...,https://github.com/pytorch/pytorch/pull/185718,373cca74b96d8a243650191c922069f0affaa7461abea76153a335bbeff34299 closes,pr,185718,issue,172428,high,pr.body,Tensor around a FakeTensor; discovering the inner fake tensor through the functorch wrapper resolves that case as well. Fixes #151196 Fixes #172428 Generated by my agent Test Plan: python -m py_compile torch/_subclasses/fake_tensor.py ninja -C build torch_cpu manual vmap(jacfw...,https://github.com/pytorch/pytorch/pull/185718,01ac989cd87267dbd7d394bc68fbd734367506e152091df16392cda6764b659c closes,pr,188114,issue,188113,high,pr.closingIssuesReferences,pr #188114 declares a closing reference to issue #188113.,https://github.com/pytorch/pytorch/pull/188114,bca08deb97bd87b12d5bdc68c90b36671191d234be1f8f9c4f97dff62b49f22f closes,pr,188114,issue,188113,high,pr.body,"x950. Added gfx1200/gfx1201. Tested on gfx1201 (RX 9070): WITHOUT patch the kernel ASSERTs, WITH patch the kernel executes correctly. Fixes #188113 cc @jeffdaily @sunway513 @jithunnair-amd @pruthvistony @ROCmSupport @jataylo @hongxiayang @naromero77amd @pragupta @jerrymannil @...",https://github.com/pytorch/pytorch/pull/188114,d47adfc584d9c7a2259db42725a1cbac9c0a9d730ea3f063a91534d7d6ba71c4 review guidance,pr,188114,pr,144777,high,pr.reviews[0].body,"cc @jataylo @drisspg I guess we dont have gfx1200, gfx1201 on ci ?",https://github.com/pytorch/pytorch/pull/188114,93b32bd731cc2c83cf46cc85b528c557389d039462e1cfbbdfb9dde8966e66df review guidance,pr,188114,pr,155103,high,pr.reviews[0].body,"cc @jataylo @drisspg I guess we dont have gfx1200, gfx1201 on ci ?",https://github.com/pytorch/pytorch/pull/188114,32508c3c2889d59360bd43464819056dac6d5337296edcc09a43b5bc284e6519 review guidance,pr,188114,pr,187267,high,pr.reviews[0].body,"cc @jataylo @drisspg I guess we dont have gfx1200, gfx1201 on ci ?",https://github.com/pytorch/pytorch/pull/188114,10c0281a3d81c0607a1510c29db7b363096d2a916f1b510eff36e9df8b2a1684 review guidance,pr,188114,issue,188113,high,pr.reviews[0].body,"cc @jataylo @drisspg I guess we dont have gfx1200, gfx1201 on ci ?",https://github.com/pytorch/pytorch/pull/188114,e9e9305329a33a9dce773126f95e5f2016a0a1bb384bc4857b6c311ed109d34c review guidance,pr,188114,pr,144777,high,pr.reviews[1].body,"@poad42 Unfortunately, as it stands right now, we cannot add any new architecture support for CK by default. Building for multiple architectures puts too much strain on the PyTorch CI. If you want this to go through we ask that you implement it as an opt-in feature gated by an environment variabl...",https://github.com/pytorch/pytorch/pull/188114,17d3134a3299dc1e7a583b1afadc44e980c76bfb8de03fe3c97bf81c6690f815 review guidance,pr,188114,pr,155103,high,pr.reviews[1].body,"@poad42 Unfortunately, as it stands right now, we cannot add any new architecture support for CK by default. Building for multiple architectures puts too much strain on the PyTorch CI. If you want this to go through we ask that you implement it as an opt-in feature gated by an environment variabl...",https://github.com/pytorch/pytorch/pull/188114,544796e95d5d4ff716ca4a7a17b1c774ee59d4bda50cb5020be9f17757217d73 review guidance,pr,188114,pr,187267,high,pr.reviews[1].body,"@poad42 Unfortunately, as it stands right now, we cannot add any new architecture support for CK by default. Building for multiple architectures puts too much strain on the PyTorch CI. If you want this to go through we ask that you implement it as an opt-in feature gated by an environment variabl...",https://github.com/pytorch/pytorch/pull/188114,c4c6f09dfdfae52064f0b924bbf7acffdfd5a0dce6a453b43744ef31e7c78fba review guidance,pr,188114,issue,188113,high,pr.reviews[1].body,"@poad42 Unfortunately, as it stands right now, we cannot add any new architecture support for CK by default. Building for multiple architectures puts too much strain on the PyTorch CI. If you want this to go through we ask that you implement it as an opt-in feature gated by an environment variabl...",https://github.com/pytorch/pytorch/pull/188114,a89c1d321a0779ca0d7903df1bfe920b7b5f5affdd82898786eb4b533298288e review guidance,pr,188237,pr,188237,high,pr.reviews[0].body,"@zhengjun-xing This looks good to me, but it would be nice to see some benchmarks to validate the change is generally a performance improvement (or neutral)",https://github.com/pytorch/pytorch/pull/188237,18677eb9c39666bf5cf0bd2f8928acbf46c105b0d66a4bd5e73a77453681d969 closes,pr,185755,issue,149857,high,pr.body,"rce_issue stayed on fmha_cutlassB_f32_aligned_64x64_k64_sm80; forcing 128x64_k65536 was slower (38.784 ms final vs 48.749 ms forced). Fixes #149857 Generated by my agent Test Plan: ninja -C build torch_python -j 16 LD_PRELOAD=""$PRELOAD"" python test/test_transformers.py -k test...",https://github.com/pytorch/pytorch/pull/185755,7ba00ed0236c9617543db44fef33f8758d194f090153cc2dad6c512fe1abf36a closes,pr,185722,issue,150915,high,pr.body,through torch.compile. Both are broader behavior changes; this keeps the fix targeted and reuses the existing wrap_inline mechanism. Fixes #150915 Generated by my agent Benchmark Results: Microbenchmark command: 30 iterations of ToyModel().compile(backend=CompileCounter()); mo...,https://github.com/pytorch/pytorch/pull/185722,7f933ba7c7fa2bb38081bab2f771b31d354d7c8a10f6e214de2992b6e2dde5cf closes,pr,184119,issue,150621,high,pr.body,logue-fusion selection because the consumer may still be a MultiTemplateBuffer before autotuning selects a concrete triton_mm choice. Fixes #150621 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/184119,e44fc61dc01d39a85e695f760acd01730572c495f1c78e2df23b21e7623b43fe review guidance,pr,184119,issue,150621,high,pr.reviews[0].body,"I still dont understand the root cause here. Wouldnt this same instruction be generated with and without the matmul. In which case, it's not a template issue ?",https://github.com/pytorch/pytorch/pull/184119,7b5e8b8dd6dc72cfe7558d4917860d9d106672d2c8154a209e25880543ee1db1 review guidance,pr,184119,pr,184119,high,pr.reviews[0].body,"I still dont understand the root cause here. Wouldnt this same instruction be generated with and without the matmul. In which case, it's not a template issue ?",https://github.com/pytorch/pytorch/pull/184119,74f713871c5a03ba5539ce2ebaad98363ba83c25595e126b20a3833ba7189109 review guidance,pr,184119,issue,150621,high,pr.reviews[1].body,"From original issue poster But when switching back to the working nightly version of PyTorch, all Triton versions work again (including one I just compiled from the tip of the Triton repo just now). deferring on this issue",https://github.com/pytorch/pytorch/pull/184119,7240d98bfc89dd81b0edc3602353b76ec452133a45bd53a1c9659eec8a153bf5 review guidance,pr,184119,pr,184119,high,pr.reviews[1].body,"From original issue poster But when switching back to the working nightly version of PyTorch, all Triton versions work again (including one I just compiled from the tip of the Triton repo just now). deferring on this issue",https://github.com/pytorch/pytorch/pull/184119,3373f142be73be22dffeaefa3aa51f1183895a5994cbd6f82d9c091c2599ce43 review guidance,pr,184119,issue,150621,high,pr.reviews[2].body,"I still dont understand the root cause here. Wouldnt this same instruction be generated with and without the matmul. In which case, it's not a template issue ?",https://github.com/pytorch/pytorch/pull/184119#pullrequestreview-4339755535,5b1b0da752288ee0bf73f30efbcba0a8cc5c1b0a9b5b7091e9c5c1c19ce8ff99 review guidance,pr,184119,pr,184119,high,pr.reviews[2].body,"I still dont understand the root cause here. Wouldnt this same instruction be generated with and without the matmul. In which case, it's not a template issue ?",https://github.com/pytorch/pytorch/pull/184119#pullrequestreview-4339755535,9ee1cb910c14594405e2059c5bc7ffa48d0eb43e654e8098f0ab567349fa1c3f review guidance,pr,184119,issue,150621,high,pr.reviews[3].body,"From original issue poster > But when switching back to the working nightly version of PyTorch, all Triton versions work again (including one I just compiled from the tip of the Triton repo just now). deferring on this issue",https://github.com/pytorch/pytorch/pull/184119#pullrequestreview-4377344204,2be8277b8e7e47ecf183ca26e792e9163d2264d037ccfd67cd1d09ee87138f95 review guidance,pr,184119,pr,184119,high,pr.reviews[3].body,"From original issue poster > But when switching back to the working nightly version of PyTorch, all Triton versions work again (including one I just compiled from the tip of the Triton repo just now). deferring on this issue",https://github.com/pytorch/pytorch/pull/184119#pullrequestreview-4377344204,3ec92cf52e7773a2d985dd52cd40a53f9451b19f34b320bd4566c229a1505ef6 review guidance,pr,184119,issue,150621,high,pr.reviews[6].body,"note that for any prologue fusion, we'll compile and skip if there is an error in triton compilation. so its not clear to me if this is worth the extra complication.",https://github.com/pytorch/pytorch/pull/184119#pullrequestreview-4403070958,e39cb49fa2adae7a3631ab00d80acd03519ca6c32981ce828c33eca723a65ed4 review guidance,pr,184119,pr,184119,high,pr.reviews[6].body,"note that for any prologue fusion, we'll compile and skip if there is an error in triton compilation. so its not clear to me if this is worth the extra complication.",https://github.com/pytorch/pytorch/pull/184119#pullrequestreview-4403070958,4e290440e60665c92814f253cd7defcafbc7f26efde0b61d9ac58964fe5e3bab closes,pr,185724,issue,150613,high,pr.body,"son as the existing scalar bucketize variants: the 0-d scalar extern path has no dynamic loop variable for that harness to assert on. Fixes #150613 Generated by my agent Test Plan: python repro for issue #150613: passed, printed tensor(2) python repro for analogous scalar buck...",https://github.com/pytorch/pytorch/pull/185724,084b635eb6b945e899933251ccc4f65d3d6ad0df4c8a0032cdb4c2516af7bb41 closes,pr,185726,issue,150465,high,pr.body,ace function with a non-graphable dict_items input and asserts that the error includes the test callsite and omits external_utils.py. Fixes #150465 Generated by my agent Test Plan: python test/dynamo/test_decorators.py DecoratorTests.test_nonstrict_trace_direct_compile_invalid...,https://github.com/pytorch/pytorch/pull/185726,8736c132e682f4664c2ed74342bcba25113bd8b9d37b8f1e65a1675cefed34ff references,pr,188951,issue,188319,medium,pr.body,"Implements the changes proposed in #188319. Changes generate_binary_build_matrix.py: add ROCM_ARCHES_FULL_VERSION mapping (7.1→7.1.1, 7.2→7.2.3) generate_docker_release_matrix.py: ad",https://github.com/pytorch/pytorch/pull/188951,7274a3c665bda3d0c50b40bfe588d4a406c9f9b256e39071fa13336988e34c90 closes,pr,185731,issue,150319,high,pr.body,he developer error after we had already carried invalid metadata forward; the prefix copy is where the stale reference is introduced. Fixes #150319 Generated by my agent Benchmark Results: A simple torch.compile eager-backend compile-time microbenchmark was run before and afte...,https://github.com/pytorch/pytorch/pull/185731,7fab83b425bf51862340d0590b5d2ecdc263417cc8ca63fbb8e9810767f98b46 closes,pr,185732,issue,150262,high,pr.body,exempts those paths while preserving eager execution for the tensor-subclass fake/functional propagation path that triggered the bug. Fixes #150262 Generated by my agent Test Plan: python test/distributed/tensor/test_dtensor_compile.py TestDTensorCompile.test_aot_autograd_over...,https://github.com/pytorch/pytorch/pull/185732,9ca7fccf707e8377709f70e91e1f1e9e3a414b62628b98f6ccd2c89f1c97fb98 closes,pr,185734,issue,150063,high,pr.body,le before the cached ShapeEnv boundary was also rejected because it regressed torch._check laziness and unhashable callable messages. Fixes #150063 Generated by my agent Test Plan: python test/test_dynamic_shapes.py TestPySymInt.test_expect_true_message_replays TestPySymInt.te...,https://github.com/pytorch/pytorch/pull/185734,dc3d8f5beeeff245ff02035f47d1d1879b3521e4279b199753366c89129418f6 closes,pr,185740,issue,149968,high,pr.body,signature -v python test/dynamo/test_misc.py -k inspect_variable_redirect -v lintrunner -a git diff --check git diff --cached --check Fixes #149968 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/185740,e77287ba6f43a9fdf676798475264eba421ab862e3594401cf6f951fc24d85c9 closes,pr,185742,issue,149963,high,pr.body,python -m pytest test/dynamo/test_flat_apply.py -q python -m pytest test/dynamo/test_decorators.py -q git diff --check lintrunner -a Fixes #149963 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/185742,6372897a16f4a0f2dd1fb1129de4330b5d1ce97b5f9864a643a1d070819769a8 closes,pr,185730,issue,150022,high,pr.body,the dynamic_shapes tree and the traced input tree had different structures; the trace-order remap fixes that path too. Fixes #150371 Fixes #150022 Generated by my agent Test Plan: python - <<'PY' ... PY issue repro for strict=False and strict=True: both exported successfully w...,https://github.com/pytorch/pytorch/pull/185730,1280b61e27f92076de65aa5eca8370a1a47db6fb0cc1a66e3b0833aa3669923d closes,pr,185730,issue,150371,high,pr.body,"match, because the dynamic_shapes tree and the traced input tree had different structures; the trace-order remap fixes that path too. Fixes #150371 Fixes #150022 Generated by my agent Test Plan: python - <<'PY' ... PY issue repro for strict=False and strict=True: both exported...",https://github.com/pytorch/pytorch/pull/185730,4af4ddfd7b54860737a4e6a88da999edb071902db60c8c7329ca1f2c15651f35 review guidance,pr,185730,issue,150022,high,pr.reviews[0].body,"(Reviewed by me, assisted by AI)",https://github.com/pytorch/pytorch/pull/185730,a35c76dd1027595a3e51e7f7eeb31aeaf3072bcee23de5691a56dc145f400667 review guidance,pr,185730,issue,150371,high,pr.reviews[0].body,"(Reviewed by me, assisted by AI)",https://github.com/pytorch/pytorch/pull/185730,f87201261cc221e70f25e96b7de7dc2a5589c6ac235e1bee2e3eae4d2c5ea885 review guidance,pr,185730,pr,185730,high,pr.reviews[0].body,"(Reviewed by me, assisted by AI)",https://github.com/pytorch/pytorch/pull/185730,e75348e163b44bd157fdfc9bfcce3efdf14b8961eb08f117090c8a7cd7c0de53 closes,pr,188783,issue,188773,high,pr.closingIssuesReferences,pr #188783 declares a closing reference to issue #188773.,https://github.com/pytorch/pytorch/pull/188783,ea5fad4c409b8c7375471859b5be0f91e5653ca6a03a88f7b54c2b97eca1829e closes,pr,188783,issue,188773,high,pr.body,"Fixes #188773 The _bmm_outer_product_cond function only checked a.is_cuda and b.is_cuda, which returns False for XPU tensors. This caused the triton over",https://github.com/pytorch/pytorch/pull/188783,7e0d9a2f6a12a6427a26e5569ac6a6f26358dba746dcb68b949942f2940c10f1 closes,pr,189043,issue,188892,high,pr.closingIssuesReferences,pr #189043 declares a closing reference to issue #188892.,https://github.com/pytorch/pytorch/pull/189043,f5b6ab80fe1892253cf4f81379d702a448884f7e79992b4947759aec07f8d45b closes,pr,189043,issue,188892,high,pr.body,"Fixes #188892 Problem cuDNN 9 is split into a dispatcher (libcudnn.so) plus a set of engine sub-libraries (libcudnn_graph.so, libcudnn_engines_*.so, libc",https://github.com/pytorch/pytorch/pull/189043,27e6a9fef059c2da0dd1cdcdb6e15a046ee5906d93020d0aad202760f3dad1b2 closes,pr,186039,issue,143308,high,pr.body,"sMapping overloads, including the frame-locals accessor path, so diagnostics follow the same access path as runtime guard evaluation. Fixes #143308 Generated by my agent Test Plan: ninja -C build torch_python -j 8 && cp build/lib/libtorch_python.so torch/lib/libtorch_python.so...",https://github.com/pytorch/pytorch/pull/186039,f21742daf2663ea39bd33c99a8373c5bf55579644394e6f5da189a3647fa224c closes,pr,187452,issue,187451,high,pr.closingIssuesReferences,pr #187452 declares a closing reference to issue #187451.,https://github.com/pytorch/pytorch/pull/187452,79c91b2c460b6f668e7acdbde8789897a4703c06c8bfcf23b5c595a118ccaea2 closes,pr,187452,issue,187451,high,pr.body,Fixes #187451 An example implementation from the RFC on dev-discuss. The goal of this PR is to demonstrate that significant speedups are achievable by ap,https://github.com/pytorch/pytorch/pull/187452,0422aa091c021c2b41ac4185b3275be19d05a3e801747eeda8f51b202e619d9c closes,pr,185902,issue,149909,high,pr.body,"fig.allow_rnn was set manually. That made torch.compile(..., fullgraph=True) fail immediately for these modules, which is the root cause of #149909. Turn the RNN tracing path on by default so these modules can be captured. Enabling that path exposed an LSTM decomposition bug o...",https://github.com/pytorch/pytorch/pull/185902,3c843d4dffdf57a8f4f9721532599001bcf08d0bb0a381f7213f489293fdd365 review guidance,pr,185902,issue,149909,high,pr.reviews[0].body,do you have performance benchmarks for this? against fullgraph=False that would have fallen back to eager?,https://github.com/pytorch/pytorch/pull/185902,353ca41e2d6d2748dbf71dbf24f35b94b41c46b64817572518cc4817b4b76bf9 review guidance,pr,185902,pr,185902,high,pr.reviews[0].body,do you have performance benchmarks for this? against fullgraph=False that would have fallen back to eager?,https://github.com/pytorch/pytorch/pull/185902,1de647d37f8e67815474c61b520ceadfdf613ff88e3fa343c7c09bd3dc00027b closes,pr,185752,issue,149894,high,pr.body,"al and doing prefix translation at _load_from_state_dict matches nn.Module semantics more closely, including nested compiled modules. Fixes #149894 Generated by my agent Test Plan: python test/dynamo/test_modules.py OptimizedModuleTest -v python test/nn/test_load_state_dict.py...",https://github.com/pytorch/pytorch/pull/185752,8426dcaf226bf1b9693ec4d7be2add4e77d3b73cd9f8b788457b6bec1dfc29ca closes,pr,185759,issue,149767,high,pr.body,"owering point lets the guard include sequence length and sparse-block compatibility without broadening the public heuristic contract. Fixes #149767 Generated by my agent Benchmark Results Command recorded in state/149767/notes. On NVIDIA B200, PyTorch 2.13.0a0+gite13b8c7, CUDA...",https://github.com/pytorch/pytorch/pull/185759,a7b9817e6622ba34548dc041e530875b56bc02925aa803637214f78d20926777 references,pr,185759,issue,149767,medium,pr.comments[2].body,"directly, just noting it's an internal contract. No bugs or correctness issues found. The change is well-scoped to the problem described in #149767. ---",https://github.com/pytorch/pytorch/pull/185759,57cc6965133cd57948e3d3e4ba4b5a2a38209df5bbb85f49908b5d611537300e review guidance,pr,185759,issue,149767,high,pr.reviews[0].body,I don't like how this code is structured,https://github.com/pytorch/pytorch/pull/185759,dee6e85a1d46aa86f708eeb6390ef40b9dc7074c2b6b953a0ab7180e221ffd4a review guidance,pr,185759,pr,185759,high,pr.reviews[0].body,I don't like how this code is structured,https://github.com/pytorch/pytorch/pull/185759,751514c25a827e332412e4ce62f6f56d39d9302e9db6b027bb28b654aea083ef review guidance,pr,185759,issue,149767,high,pr.reviews[1].body,I don't like how this code is structured,https://github.com/pytorch/pytorch/pull/185759#pullrequestreview-4410124480,fdc31beb26c471fff3fcf358fae0de67af4557233f2b90c455f32903b94c87a9 review guidance,pr,185759,pr,185759,high,pr.reviews[1].body,I don't like how this code is structured,https://github.com/pytorch/pytorch/pull/185759#pullrequestreview-4410124480,13d2b1ac7d5f98fb7386410d6594e29322eb6dc302bafe04d5d1a24628c7add1 closes,pr,185762,issue,149586,high,pr.body,round the compiled region. Graph deduplication treats those helpers like other global-state enter/exit nodes to keep ordering stable. Fixes #149586 Generated by my agent Test Plan: python test/dynamo/test_ctx_manager.py -k auto_dispatch_below_autograd python test/dynamo/test_g...,https://github.com/pytorch/pytorch/pull/185762,0859b47a72d97725fdbc02d8c0319133c5063a1d03ef477d21a980bc47a5a7d2 closes,pr,185763,issue,149566,high,pr.body,with disable is left unchanged. This keeps the fix local to the diagnostic path instead of weakening skip-file handling more broadly. Fixes #149566 Generated by my agent Test Plan: python test/dynamo/test_error_messages.py ErrorMessagesTest.test_skipfile_dynamo_call ErrorMessa...,https://github.com/pytorch/pytorch/pull/185763,3176e21c89d5b2aaecb491b8fd851ff5f7c9ae4890d256fa96321b8d8d7d812f closes,pr,185766,issue,149556,high,pr.body,ch.jit.isinstance calls instead of changing the general builtin isinstance implementation or the broader torch.jit skiplist behavior. Fixes #149556 Generated by my agent Benchmark Results: Before: the fullgraph torch.jit.isinstance tensor repro failed before execution with tor...,https://github.com/pytorch/pytorch/pull/185766,ba307e21c0457606271cfd70cbe0141108f5307778277c56327cc6b513f0be30 closes,pr,185768,issue,149047,high,pr.body,"dict mutations and RemovableHandle lifetime/removal semantics. This change keeps the fix scoped to the current unsupported behavior. Fixes #149047 Generated by my agent Benchmark Results: Command: temporarily reverse-applied the staged patch, then ran a Python heredoc that com...",https://github.com/pytorch/pytorch/pull/185768,e850f15b64662b61a4b53163bc77957149f675c8022c48613703694a3df9ec59 closes,pr,185769,issue,149037,high,pr.body,missing values and duplicates the VariableTracker handling. A shared helper keeps the CPython StopIteration.value semantics explicit. Fixes #149037 Generated by my agent Test Plan: python test/dynamo/test_generator.py -k generator_return -q python test/dynamo/test_misc.py -k t...,https://github.com/pytorch/pytorch/pull/185769,f8ee214f3210efb834fcf0899453e9a0f1bd8b63508e035147901f12d3472311 closes,pr,185771,issue,149010,high,pr.body,"al Dynamo modeling gap in stdlib string handling while inlining, and the same path is useful for other error-message formatting code. Fixes #149010 Generated by my agent Test Plan: python test/dynamo/test_misc.py -k custom_strip python test/dynamo/test_misc.py -k textwrap_inde...",https://github.com/pytorch/pytorch/pull/185771,d07fe3c3cce426af48fb041de98de84b91a8dd97979a9ba360af248b0f03a707 closes,pr,186469,issue,123374,high,pr.body,s. Test Plan: python test/dynamo/test_bytecode_utils.py -k test_runtime_error python test/dynamo/test_bytecode_utils.py lintrunner -a Fixes #123374 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/186469,e1b5c0436476194f0f2ecce4fe2f102cd9c3dee6173a17390143aeb133e5bf89 closes,pr,185772,issue,148859,high,pr.body,"ding differences everywhere, but it would give up fusion for a common operator; this patch keeps the fix scoped to the reported path. Fixes #148859 Generated by my agent Benchmark Results: Compiled microbenchmark on CPU for input shape (1, 3, 16, 16), output shape (32, 32), to...",https://github.com/pytorch/pytorch/pull/185772,67a2f83a5000fcf83461496ffc428eb6e82715bff512d279eb72a4f7c0ed2311 closes,pr,185773,issue,148842,high,pr.body,e many other callers intentionally want the current concrete hint for heuristics unrelated to compile-time max-autotune benchmarking. Fixes #148842 Generated by my agent Test Plan: python test/inductor/test_select_algorithm.py TestAutotuneDynamicMaxSize.test_benchmark_inputs_u...,https://github.com/pytorch/pytorch/pull/185773,416792e4f7be2dc5b3fd190d358f4970bc2d1e3db1a80577c75a3e24a9525866 closes,pr,185774,issue,148779,high,pr.body,subclass failure in place. Redirecting Tensor.numpy itself fixes the common entry point and matches the established Dynamo strategy. Fixes #148779 Generated by my agent Test Plan: python test/export/test_export.py TestExport.test_non_strict_export_tensor_numpy TestExport.test_...,https://github.com/pytorch/pytorch/pull/185774,8cd8653de39662a1b15166730b3c11dd3e92a00b50ab760ae59d9d34d0c459ec closes,pr,185456,issue,156614,high,pr.body,other keyword-shaped targets or kwargs. Fixing FX keeps the behavior general and lets Dynamo use its established generic method path. Fixes #156614 Generated by my agent Test Plan: python test/test_fx_graph_print.py -k test_codegen_torch_op_overload_with_keyword_attribute pyth...,https://github.com/pytorch/pytorch/pull/185456,4a3d43813da693df446a1961b9330cb6c63374c20133f7a1e51c613f0cc6aa14 closes,pr,184130,issue,148651,high,pr.body,"re any torch import starts native threads. Defer torch key, Triton path, and async-compile setup to the actual worker initializer.\n\nFixes #148651\nGenerated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nr...",https://github.com/pytorch/pytorch/pull/184130,cc16250a3e039192c06f3ba86e3733e627f2ba22519cb7d2f8db87d73bc0f98e closes,pr,185781,issue,148199,high,pr.body,ting CPU qlinear/select-algorithm test file. The test checks both correctness and that qlinear_binary lowering is actually exercised. Fixes #148199 Generated by my agent Test Plan: Reproduced the original dynamic-shape qlinear_binary repro before the fix: generated code return...,https://github.com/pytorch/pytorch/pull/185781,ce548ecab1ab627331b701cac5da9b1504c29110a0f35595a4dbe9e6126c3daa closes,pr,185408,issue,159160,high,pr.body,"l Dynamo or symbolic-shape shutdown log calls because the bug is caused by handler stream lifetime, not by those specific call sites. Fixes #159160 Generated by my agent Test Plan: TORCH_LOGS=""+dynamo"" pytest test/dynamo/test_subgraphs.py -k test_control_flow5 pytest test/test...",https://github.com/pytorch/pytorch/pull/185408,15e3d43740e153645d47eaeedcd2d30209b62b3a0f941a025686674d8e5e3edc closes,pr,185526,issue,154282,high,pr.body,nstead of the generated GraphModule forward. Materializing at the GraphModule boundary is narrower and matches the existing contract. Fixes #154282 Generated by my agent Test Plan: python test/dynamo/test_modules.py -k test_lazy_graph_module_fullgraph_call python test/dynamo/t...,https://github.com/pytorch/pytorch/pull/185526,d7ecf0c8f6b8ae969e13cea3d93dc56b4f5c87c144209f5ae726321c80e78983 review guidance,pr,185526,issue,154282,high,pr.reviews[0].body,"It's unclear to me when force_lazy_graph_module_recompile needs to be called. Can we document this in more detail, or refactor so that there are fewer calls scattered around the codebase?",https://github.com/pytorch/pytorch/pull/185526,a8f8ac5d83e61ad75744ca7daf98f9505e5d7b6fa6d93d8f9f5bf2c5643eaee0 review guidance,pr,185526,pr,185526,high,pr.reviews[0].body,"It's unclear to me when force_lazy_graph_module_recompile needs to be called. Can we document this in more detail, or refactor so that there are fewer calls scattered around the codebase?",https://github.com/pytorch/pytorch/pull/185526,7de23d01567ba55ebf10c4ecde62526d3c8b9dc6a12eb0151d639916a6630eae closes,pr,185784,issue,148112,high,pr.body,"be the scatter payload. Fixing it there keeps scalar captured gradients unchanged and handles vector captured gradients consistently. Fixes #148112 Generated by my agent Benchmark Results: Existing scalar captured-gradient path, baseline main loaded from HEAD:torch/_inductor/s...",https://github.com/pytorch/pytorch/pull/185784,b03a805be56f4345b16d72a12b72f707eff68872be137411f0d2991c051bdcb2 closes,pr,185786,issue,136628,high,pr.body,"ter dispatch rather than weakening pow_by_natural itself so explicit PowByNatural nodes retain their stronger contract. Fixes #148003 Fixes #136628 Generated by my agent Benchmark Results: Microbenchmark: 7 samples of 10,000 bound_sympy calls each, on x2 and xy, plus one x**-1...",https://github.com/pytorch/pytorch/pull/185786,0e28d15b9d9c888d489773361890168ddb062c407aed202b1c1ba2c6fece942a closes,pr,185786,issue,148003,high,pr.body,"n the interpreter dispatch rather than weakening pow_by_natural itself so explicit PowByNatural nodes retain their stronger contract. Fixes #148003 Fixes #136628 Generated by my agent Benchmark Results: Microbenchmark: 7 samples of 10,000 bound_sympy calls each, on x2 and xy,...",https://github.com/pytorch/pytorch/pull/185786,5ba7af5890d774376ca1d7cc681371efc24b1a46605700ed8ddbf6809d6218b1 closes,pr,183904,issue,176929,high,pr.body,"r with the CPU FMA contraction order that matches eager, while preserving eager validation and autograd behavior for supported cases. Fixes #176929 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...",https://github.com/pytorch/pytorch/pull/183904,6ee1687f7f230b0332c322c2135da3ae35e50e4736f9c4572fb8604d1759a92a closes,pr,185795,issue,147843,high,pr.body,"ixes the root cause instead of disabling cloning for only the optimized model, which previously regressed accuracy minifier behavior. Fixes #147843 Generated by my agent Test Plan: python test/dynamo/test_misc.py MiscTests.test_clone_input_preserves_expanded_input_stride MiscT...",https://github.com/pytorch/pytorch/pull/185795,227f9c0e9dc273a2ca26c4e1507973adab6c440e95b412e72519ff4fdf6ceb52 closes,pr,185799,issue,147530,high,pr.body,eaves for those public helpers and uses private dataclass-aware mapping only in FX internals that own graph bookkeeping or execution. Fixes #147530 Generated by my agent Benchmark Results: Command: Python microbenchmark constructing 40 FX Graphs with 1000 chained operator.add...,https://github.com/pytorch/pytorch/pull/185799,55d6b776cdbcd4a083617629a0f7aab32deab004596925a23985ec7e532fc95a closes,pr,185803,issue,147380,high,pr.body,havior for true unused parameters. Keeping the fix in export's tied-alias handling preserves the existing unused-parameter assertion. Fixes #147380 Generated by my agent Benchmark Results: On a 4-layer Linear Module benchmark measuring _export_forward_backward over 10 iteratio...,https://github.com/pytorch/pytorch/pull/185803,bbc33093f3ad60d55ebb076ebea756268d75e6ed8d672808ca5c6328ed7aab2d closes,pr,185914,issue,147135,high,pr.body,hat was not enough because this early error path did not actually emit the metadata into TORCH_TRACE; this change logs it explicitly. Fixes #147135 Generated by my agent Test Plan: python repro before the fix: reproduced RuntimeError containing fw_metadata=ViewAndMutationMeta(...,https://github.com/pytorch/pytorch/pull/185914,b096ad9203314cc185931c4cc73cc8bc924f68002445702a192c407c9edf2891 closes,pr,185809,issue,147077,high,pr.body,iled graph that is first run with a valid backing storage then called with a smaller backing storage that must still fail at runtime. Fixes #147077 Generated by my agent Benchmark Results: CPU microbenchmark of a valid compiled as_strided_copy after one compile call and 10 war...,https://github.com/pytorch/pytorch/pull/185809,bb4156b519973d1e1cb1a997e699057460d50a457f0f2463e56ced9014bc3082 closes,pr,186522,issue,113899,high,pr.body,"s. The explicit branch follows the existing proxy-tracking pattern used for pre-dispatch Python APIs that need proxy-aware arguments. Fixes #113899 Generated by my agent Benchmark Results: Microbenchmark: 100 iterations of make_fx(f, pre_dispatch=True)(x) where f clones a 16-e...",https://github.com/pytorch/pytorch/pull/186522,719f6d9456633052c948170e062628f2ebdd4c9a0dbdb317f0eb97fc45d04a3e review guidance,pr,186522,issue,113899,high,pr.reviews[0].body,"not an aten operator, I don't think it should go into the pre-dispatch graph...",https://github.com/pytorch/pytorch/pull/186522,ef2f0972bba56179f8ae9b558ed0afdb39beea478f3847de919af59bf849e63b review guidance,pr,186522,pr,186522,high,pr.reviews[0].body,"not an aten operator, I don't think it should go into the pre-dispatch graph...",https://github.com/pytorch/pytorch/pull/186522,2e6a9b6ee76007f07188d25caae09c60a49a4209d5f62e582bb41825202085da review guidance,pr,186522,issue,113899,high,pr.reviews[1].body,"not an aten operator, I don't think it should go into the pre-dispatch graph...",https://github.com/pytorch/pytorch/pull/186522#pullrequestreview-4500435392,7404ba3b42dac6707cc68819edbf2af440a82727c114da69d8d2301f1369eeda review guidance,pr,186522,pr,186522,high,pr.reviews[1].body,"not an aten operator, I don't think it should go into the pre-dispatch graph...",https://github.com/pytorch/pytorch/pull/186522#pullrequestreview-4500435392,84b60ecd9ab8c0ff361b14c2c2a1c666f8904cda3087617665082b9f1d4737d2 closes,pr,185916,issue,146896,high,pr.body,"the index there, but that would only paper over one shape mismatch and would still leave the intermediate scatter reduced too early. Fixes #146896 Generated by my agent Test Plan: Reproduced the issue before the fix with the CUDA repro from #146896; backward failed with Assert...",https://github.com/pytorch/pytorch/pull/185916,c94519d7a094018dd59cef8308804dc20ea41cd37482afbcd36b2ccb63d7f849 closes,pr,185823,issue,146673,high,pr.body,"xtension descriptors, keyword arguments, subclass override semantics, and inherited descriptor access through subclass class objects. Fixes #146673 Generated by my agent Test Plan: python test/dynamo/test_repros.py -k test_class_method_descriptor_call git diff --check lintrunn...",https://github.com/pytorch/pytorch/pull/185823,026d983af83a430731481c6224df568a304e5dfb6733635ad4b0c44da3063172 closes,pr,187633,issue,108406,high,pr.closingIssuesReferences,pr #187633 declares a closing reference to issue #108406.,https://github.com/pytorch/pytorch/pull/187633,fa634bdca4b446b73db4132200d7edf58b3ba4b08f30f83c4f5a77f4ccc566a9 closes,pr,187633,issue,108406,high,pr.body,"and venv environments on both Linux and macOS, including a quoting bug where the shell command substitution was not being evaluated Closes #108406",https://github.com/pytorch/pytorch/pull/187633,f61950e16bc7652e3c1222f8023e7ca85da393e2f1f60456e6fb09eb8ee722b5 references,pr,187633,pr,180247,medium,pr.reviews[0].body,"ect is that the docs lean on setup.py develop, which is on its way out. The build system is migrating from setuptools to scikit-build-core (#180247 and the stack on top of it), and setup.py is removed as part of that. So the new advice to run python setup.py develop --uninstal...",https://github.com/pytorch/pytorch/pull/187633,7e83fefc394310a8c2de11189d7bdedd97a1f9692704093db01da7fce979776d review guidance,pr,187633,issue,108406,high,pr.reviews[0].body,"Thanks for taking this on, @w1ndcn, and thanks for the ping, @albanD. On the direction question first: yes, better developer-build docs are very welcome, and a few pieces here are good as they stand - the CMAKE_PREFIX_PATH quoting fix is a real bug catch, and the ""Verify Your Installation"" and tr...",https://github.com/pytorch/pytorch/pull/187633,9ff9e184347f6ea63d58776597c9528100d6fc83bb0dc9647e5dd5954cef8856 closes,pr,185920,issue,146628,high,pr.body,adata. The schema checks keep the rejection close to each mutating entry point while still using the existing SideEffects error path. Fixes #146628 Generated by my agent Test Plan: python test/dynamo/test_generator.py GeneratorTests.test_reconstruct_generator_tensor_mutation G...,https://github.com/pytorch/pytorch/pull/185920,3c1b4fcf9c9f7c6126487c5e55c301cf8faf5b7ec8630125712f37c2306cba06 closes,pr,185691,issue,151510,high,pr.body,"Stack from ghstack (oldest at bottom): -> #185691 Fixes #151510 The repro passes torch.iinfo(torch.int64).max as the posinf replacement to nan_to_num, whose API stores replacement values as double. That",https://github.com/pytorch/pytorch/pull/185691,132e1cc61911a5cf6570a2e5b3fd5ba6c0582db00723af97240d74d887b44fcc closes,pr,184153,issue,146569,high,pr.body,nvalidate live .grad storage. Clone backward outputs before handing them to autograd so persistent grads have normal tensor lifetime. Fixes #146569 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/184153,0449a0cb7acc7bb047ae81cb3d841e5e87289bbff1ac23e5f7813abb6fcf6bc3 closes,pr,184161,issue,145093,high,pr.body,"stant loads, including input mutation copyback paths. Conjugated complex inputs continue to use fallback rather than generated loads. Fixes #145093 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...",https://github.com/pytorch/pytorch/pull/184161,7beb8cd9c6dade4431239fc196179026dff46906bd3425d15e2d7b42781ca831 references,pr,184161,issue,186031,medium,pr.comments[9].body,My agent says the remaining TorchTitan failure matches the known torchcomms TP+PP+compile issue tracked in #186031 (`RuntimeError: Could not resolve the process group registered under the name ...`). I do not see evidence that this is caused by the lazy,https://github.com/pytorch/pytorch/pull/184161,8ff7539a740a58d1703bbee3a697a12b8289e9e6e5054bff789502abcf6ec48f review guidance,pr,184161,issue,145093,high,pr.reviews[0].body,"aotautograd should be removing the neg tensor semantics, as discussed in the issue",https://github.com/pytorch/pytorch/pull/184161,c119b498288ae84f5ebf8f8296403c825f02ec2e5fb1c55162fcad59cadba121 review guidance,pr,184161,pr,184161,high,pr.reviews[0].body,"aotautograd should be removing the neg tensor semantics, as discussed in the issue",https://github.com/pytorch/pytorch/pull/184161,3e03a3262a6e5a912a6cb6006e9f600a70a69f43fe344a26287bf919f772de0c review guidance,pr,184161,issue,145093,high,pr.reviews[1].body,"Read the tests and they def look fine to me. For the rest of the code; I'd defer to someone else, but I found no glaring issues.",https://github.com/pytorch/pytorch/pull/184161,ec01ebaeaf29ecd162548bca2d5b8da9eb71a0f97b96b444fae3c87f54a46d45 review guidance,pr,184161,pr,184161,high,pr.reviews[1].body,"Read the tests and they def look fine to me. For the rest of the code; I'd defer to someone else, but I found no glaring issues.",https://github.com/pytorch/pytorch/pull/184161,40a1978b2a77272479b89514ecd5446b549b4c3949666b2d8ceb32dd638cef22 review guidance,pr,184161,issue,145093,high,pr.reviews[2].body,"aotautograd should be removing the neg tensor semantics, as discussed in the issue",https://github.com/pytorch/pytorch/pull/184161#pullrequestreview-4410086434,8bcaadf405f5ca38eaa4677b40041be263e23e0e15ec28f05e187820de4550f1 review guidance,pr,184161,pr,184161,high,pr.reviews[2].body,"aotautograd should be removing the neg tensor semantics, as discussed in the issue",https://github.com/pytorch/pytorch/pull/184161#pullrequestreview-4410086434,5603b3dbd822e43f0e0095960ca4603ee4b96c5e7431c185dcc23a92b70c87a6 review guidance,pr,184161,issue,145093,high,pr.reviews[3].body,"Read the tests and they def look fine to me. For the rest of the code; I'd defer to someone else, but I found no glaring issues.",https://github.com/pytorch/pytorch/pull/184161#pullrequestreview-4442431668,d41d93b8a4ff99491da967522c80c7f9402268888db0e0d0a38a317648674f8b review guidance,pr,184161,pr,184161,high,pr.reviews[3].body,"Read the tests and they def look fine to me. For the rest of the code; I'd defer to someone else, but I found no glaring issues.",https://github.com/pytorch/pytorch/pull/184161#pullrequestreview-4442431668,0704c24463a4e01819c1b5463aff40ea8ddf0247e69db85c55bfee6557cafece references,pr,181728,pr,181726,medium,pr.body,"don't have ghstack permission, I manually created the following stacked PRs for review. I also created a combined PR #187318 to test on CI. #181726 [xpu][1/4]Implement scaled_mm_v2 for MXFP8/MXFP4/NVFP4 on XPU #181727 [xpu][2/4]Implement scaled_mm_v1 for MXFP8/MXFP4/NVFP4 on X...",https://github.com/pytorch/pytorch/pull/181728,7a4c8fca546d3691abbe70e695c073fb6575892ea97bb67ca463094a35fbdfa1 references,pr,181728,pr,181727,medium,pr.body,d PRs for review. I also created a combined PR #187318 to test on CI. #181726 [xpu][1/4]Implement scaled_mm_v2 for MXFP8/MXFP4/NVFP4 on XPU #181727 [xpu][2/4]Implement scaled_mm_v1 for MXFP8/MXFP4/NVFP4 on XPU #187315 [xpu][3/4] inductor: route MX scaled_mm_v2 fallback through...,https://github.com/pytorch/pytorch/pull/181728,a79db783a090c751492289799fd7e9dac7c2e55e3f3f18dde4e44854dc66dd87 references,pr,181728,pr,187315,medium,pr.body,#181726 [xpu][1/4]Implement scaled_mm_v2 for MXFP8/MXFP4/NVFP4 on XPU #181727 [xpu][2/4]Implement scaled_mm_v1 for MXFP8/MXFP4/NVFP4 on XPU #187315 [xpu][3/4] inductor: route MX scaled_mm_v2 fallback through v2 aten kernel #181728 [xpu][4/4]Enable MXFP8/MXFP4/NVFP4 tests for X...,https://github.com/pytorch/pytorch/pull/181728,502fcb32c772b0940c040aa6364fa95c2f26ecc91f3f3b7c891828bbbbb10e68 references,pr,181728,pr,187318,medium,pr.body,"sts. PR Stack: Since I don't have ghstack permission, I manually created the following stacked PRs for review. I also created a combined PR #187318 to test on CI. #181726 [xpu][1/4]Implement scaled_mm_v2 for MXFP8/MXFP4/NVFP4 on XPU #181727 [xpu][2/4]Implement scaled_mm_v1 for...",https://github.com/pytorch/pytorch/pull/181728,a55ee4ae9c9799561e7eb2cac2f5a4e2b04e4969849e729cbac3a235f47beb6b review guidance,pr,181728,pr,181726,high,pr.reviews[0].body,"Overall LGTM, @carsonwang could you merge three prs into one draft PR and launch a CI for full test? So we could show the CI result",https://github.com/pytorch/pytorch/pull/181728,6c43b7eb30f13b290a8f8ff39a5e387325370197f77eb97536c307fe72b81e2f review guidance,pr,181728,pr,181727,high,pr.reviews[0].body,"Overall LGTM, @carsonwang could you merge three prs into one draft PR and launch a CI for full test? So we could show the CI result",https://github.com/pytorch/pytorch/pull/181728,46a2230fa29c5059250803b904e6f0fc355db3e93fdccc6f21ef5590a8583a94 review guidance,pr,181728,pr,181728,high,pr.reviews[0].body,"Overall LGTM, @carsonwang could you merge three prs into one draft PR and launch a CI for full test? So we could show the CI result",https://github.com/pytorch/pytorch/pull/181728,78acc8e83aa1f24eabb6e9c159e579e688fa0f247fe747f526592eb0e459bda1 review guidance,pr,181728,pr,187315,high,pr.reviews[0].body,"Overall LGTM, @carsonwang could you merge three prs into one draft PR and launch a CI for full test? So we could show the CI result",https://github.com/pytorch/pytorch/pull/181728,7f0748b8a47a6f7d626be7766c80830eeb71d08b105c06891f77292bacd06bbb review guidance,pr,181728,pr,187318,high,pr.reviews[0].body,"Overall LGTM, @carsonwang could you merge three prs into one draft PR and launch a CI for full test? So we could show the CI result",https://github.com/pytorch/pytorch/pull/181728,e2c351ad60d59bb1f4d1c89dbf0806499909da7693199d7ec53ee1fffd907faa references,pr,187315,pr,181726,medium,pr.body,"don't have ghstack permission, I manually created the following stacked PRs for review. I also created a combined PR #187318 to test on CI. #181726 [xpu][1/4]Implement scaled_mm_v2 for MXFP8/MXFP4/NVFP4 on XPU #181727 [xpu][2/4]Implement scaled_mm_v1 for MXFP8/MXFP4/NVFP4 on X...",https://github.com/pytorch/pytorch/pull/187315,17a5791de8fad7fcc701d6380e275c86ca579c56850c210210cfaf79cf474daa references,pr,187315,pr,181727,medium,pr.body,d PRs for review. I also created a combined PR #187318 to test on CI. #181726 [xpu][1/4]Implement scaled_mm_v2 for MXFP8/MXFP4/NVFP4 on XPU #181727 [xpu][2/4]Implement scaled_mm_v1 for MXFP8/MXFP4/NVFP4 on XPU #187315 [xpu][3/4] inductor: route MX scaled_mm_v2 fallback through...,https://github.com/pytorch/pytorch/pull/187315,d979a0414f1a8a3115b97da56141facb951d81beda5d851848168d1cb5862daf references,pr,187315,pr,181728,medium,pr.body,[2/4]Implement scaled_mm_v1 for MXFP8/MXFP4/NVFP4 on XPU #187315 [xpu][3/4] inductor: route MX scaled_mm_v2 fallback through v2 aten kernel #181728 [xpu][4/4]Enable MXFP8/MXFP4/NVFP4 tests for XPU Authored with Claude. cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Ch...,https://github.com/pytorch/pytorch/pull/187315,6df09356dab95406151aa9cad2bdabecbcda875a9bb5d39f02d02c9284d2deca references,pr,187315,pr,187318,medium,pr.body,"fic. PR Stack: Since I don't have ghstack permission, I manually created the following stacked PRs for review. I also created a combined PR #187318 to test on CI. #181726 [xpu][1/4]Implement scaled_mm_v2 for MXFP8/MXFP4/NVFP4 on XPU #181727 [xpu][2/4]Implement scaled_mm_v1 for...",https://github.com/pytorch/pytorch/pull/187315,7f8753fc658d9a9f7b788e92700b2f740600d76bfe3e489d9fd35f6d9b11cc26 references,pr,187315,issue,188721,medium,pr.comments[0].body,"akiness on trunk: B200 Smoke Tests / linux-jammy-cuda13.0-py3.12-gcc11-sm100 / test-osdc (smoke_b200, 1, 1, mt-l-x86iamx-22-225-b200) (gh) (#188721) test/inductor/test_flex_flash.py::TestFlexFlashCUDA::test_cutedsl_captured_alias_views_keep_distinct_layouts_case_offset_cuda Th...",https://github.com/pytorch/pytorch/pull/187315,9444e4c8c45c72ff9bb31953755564c17ae0d7b6ac0031c0e49a5e178f320743 closes,pr,186583,issue,186537,high,pr.body,"shapes to exceed the default recompile limit, and asserts both the forward and recompute phases see the caller's raised Dynamo limit. Fixes #186537 Generated by my agent Benchmark Results: Post-fix Python autograd.Function backward microbenchmark, 20,000 backward calls per rep...",https://github.com/pytorch/pytorch/pull/186583,e80442fb1198a695139d509e398d3ee6358e7f0bdb154f5e6753da3894aa690d review guidance,pr,186583,issue,186537,high,pr.reviews[0].body,"Approving since this fixes the bug, but two notes (can be addressed here or in follow up PRs) We should remove these lines pytorch/torch/_functorch/_aot_autograd/runtime_wrappers.py Lines 3246 to 3250 in c35920b lines.append("" if (torch._C._is_key_in_tls('context')"") lines.append( "" and (_cc := t...",https://github.com/pytorch/pytorch/pull/186583,fb91aa7da92e849c9d6de65193afd1028b7a5fe3d28f97a9c961968201e560ed review guidance,pr,186583,pr,186583,high,pr.reviews[0].body,"Approving since this fixes the bug, but two notes (can be addressed here or in follow up PRs) We should remove these lines pytorch/torch/_functorch/_aot_autograd/runtime_wrappers.py Lines 3246 to 3250 in c35920b lines.append("" if (torch._C._is_key_in_tls('context')"") lines.append( "" and (_cc := t...",https://github.com/pytorch/pytorch/pull/186583,bbdad869f63e4233639c05be902985626a5fbbdfdfc64c5d80f407517ced0112 closes,pr,185922,issue,146274,high,pr.body,"re real tensor values, avoiding data-dependent fake execution while still matching eager for the reported static invalid class count. Fixes #146274 Generated by my agent Test Plan: ninja -C build torch_python ninja -C build install python test/test_fake_tensor.py -k test_one_h...",https://github.com/pytorch/pytorch/pull/185922,ba9ee244641d6c61632c5a3301f0c7f81d665f0a7d37958a958b2731176024ef closes,pr,185829,issue,145899,high,pr.body,"t trace_autograd_ops=False path for correctness. The trace_autograd_ops=True path, which traces autograd.grad directly, is unchanged. Fixes #145899 Generated by my agent Test Plan: python test/dynamo/test_fwd_loss_bwd.py TestForwardLossBackward.test_autograd_grad_kwarg_after_g...",https://github.com/pytorch/pytorch/pull/185829,d3e879ccb6c059c163b2eb286dd4dc1bb102cb68a7c83d5209d93e21e57a1af7 review guidance,pr,185829,issue,145899,high,pr.reviews[0].body,This looks correct to me after some digging. I think it would be a good idea to get a second human opinion here. @williamwen42 @jansel (not the agent),https://github.com/pytorch/pytorch/pull/185829,7e7be7719da426171b184931c714678b72427908d341ca9349592ddb9109628d review guidance,pr,185829,pr,185829,high,pr.reviews[0].body,This looks correct to me after some digging. I think it would be a good idea to get a second human opinion here. @williamwen42 @jansel (not the agent),https://github.com/pytorch/pytorch/pull/185829,96e96bc619aaf256b97605763c913309d5529900e26aaf795241bead9e4f8102 closes,pr,184178,issue,143412,high,pr.body,his prevents padding and tail lanes from creating out-of-bounds pointer values while preserving the original masks and store indices. Fixes #143412 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/184178,787cdffff7908c664f559e1dd0477b67bd3fa8ec694bb52d2d95313a7238beef closes,pr,182273,issue,178977,high,pr.closingIssuesReferences,pr #182273 declares a closing reference to issue #178977.,https://github.com/pytorch/pytorch/pull/182273,0b69bb4d33378ae9a101f65fdc38f054e4aa77ee3c0290e426eab6c905f1611e closes,pr,182273,issue,178977,high,pr.body,nd override both in ProcessGroupWrapper to forward to the wrapped backend. Also expose bound_device_id on the Backend Python binding. Fixes #178977,https://github.com/pytorch/pytorch/pull/182273,0e583cb46046001d458bce7432b2b2a107143e5da53a5c3e9ba6631ee7b61e79 review guidance,pr,185586,pr,185586,high,pr.reviews[0].body,"SGTM, thank you. I think there are some conflicts with this PR, so could you please resolve the conflicts frist.",https://github.com/pytorch/pytorch/pull/185586,189c04a383d68b7e4394aaa07e7d8ab7052eb957a6260227374d216d63ce228c closes,pr,188906,issue,162980,high,pr.closingIssuesReferences,pr #188906 declares a closing reference to issue #162980.,https://github.com/pytorch/pytorch/pull/188906,49563eb9f191abb944a571f907b0006d1b4ddbb3d2849bf4ccf7d47309574f61 closes,pr,188906,issue,162980,high,pr.body,Fixes #162980 Summary The determinism_check argument of torch.utils.checkpoint.checkpoint looks up its metadata function in the module-private dict _allo,https://github.com/pytorch/pytorch/pull/188906,7e0a213c9842a8aedf5db221e9dd57302800a002857b7ad0a1e696dab00a3e5e closes,pr,184820,issue,175292,high,pr.body,ing default construction path unchanged and routes only custom metaclass calls through the special-method dispatch they actually use. Fixes #175292 Generated by my agent Test Plan: python test/dynamo/test_enum.py EnumTests.test_metaclass_custom_call python test/dynamo/test_enu...,https://github.com/pytorch/pytorch/pull/184820,cd30296312f43031dc75982e21afe33974febf323dc80657d3c1fb8d057d2fff review guidance,pr,184820,issue,175292,high,pr.reviews[0].body,Looks fine to me aside from a few comments.,https://github.com/pytorch/pytorch/pull/184820,f8b164230be523161716935b3a226f1a409e81098517f1c94a24687668cc7044 review guidance,pr,184820,pr,184820,high,pr.reviews[0].body,Looks fine to me aside from a few comments.,https://github.com/pytorch/pytorch/pull/184820,aafaa500de7e7708d2567232a815beae52355735581cc80bf2bfb104912d0b25 closes,pr,186043,issue,143157,high,pr.body,fore torch._check can see it. Add an error hint for that path pointing users to torch.sym_not() or an equivalent symbolic comparison. Fixes #143157 Generated by my agent Test Plan: python test/dynamo/test_misc.py MiscTests.test_torch_check_symbool_python_not -v python test/exp...,https://github.com/pytorch/pytorch/pull/186043,1e48548777aa3f485ddda1e405a62df9b4a905ef772fe37f3ce404429d424c73 closes,pr,186051,issue,142358,high,pr.body,"re rather than matching source debug-name strings, so unrelated locals named _forward_pre_hooks still honor torch.compiler.disable(). Fixes #142358 Generated by my agent Test Plan: python test/dynamo/test_hooks.py HooksTests.test_disabled_noop_named_forward_pre_hooks_still_gra...",https://github.com/pytorch/pytorch/pull/186051,2a176b99b4f51946af44e424405a884f3d1da0dfdcd67e34325e83e37ec429db closes,pr,186052,issue,142321,high,pr.body,"er adds about 0.13 microseconds per emitted pybinding source in this microbenchmark, which is negligible relative to C++ compilation. Fixes #142321 Generated by my agent Test Plan: python test/inductor/test_compile.py -k 'cpp_pybinding_source_literal' python test/inductor/test...",https://github.com/pytorch/pytorch/pull/186052,89e66ff9206acd4ffa02f06549375477bd2db8a2c8e41fa1dc7bec679b761e94 closes,pr,186554,issue,91468,high,pr.body,"an intentional correctness trade-off for this graph shape, and it is tracked by the aot_autograd graph_has_dependent_outputs counter. Fixes #91468 Generated by my agent Test Plan: python test/dynamo/test_hooks.py -k ""output_hook_on_dependent_output"" python test/dynamo/test_hoo...",https://github.com/pytorch/pytorch/pull/186554,0ffbecab64cf8d31715793915d48133b2f9e803717af08234d606e5492133baf closes,pr,185899,issue,149911,high,pr.body,"de fullgraph failures look better but broke catchability, so the observed-exception path is used with an uncaught conversion instead. Fixes #149911 Generated by my agent Test Plan: python test/dynamo/test_input_attr_tracking.py TestInputAttrTracking.test_set_data_on_input_tens...",https://github.com/pytorch/pytorch/pull/185899,9e5f07852b8950656181b5cc6b1b74b6049057380d178d3c4d7b1c197c564fb9 review guidance,pr,185899,issue,149911,high,pr.reviews[0].body,"@mlazos has a change that might affect this, would wait until that gets merged",https://github.com/pytorch/pytorch/pull/185899,4b0d01f50c3f3cb715e329298b7676b891750982f2074e972dcf2f62151e27ef review guidance,pr,185899,pr,185899,high,pr.reviews[0].body,"@mlazos has a change that might affect this, would wait until that gets merged",https://github.com/pytorch/pytorch/pull/185899,cedb1ae554e9f01eb5e24847f2dc59aaf4ddcf3cbcd46a0a7199ed079851dcd0 closes,pr,187589,issue,126024,high,pr.body,ing state instead of a removed global. This is a focused version of the thread-safety fix related to the stale prior work in #168999. Fixes #126024 Generated by my agent Test Plan: python test/dynamo/test_misc.py -k test_fx_tracing_flag_is_thread_local_for_compile python test/...,https://github.com/pytorch/pytorch/pull/187589,3ed65f581f49a028a04fa066ac7f5a4f61107dc6d85aa539516db81c6673fc10 competes with,pr,187589,pr,168999,medium,pr.body,hread's FX tracing state instead of a removed global. This is a focused version of the thread-safety fix related to the stale prior work in #168999. Fixes #126024 Generated by my agent Test Plan: python test/dynamo/test_misc.py -k test_fx_tracing_flag_is_thread_local_for_compi...,https://github.com/pytorch/pytorch/pull/187589,25ec4226faa2bb3031396e9cbb8f64b7f387d418b92dda13453b23ea01327f13 closes,pr,184185,issue,141916,high,pr.body,eping the reduction tile live materially reduces global memory traffic. Also fix the memory stats addition used by that heuristic.\n\nFixes #141916\nGenerated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nr...,https://github.com/pytorch/pytorch/pull/184185,4baeb3bf4ea58da96e3f8c4f791fc11a13ccaaf4a211646f6a396137f4d75f3a references,pr,184185,issue,141916,medium,pr.comments[10].body,"per? The device guard limits blast radius, but it would be good to have positive evidence on the target hardware before merging. The issue (#141916) was reported on H100 which suggests this is the right target, but explicit benchmark numbers would strengthen confidence. --- ##...",https://github.com/pytorch/pytorch/pull/184185,e812b217a52db8ef012a11fb93b5f034bf648f0f29555f7d525003874b288b90 review guidance,pr,184185,issue,141916,high,pr.reviews[0].body,Probably hold off on this before better benchmarking evidence across variety of shapes/ideally hardware,https://github.com/pytorch/pytorch/pull/184185,f497692ff7dd31e0b2cb56563eb52ae92f1a7e9d0fad9b59b8668d05525d98c1 review guidance,pr,184185,pr,184185,high,pr.reviews[0].body,Probably hold off on this before better benchmarking evidence across variety of shapes/ideally hardware,https://github.com/pytorch/pytorch/pull/184185,59192ef3226244ddfb4a599527063847e89c6d1c3870a6ce64996e42ee92d797 review guidance,pr,184185,issue,141916,high,pr.reviews[1].body,would prefer to defer on heuristic validation is better (which i am actively working on ).,https://github.com/pytorch/pytorch/pull/184185,fede9ec3dff9df7b389b86a5f94ec80930f842f3de72d8198198df3f6acdedd3 review guidance,pr,184185,pr,184185,high,pr.reviews[1].body,would prefer to defer on heuristic validation is better (which i am actively working on ).,https://github.com/pytorch/pytorch/pull/184185,3b4794376a6abeba9f5077465b170227bad0aa0bd8660219a336d48e6d8332d9 review guidance,pr,184185,issue,141916,high,pr.reviews[2].body,Probably hold off on this before better benchmarking evidence across variety of shapes/ideally hardware,https://github.com/pytorch/pytorch/pull/184185#pullrequestreview-4337738483,ee2c6ba99cee0e6090e505f6ce37bba4bba7fae70f6d7918569f678bffba6e15 review guidance,pr,184185,pr,184185,high,pr.reviews[2].body,Probably hold off on this before better benchmarking evidence across variety of shapes/ideally hardware,https://github.com/pytorch/pytorch/pull/184185#pullrequestreview-4337738483,dfb5a7237b3aa836f3210a3773573ad6d005a69bd1655f3b5346a6803fd1e65f review guidance,pr,184185,issue,141916,high,pr.reviews[3].body,would prefer to defer on heuristic validation is better (which i am actively working on ).,https://github.com/pytorch/pytorch/pull/184185#pullrequestreview-4479138773,f1755fd88886c35a872a3ef57ef9293a742ad029d4200ba5f897571fa708716d review guidance,pr,184185,pr,184185,high,pr.reviews[3].body,would prefer to defer on heuristic validation is better (which i am actively working on ).,https://github.com/pytorch/pytorch/pull/184185#pullrequestreview-4479138773,f2733b4250ee44c107e0ae28e966728a640b22e05cf30c3a669395dd119ccea1 closes,pr,186053,issue,141790,high,pr.body,"forward/backward cache-entry contract rather than introducing partial AOTAutograd entries or broadly eager-compiling lazy backwards. Fixes #141790 Generated by my agent Benchmark Results: On the issue reproducer, measured torch.compile(fn, backend=""inductor"", fullgraph=True)(a...",https://github.com/pytorch/pytorch/pull/186053,668f64a54f221c83a2a8c72cd810c2f6d4eb1454c3ef602b8b8b4465d28f26a2 closes,pr,184189,issue,141744,high,pr.body,"load time and keys async/Python caches with the runtime device properties, avoiding stale architecture metadata in standalone files. Fixes #141744 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...",https://github.com/pytorch/pytorch/pull/184189,201fa47da86fe3ee29327c749163365d38a60ea6b7d560deca677ee3b56503dd review guidance,pr,184189,issue,141744,high,pr.reviews[0].body,"Needs to demonstrate that we're actually able to run across hardware. Also, I don't know if we should commit to this, in case we emit specialized inline asm etc",https://github.com/pytorch/pytorch/pull/184189,7ac7d9f01227c43fd6bdcb4733e563093fd75830b06ae1cd60e864d51a6ae3be review guidance,pr,184189,pr,184189,high,pr.reviews[0].body,"Needs to demonstrate that we're actually able to run across hardware. Also, I don't know if we should commit to this, in case we emit specialized inline asm etc",https://github.com/pytorch/pytorch/pull/184189,6cce26ddd102ee7a1ae6bc6119c31144c48e3949c56bab42d7c5e5467daa1217 closes,pr,186058,issue,141640,high,pr.body,alue metadata matches the intended FX metadata behavior while keeping value metadata under the transform that produced the new nodes. Fixes #141640 Generated by my agent Test Plan: python test/dynamo/test_export.py -k test_export_preserves_metadata_during_normalization python...,https://github.com/pytorch/pytorch/pull/186058,11ce14773b13749aaa8ab68e2e63bfacc4fab3d42056aaeb52fa110bacea1eef closes,pr,186063,issue,141473,high,pr.body,"r, super() and class identity, in-method forward mutation, concurrent delegated calls, and the previous patched-init robustness case. Fixes #141473 Generated by my agent Benchmark Results: Delegated method lookup microbenchmark, 1,000,000 iterations: old equivalent getattr(mod...",https://github.com/pytorch/pytorch/pull/186063,80fe2da9e47179b0a3a412931b73db7e88a449629c9340c94528c6f834543952 closes,pr,186069,issue,141258,high,pr.body,"ay semantics. This avoids the abandoned approach in #141493, which added round to torch._numpy.ndarray and changed torch_np behavior. Fixes #141258 Generated by my agent Test Plan: python test/dynamo/test_misc.py MiscTests.test_numpy_scalar_round MiscTests.test_numpy_scalar_dt...",https://github.com/pytorch/pytorch/pull/186069,c3a714fead1fde308ca004beb37ea0a87bfd8b812d922c357b977ea698ce1fa1 review guidance,pr,183613,pr,182696,high,pr.reviews[0].body,lgtm,https://github.com/pytorch/pytorch/pull/183613,1406619077cd759093957789f6809e57cbf656710b2e24f4c661ec6d30d0c96c closes,pr,185748,issue,149921,high,pr.body,"existing hook replacement/reset cases, and public dict/set container operations graph break when a runtime handle id could be stale. Fixes #149921 Generated by my agent Test Plan:\n- python test/dynamo/test_hooks.py HooksTests.test_register_forward_hook_inside_compiled_region...",https://github.com/pytorch/pytorch/pull/185748,5cbc13efe43b83b92cf77a2e08e76252a65bd80c928b11b7d248c0f0e10d599a references,pr,187318,pr,181726,medium,pr.body,The purpose of this PR is to test on CI only. I have split the code into the following 4 PRs for review: #181726 [xpu][1/4]Implement scaled_mm_v2 for MXFP8/MXFP4/NVFP4 on XPU #181727 [xpu][2/4]Implement scaled_mm_v1 for MXFP8/MXFP4/NVFP4 on XPU #187315,https://github.com/pytorch/pytorch/pull/187318,a8fb244e2440f827e7395f16d9f2160e1f832a63af5f4de9438e89adfd622309 references,pr,187318,pr,181727,medium,pr.body,n CI only. I have split the code into the following 4 PRs for review: #181726 [xpu][1/4]Implement scaled_mm_v2 for MXFP8/MXFP4/NVFP4 on XPU #181727 [xpu][2/4]Implement scaled_mm_v1 for MXFP8/MXFP4/NVFP4 on XPU #187315 [xpu][3/4] inductor: route MX scaled_mm_v2 fallback through...,https://github.com/pytorch/pytorch/pull/187318,a1617317ce0dbf4c98a1efb9a1d45d0d0d92e71ad16e5f92b8db339379378041 references,pr,187318,pr,181728,medium,pr.body,[2/4]Implement scaled_mm_v1 for MXFP8/MXFP4/NVFP4 on XPU #187315 [xpu][3/4] inductor: route MX scaled_mm_v2 fallback through v2 aten kernel #181728 [xpu][4/4]Enable MXFP8/MXFP4/NVFP4 tests for XPU cc @jgong5 @mingfeima @XiaobingSuper @sanchitintel @ashokei @jingxu10 @aditew01...,https://github.com/pytorch/pytorch/pull/187318,7719e19493dbf0432b144ccbb309cc6558beb2c62284876e8d4a49ad768585bc references,pr,187318,pr,187315,medium,pr.body,#181726 [xpu][1/4]Implement scaled_mm_v2 for MXFP8/MXFP4/NVFP4 on XPU #181727 [xpu][2/4]Implement scaled_mm_v1 for MXFP8/MXFP4/NVFP4 on XPU #187315 [xpu][3/4] inductor: route MX scaled_mm_v2 fallback through v2 aten kernel #181728 [xpu][4/4]Enable MXFP8/MXFP4/NVFP4 tests for X...,https://github.com/pytorch/pytorch/pull/187318,bec700fbfe9d29d51f2815f9516786b28dfe02dda70129942df984d13b7b4dfa closes,pr,189009,issue,188131,high,pr.closingIssuesReferences,pr #189009 declares a closing reference to issue #188131.,https://github.com/pytorch/pytorch/pull/189009,db9348f01f00be2759f515b93f24cb1521462262ca66ffbe39f2f15117e0394c closes,pr,189009,issue,188131,high,pr.body,"//2=16 after reshape. The XPU sycltla flash attention kernel only supports specific head_dim values and crashes for unsupported ones. Fixes #188131 (DISABLED test_fallback_kernel_with_symexpr_output_xpu). Solution Change tensor_shape to (4,128,4,4) so head_dim=64, which is sup...",https://github.com/pytorch/pytorch/pull/189009,82d267989e20e976948bb5ce3cd759f006343f1e783052b6b80ef26c77059884 review guidance,pr,178768,pr,3077,high,pr.reviews[0].body,"Found two things below: Enable the skipped XPU test in test_dlpack.py:527 Verify: Confirm that the Module.cpp:769 call site doesn't need the same fix (appears safe, but worth explicit verification)",https://github.com/pytorch/pytorch/pull/178768,8d86b09f98c4a239e3805a820901b809847e2938f3ac46ea26e58c11ebd003a0 closes,pr,184809,issue,175408,high,pr.body,tic change for custom ops and tensor subclasses; this patch keeps the existing intended behavior and makes the diagnostic actionable. Fixes #175408 Generated by my agent Test Plan: python test/dynamo/test_error_messages.py ErrorMessagesTest.test_custom_op_tensor_subclass_data_...,https://github.com/pytorch/pytorch/pull/184809,259e8bfa4102e142868a6366b5416487cee490eee63d71f46cae06b3df2639de closes,pr,184336,issue,183971,high,pr.closingIssuesReferences,pr #184336 declares a closing reference to issue #183971.,https://github.com/pytorch/pytorch/pull/184336,09fba1c12a3ff7d090d07bafd952609325702be1dae84fc4c95387736d164dc0 closes,pr,184336,issue,183971,high,pr.body,"dy-built wheel. The wheel already has libgomp in torch/lib/, so copy_libraries picks it up. The only missing piece was the rpath fix. Fixes #183971 Created with Assistance from Claude",https://github.com/pytorch/pytorch/pull/184336,732baec27463e228f553c394403bc8fc9434cee418f9ac3242fab3f732b32e4f review guidance,pr,184336,pr,174753,high,pr.reviews[0].body,Asked a question. Has libtorch CI been triggered?,https://github.com/pytorch/pytorch/pull/184336,566f935c77f1003031f5013e0c84362c75ef85e365106d71a1c9d1c650ab121c review guidance,pr,184336,issue,183971,high,pr.reviews[0].body,Asked a question. Has libtorch CI been triggered?,https://github.com/pytorch/pytorch/pull/184336,c9f2d3c637db9b819e0d31bd23ed4b322f6b1519476f77b6e211d1fd79bf2647 closes,pr,186614,issue,141184,high,pr.body,"d regression coverage for RNN, GRU, and LSTM modules that verifies Dynamo compiles both the pre-recurrent and post-recurrent regions. Fixes #141184 Generated by my agent Test Plan: python test/dynamo/test_modules.py -k test_rnn_graph_break_resumes_after_call python test/export...",https://github.com/pytorch/pytorch/pull/186614,6561c38395af1c051a0072dec1ddfe477fbc324fb54feefa1b3dad38874cd4e0 closes,pr,186072,issue,141162,high,pr.body,e by default. Keeping the default decision inside the logging argument checks preserves the narrower behavior requested by the issue. Fixes #141162 Generated by my agent Test Plan: python test/dynamo/test_reorder_logs.py ReorderLogsTests.test_reorder_constant_warnings_by_defau...,https://github.com/pytorch/pytorch/pull/186072,0080b5697b905ec26cbec05ef3a210d29e33950a5cebdf2d41821736378ddb17 closes,pr,186898,issue,141115,high,pr.body,"tput check entry point. The more expensive exact dtype, size, stride, and device validation remains behind AOTI_RUNTIME_CHECK_INPUTS. Fixes #141115 Generated by my agent Benchmark Results: Before/after on a tiny CPU AOTI package model x + 1, 7 samples of 20000 calls each, meas...",https://github.com/pytorch/pytorch/pull/186898,f997fe2e449395588094c3a7e8b3b090e159dc1def315a0329d5725dbc356588 closes,pr,185866,issue,185589,high,pr.body,ion import shim ... PY; ran 3 tests OK. lintrunner -a; ok No lint issues. Fresh review subagents; final quick review found no issues. Fixes #185589 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/185866,c8c2390d47d6c145b4f95b161c40716939d0150facb21661cab8806b598e23f1 closes,pr,186074,issue,141017,high,pr.body,"ch.compile backend=aot_eager regression covering the reported batched mismatch across reductions and the unbatched scalar shape case. Fixes #141017 Generated by my agent Benchmark Results: CPU microbenchmark for torch.compile backend=aot_eager with reduction=none, N=64, C=128...",https://github.com/pytorch/pytorch/pull/186074,c035146fe9f343ece818295ddd993e9df70f9f6efc90b573a91cd2a26e1fe185 closes,pr,186078,issue,141005,high,pr.body,eps user-facing direct access consistent with _orig_mod while containing the inherited nn.Module paths that need wrapper-local state. Fixes #141005 Generated by my agent Test Plan: python test/dynamo/test_modules.py -k 'OptimizedModuleTest.test_parameters_attr' python test/dyn...,https://github.com/pytorch/pytorch/pull/186078,0e95751180e7bd152bda663359e8ca11e67190c58c60605027feab0b8e16e262 closes,pr,184460,issue,102870,high,pr.body,"e already passed as explicit graph inputs. Keep the C++ wrapper path unchanged, since it still owns its size/stride binding behavior. Fixes #102870 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...",https://github.com/pytorch/pytorch/pull/184460,c01ccdf332ef453c715d8381fa7e2abcd72a51a424eff6973e02e254e530ced7 closes,pr,185749,issue,149895,high,pr.body,"ynamo side effects, but Dynamo's own helper checks no longer accidentally observe the uninitialized backing object through user code. Fixes #149895 Generated by my agent Test Plan: python test/dynamo/test_user_defined_object.py -k test_instantiate python test/dynamo/test_user_...",https://github.com/pytorch/pytorch/pull/185749,99a3ba64b4860c3e9e0fc8048f8f6ff078710044d8a9b38aad730251675ce96a closes,pr,188556,issue,188545,high,pr.closingIssuesReferences,pr #188556 declares a closing reference to issue #188545.,https://github.com/pytorch/pytorch/pull/188556,e3c9cf7d85e06e307095fd5c949da10e39cc8398c7b43c407e59b269a56e2102 closes,pr,188556,issue,188545,high,pr.body,Fixes #188545 Summary Preserve eager-mode semantics for infinite inputs in the Triton lowering of torch.special.bessel_j0 torch.special.bessel_j1 torch.s,https://github.com/pytorch/pytorch/pull/188556,34f15e53b6aecea93c0ba3a4c691fd71c5356fd08b8e37b5ebb81c39079b02b8 closes,pr,187353,issue,187336,high,pr.body,"ng on compiler unary-minus behavior, but it preserves signed zero and only adds a negligible cost in a memory-bound pointwise kernel. Fixes #187336 Generated by my agent Benchmark Results: Compiled CUDA fp32 unary neg on NVIDIA B200, 67,108,864 elements, 200 timed iterations,...",https://github.com/pytorch/pytorch/pull/187353,c83e322a5553bb23735d0e65198d5bc261b8106a3d22537ebd7e4f9074bfd0d0 closes,pr,185333,issue,160247,high,pr.body,"kend=eager, 10 iterations: before median_ms 0.715 and median_peak_extra_mib 8.0; after median_ms 0.313 and median_peak_extra_mib 1.0. Fixes #160247 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...",https://github.com/pytorch/pytorch/pull/185333,15ceaf9bdab16cd3c46a5a8bc5833573f8c4ec254545dc21db22309e8523a184 review guidance,pr,185333,issue,160247,high,pr.reviews[0].body,"I think this needs discussion. we're turning something that used to be a full graph into a graph break. I get the motivation, but this is usually not the direction we go in unless there's silent incorrectness",https://github.com/pytorch/pytorch/pull/185333,c3fc274c797c1234ea7e22faf0f8dafaa673f1b950219633b220f227466c4153 review guidance,pr,185333,pr,185333,high,pr.reviews[0].body,"I think this needs discussion. we're turning something that used to be a full graph into a graph break. I get the motivation, but this is usually not the direction we go in unless there's silent incorrectness",https://github.com/pytorch/pytorch/pull/185333,b7c9640ba3d127572e3f16f528bbc876f137033d7f52e20681c926ea47c2fa06 review guidance,pr,185333,issue,160247,high,pr.reviews[1].body,"My agent says I addressed this by tightening the graph break to the actual lifetime hazard: it now only triggers when a mutation displaces an existing tensor entry from a pre-existing list, while tensor assignment into non-tensor slots can still compile fullgraph. I also added regression coverage...",https://github.com/pytorch/pytorch/pull/185333,0316348a1f5b1ba8f75d2958e4495c76f5dd5359067dbcb17f0aac0a663d6875 review guidance,pr,185333,pr,185333,high,pr.reviews[1].body,"My agent says I addressed this by tightening the graph break to the actual lifetime hazard: it now only triggers when a mutation displaces an existing tensor entry from a pre-existing list, while tensor assignment into non-tensor slots can still compile fullgraph. I also added regression coverage...",https://github.com/pytorch/pytorch/pull/185333,2c0fc7996eb9a8d3fccb3d97c3980b4e949f7601c7dfd839af7e3d178c5f401e review guidance,pr,185333,issue,160247,high,pr.reviews[2].body,"I think this needs discussion. we're turning something that used to be a full graph into a graph break. I get the motivation, but this is usually not the direction we go in unless there's silent incorrectness",https://github.com/pytorch/pytorch/pull/185333#pullrequestreview-4480282157,3f07cd63d2edf037d96b8638952dcebc53e7ebe3a3253df5e44573845f9bd1f8 review guidance,pr,185333,pr,185333,high,pr.reviews[2].body,"I think this needs discussion. we're turning something that used to be a full graph into a graph break. I get the motivation, but this is usually not the direction we go in unless there's silent incorrectness",https://github.com/pytorch/pytorch/pull/185333#pullrequestreview-4480282157,6169cdeff4f5a45d1a21ae2fcdd172e08fbf9c2fd644277492c03c24017df4d9 review guidance,pr,185333,issue,160247,high,pr.reviews[3].body,"My agent says I addressed this by tightening the graph break to the actual lifetime hazard: it now only triggers when a mutation displaces an existing tensor entry from a pre-existing list, while tensor assignment into non-tensor slots can still compile fullgraph. I also added regression coverage...",https://github.com/pytorch/pytorch/pull/185333#pullrequestreview-4490968881,fea015fbc402421d212c6f0a43a443739101c0f53c5350cd38b39255cd1414f2 review guidance,pr,185333,pr,185333,high,pr.reviews[3].body,"My agent says I addressed this by tightening the graph break to the actual lifetime hazard: it now only triggers when a mutation displaces an existing tensor entry from a pre-existing list, while tensor assignment into non-tensor slots can still compile fullgraph. I also added regression coverage...",https://github.com/pytorch/pytorch/pull/185333#pullrequestreview-4490968881,472bc40bb773365f39ae023c553429133750031c25506cd4a5be842782c49a8f closes,pr,186094,issue,137520,high,pr.body,"g only actually used unbacked symbol definitions, instead of pinning every node that defines an unbacked symbol in DCE. Fixes #140842 Fixes #137520 Generated by my agent Benchmark Results: A small Python microbenchmark compared the previous scanner predicate (SymTypes, Expr) w...",https://github.com/pytorch/pytorch/pull/186094,34c3b6357c006647557d957d73edac26919a8e88e916b41c94e6ab41df3c2aa0 closes,pr,186094,issue,140842,high,pr.body,"for preserving only actually used unbacked symbol definitions, instead of pinning every node that defines an unbacked symbol in DCE. Fixes #140842 Fixes #137520 Generated by my agent Benchmark Results: A small Python microbenchmark compared the previous scanner predicate (SymT...",https://github.com/pytorch/pytorch/pull/186094,be781e47e3050c0f6a717600ef1c25ca1a7aff3b1f18d73675213f1b2ea80be4 references,pr,186094,issue,186031,medium,pr.comments[4].body,"My agent says the remaining torchtitan_features_integration failure is the known torchcomms FSDP+TP+PP+compile issue tracked in #186031, with the same `Could not resolve the process group registered under the name` traceback. This is unrelated to this AOTI/Inductor symbol de",https://github.com/pytorch/pytorch/pull/186094,835fecfae99f2ad74b96eda89017abbd92f375d39a31b27684429ef8b2f8a27a review guidance,pr,186094,issue,137520,high,pr.reviews[0].body,Nice,https://github.com/pytorch/pytorch/pull/186094,737e35a800103ef5628d4ac85dbb1afdb6673050cb5fc94b3e0008c7c653d680 review guidance,pr,186094,issue,140842,high,pr.reviews[0].body,Nice,https://github.com/pytorch/pytorch/pull/186094,0c72788290be7a0b18aa31212cdf050c99b8748a9df882f1fc9b39114e9d1280 review guidance,pr,186094,pr,186094,high,pr.reviews[0].body,Nice,https://github.com/pytorch/pytorch/pull/186094,bb668e7aa6e9e83b762b3ec65bbbf867bb05bed3d8e5ad9e5d778cc2f3cc9224 review guidance,pr,186094,issue,137520,high,pr.reviews[2].body,Nice,https://github.com/pytorch/pytorch/pull/186094#pullrequestreview-4459917208,b9f4a82770e2f5653c53867546a9d08ad794f762f7fcf0e4f55418e3ea416685 review guidance,pr,186094,issue,140842,high,pr.reviews[2].body,Nice,https://github.com/pytorch/pytorch/pull/186094#pullrequestreview-4459917208,877eaef795fee2e6185c6675d13a90d4414dbee120c32cf95172850c1659d6b4 review guidance,pr,186094,pr,186094,high,pr.reviews[2].body,Nice,https://github.com/pytorch/pytorch/pull/186094#pullrequestreview-4459917208,1dd18a873e7366ef9f8fc26577365b8b25353ef6a9ceea58bc997a330cd63694 closes,pr,184207,issue,140795,high,pr.body,"ssary host memory after compilation. Add regression coverage for module unloading, cleanup on exceptions, and precompile-cache reuse. Fixes #140795 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...",https://github.com/pytorch/pytorch/pull/184207,3e4a8208598eb8ce0531ce1b9f69199a5897bce20a8eb07438ffa6c6b0733916 closes,pr,186111,issue,140707,high,pr.body,returned output is later differentiated. This change stays scoped to the original no-alias failure and explicit alias-output safety. Fixes #140707 Generated by my agent Test Plan: python test/functorch/test_codegen_runtime_wrapper.py TestCodegenRuntimeWrapper.test_training_non...,https://github.com/pytorch/pytorch/pull/186111,ac2382ced7d3070c403f82873fc00895b7e90b2255194c9ac06ce9dfaf782dbb review guidance,pr,187931,pr,187931,high,pr.reviews[0].body,"Pull request overview This PR refines XPU graph capture detection in XPUCachingAllocator by separating “allocation routed to a private pool” from “a real SYCL command-graph capture is actively recording,” and wires the new capture markers into XPUGraph. Changes: Add markCaptureBegin/markCaptureEn...",https://github.com/pytorch/pytorch/pull/187931,512539bcf6c22bb741c1e7582fd3fb609b0c480cb41ac808e5bd88b66aab3d20 closes,pr,186118,issue,140683,high,pr.body,test_frame_init.py python -m py_compile test/dynamo/test_frame_init.py && git diff --check && git diff --cached --check lintrunner -a Fixes #140683 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/186118,f7f32b7b16bebc6082bf0bea9a856788788c94df03bd710489224449048c6f57 closes,pr,186900,issue,140615,high,pr.body,"nsidered, but the root failure is in shared TreeSpec deserialization and the same guidance applies to any serialized pytree consumer. Fixes #140615 Generated by my agent Test Plan: python test/inductor/test_aot_inductor_package.py -k custom_output_type_missing_pytree_registrat...",https://github.com/pytorch/pytorch/pull/186900,35a85bcafddb125ab6a2707b4e3925b65d3e1fb8ba2a27af57acf919151384a6 closes,pr,186942,issue,186875,high,pr.body,"opt-in and enable it for fmod and remainder. This keeps mul/div unchanged, since eager uses opmath precision for those scalar paths. Fixes #186875 Generated by my agent Test Plan: TORCHINDUCTOR_FORCE_DISABLE_CACHES=1 python heredoc comparing eager vs torch.compile fmod/remaind...",https://github.com/pytorch/pytorch/pull/186942,3e38a4a20e0a8e5f1418cefc9b3027f3cf098a288af3fff31ffe0d83f90d2189 references,pr,186942,issue,186875,medium,pr.comments[8].body,"riants. I removed that assertion while keeping the exact `atol=0, rtol=0` regression check that catches the scalar-rounding bug reported in #186875. Targeted CPU/CUDA base, dynamic-shape, and codegen dynamic variants pass locally with the minimal torchvision stub; `lintrunner...",https://github.com/pytorch/pytorch/pull/186942,243476f9824fdcb7ad63bca2f4391c32a1d7934b98e336a726d3a302db87ef51 closes,pr,186119,issue,140607,high,pr.body,rt regression test that reproduces the malformed positional-args path without downloading the external model from the original issue. Fixes #140607 Generated by my agent Test Plan: python test/dynamo/test_error_messages.py ErrorMessagesTest.test_unpack_sequence_with_wrong_leng...,https://github.com/pytorch/pytorch/pull/186119,86ab29ea2ef99d485df57c3d525e0714b0b50b6d73c02d6eb523cc2d5cb05593 closes,pr,186160,issue,139756,high,pr.body,"tempted to render the dual tensor, which unpacked forward AD metadata through native::_fw_primal and hit the internal assertion reported in #139756. Allow plain meta tensors in this path to continue through the existing FakeTensorConverter, which already supports fakifying met...",https://github.com/pytorch/pytorch/pull/186160,458a431b75391016e3864e2f7ecca9cbb95c84aab522f58608506f810125de0e review guidance,pr,186160,issue,139756,high,pr.reviews[0].body,"seems like something else is very wrong, forward ad should be producing FakeTensors",https://github.com/pytorch/pytorch/pull/186160,1c8d6dc4f313f7396b065df9be5a661ac677dad992e0da309481778b418b4ba7 review guidance,pr,186160,pr,186160,high,pr.reviews[0].body,"seems like something else is very wrong, forward ad should be producing FakeTensors",https://github.com/pytorch/pytorch/pull/186160,843c34ece623549c35ea50f00284de5c86d94759ca893027a364dd1846959444 review guidance,pr,186160,issue,139756,high,pr.reviews[1].body,"seems like something else is very wrong, forward ad should be producing FakeTensors",https://github.com/pytorch/pytorch/pull/186160#pullrequestreview-4439033938,2383bc65682d55c73e852fb22d95f554ced18d3ac1c845f6fc6b4fec4c2a2172 review guidance,pr,186160,pr,186160,high,pr.reviews[1].body,"seems like something else is very wrong, forward ad should be producing FakeTensors",https://github.com/pytorch/pytorch/pull/186160#pullrequestreview-4439033938,b331c6c887dd84a1b0696d57526f045365bfdbab452cf109b125da70ae044764 closes,pr,186166,issue,139707,high,pr.body,ed recursive device dispatch. Redirecting only the Python registration path keeps the change scoped to the problematic dispatch case. Fixes #139707 Generated by my agent Test Plan: ninja -C build torch_python && cp build/lib/libtorch_python.so torch/lib/libtorch_python.so pyth...,https://github.com/pytorch/pytorch/pull/186166,27f9b6f91831d140d55ba0139990c02a1720404a62ddad73682f5b668eb6aa8a closes,pr,186044,issue,142489,high,pr.body,"ore patterns, but it would also be noisy and harder to maintain; the syntactic rule covers the observed root-cause patterns directly. Fixes #142489 Generated by my agent Test Plan: python -m unittest tools.test.test_import_linter python tools/linter/adapters/import_linter.py -...",https://github.com/pytorch/pytorch/pull/186044,f8cb63012b85e849545d133713df95f188e2b691b3473b79086e2975285fb4e7 closes,pr,186168,issue,139603,high,pr.body,lse so it exercises a set of cuda-device values without requiring CUDA; the old failure occurs before fork_rng can branch on enabled. Fixes #139603 Generated by my agent Test Plan: python test/dynamo/test_repros.py -k test_fork_rng_with_set_devices Manual repro before fix fail...,https://github.com/pytorch/pytorch/pull/186168,1c6950cd6a79101e481cd5a215426b877aed7410b6519d3939b6b0668c13f560 closes,pr,186180,issue,137869,high,pr.body,d. Checking the profiler slot first makes the external-profiler case explicit while preserving nested Dynamo profiling. Fixes #139232 Fixes #137869 Generated by my agent Test Plan: TORCH_COMPILE_CPROFILE=1 python - <<'PY' ... cProfile.Profile().runcall(run) ... PY TORCH_COMPIL...,https://github.com/pytorch/pytorch/pull/186180,9fb8586c021873e12bdaf490e7fa8e5bcfe4804a2bba26aac19d9560f71ab3e7 closes,pr,186180,issue,139232,high,pr.body,is unsupported. Checking the profiler slot first makes the external-profiler case explicit while preserving nested Dynamo profiling. Fixes #139232 Fixes #137869 Generated by my agent Test Plan: TORCH_COMPILE_CPROFILE=1 python - <<'PY' ... cProfile.Profile().runcall(run) ... PY...,https://github.com/pytorch/pytorch/pull/186180,ac09d538b01825a5f7c6faf42a176c0abd71e9d259401cca181e52f88f50b87d closes,pr,186188,issue,139092,high,pr.body,"the requested fake device. Add focused FakeTensor coverage for default meta, explicit meta, and explicit meta:0 tensor construction. Fixes #139092 Generated by my agent Test Plan: ninja -C build torch_python && ninja -C build install python reproduction script for FakeTensorMo...",https://github.com/pytorch/pytorch/pull/186188,5131a648fc24dcf5b282da9fe8cb5c491cbd25473a5a004a8851104df37f88d9 closes,pr,186920,issue,184195,high,pr.body,"guments, and omitted Tensor returns, and fix the dispatcher bridge to match popped boxed returns with the correct schema return type. Fixes #184195 Generated by my agent Test Plan: python torchgen/gen.py --update-aoti-c-shim (passed) python -m py_compile torchgen/aoti/fallback...",https://github.com/pytorch/pytorch/pull/186920,ef45180d8d8b5f43f3cb27b53a5d4b839a7e1d08bd980e22a4cb4a20bd503595 review guidance,pr,182725,pr,182725,high,pr.reviews[0].body,Can we make this a bit more general rather than putting this logic into tcpstore? I'm thinking something similar to glog LOG_EVERY_T,https://github.com/pytorch/pytorch/pull/182725,d7eed39fb2fe14d15ff91f1b0f09cef0646f3f290e19430e5b5e831a2fe66067 closes,pr,184367,issue,184002,high,pr.closingIssuesReferences,pr #184367 declares a closing reference to issue #184002.,https://github.com/pytorch/pytorch/pull/184367,d0979f9075807070a3217edc34428852054ad612e2fa6fee0e927f0aa12244ed closes,pr,184367,issue,184002,high,pr.body,Fixes #184002 Compiled DTensor view outputs no longer fall back to unsupported as_strided. Added test to cover both generic wrapper subclass case and DTe,https://github.com/pytorch/pytorch/pull/184367,43f85f6725b61efae2a5ddb946c18a91c6853d0e7af204c398a682714900aee6 closes,pr,178844,issue,178798,high,pr.closingIssuesReferences,pr #178844 declares a closing reference to issue #178798.,https://github.com/pytorch/pytorch/pull/178844,d4a8107a88f5e9cdecc649ed746ef632052808f86078e532a793948e454dba79 closes,pr,178844,issue,178798,high,pr.body,Summary Fixes #178798 ProcessGroupGloo::_allgather_base chunked the output tensor along dim 0 and required every chunk to match the input shape exactly. For stac,https://github.com/pytorch/pytorch/pull/178844,89a86cd4fc93f78e14b71dfb23c3a602db7ff84e51f716a98e1acf581a71ad9c closes,pr,186192,issue,138961,high,pr.body,"saved_compile_context -q lintrunner -a Original CUDA multithreaded repro with torch.compile(..., options={""triton.cudagraphs"": True}) Fixes #138961 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...",https://github.com/pytorch/pytorch/pull/186192,40f77b693f047fdfd7bc3e7e852b26bdd84dfdec2a20396ceec672efa09235c9 closes,pr,186194,issue,138926,high,pr.body,"Stack from ghstack (oldest at bottom): -> #186194 Issue #138926 reported a torch.cond metadata mismatch where one branch returned a tensor sized with TruncToInt(IntTrueDiv(s, 1)) while the other returned",https://github.com/pytorch/pytorch/pull/186194,df258de5cd11ccf168c81dc512592ac879743efb5b68874b54b9295fa933b820 references,pr,188939,pr,187898,medium,pr.body,Stack from ghstack (oldest at bottom): #187898 -> #188939 #188019 #188849 #188942 Add two more counter sources to the chrome-JSON exporter's GPU counters (built on the plumbing from the,https://github.com/pytorch/pytorch/pull/188939,924a77e6f99fed5babc87c2ff9f90df5cf661e6bb00997d4b566d591dcbe2b7e competes with,pr,187898,pr,188939,medium,pr.body,Stack from ghstack (oldest at bottom): -> #187898 #188939 #188019 #188849 #188942 export_chrome_trace(path) with a .pftrace path makes the cupti_monitor backend emit a Perfetto-native trace instead,https://github.com/pytorch/pytorch/pull/187898,520644dc1fc77170a8c5616a0f01dc431bc53cc2a4458cf0334153a770fbe754 closes,pr,184256,issue,138301,high,pr.body,ction heuristic XBLOCK growth to the configured Triton maximum so generated configs preserve the indexing invariants for large grids. Fixes #138301 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/184256,4eff5fa957297fbdcb46066101db224b9056b3fc519ce99e93be4e3a280594a1 review guidance,pr,188867,pr,188867,high,pr.reviews[0].body,"Pull request overview NoteCopilot couldn't run its full agentic review because no GitHub Actions runner was available. Make sure your repository has a runner available to run Copilot's review, or add a copilot-setup-steps.yml file specifying one with the runs-on attribute. See the docs for more d...",https://github.com/pytorch/pytorch/pull/188867,b90f3c4beb36311071e71abc15d1fcfa4ecf09137766c251f8b85ddc1c023188 closes,pr,186199,issue,138264,high,pr.body,e trailing functional returns. This keeps the fix in the shared functionalization codegen path rather than special-casing batch norm. Fixes #138264 Generated by my agent Test Plan: ninja -C build torch_python ninja -C build install python - <<'PY' ... torch.compile(torch._nati...,https://github.com/pytorch/pytorch/pull/186199,43a8537b0fae85e693b32c62d5d475217657e2caa96e59530f6edb23f6a60b11 closes,pr,186201,issue,138229,high,pr.body,Stack from ghstack (oldest at bottom): -> #186201 Fixes #138229 Generated by my agent AOT/export runtime assertion insertion can CSE symbolic size nodes that represent the same symbolic dimension. In the,https://github.com/pytorch/pytorch/pull/186201,fac1ce91af33c0ad7c5d311129330e1698f3e1759b39c4a8ccb6e820a78f7aa4 closes,pr,185782,issue,113040,high,pr.closingIssuesReferences,pr #185782 declares a closing reference to issue #113040.,https://github.com/pytorch/pytorch/pull/185782,199a4adc477f4949279a1b21e00a2ab874991e5ea499ea4e1ce69721e03372fc closes,pr,185782,issue,113040,high,pr.body,aN values. Write the recompiles column to the expected CSV when it is present in the downloaded data. Updated module docstring. Issue Fixes #113040 Issue URL: #113040 Changes benchmarks/dynamo/check_graph_breaks.py | 41 +++++++++++++++++++--- benchmarks/dynamo/ci_expected_accu...,https://github.com/pytorch/pytorch/pull/185782,e94efb3373602306d9ea4c2a0ba3d3c2e34e837bd79897aade7796dd7503c621 review guidance,pr,185782,issue,113040,high,pr.reviews[0].body,"Please make sure that the expected csv files are also updated with the number of recompiles. As a followup, I think it would be good if we also tracked the number of fallbacks to eager; this is because additional graph breaks due to more code traced/fewer fallbacks to eager is acceptable.",https://github.com/pytorch/pytorch/pull/185782,616c0f0c0f3f79682d7c75c9c4eb00327b8b2643509e56a3f45f8d4c47e2630c references,pr,185486,issue,181093,medium,pr.body,s eager test_inductor_matches_eager — Fused mul+add matches eager test_inductor_output_device — Output stays on openreg device Part of RFC: #181093,https://github.com/pytorch/pytorch/pull/185486,4d9cd5074c62b10de6c0bd5d77e310d68a26f605e9df68db01aca1b52f2649b7 closes,pr,186501,issue,119191,high,pr.body,"ion, stream assignment, and partitioner mutation handling to treat aten.foreach_copy as the grouped form of epilogue input mutations. Fixes #119191 Generated by my agent Benchmark Results: A 100-tensor foreach input-mutation micro-benchmark compared the current fold with graph...",https://github.com/pytorch/pytorch/pull/186501,e8464a9bee129d34df946ab7e745d91a5b940cdb227a190d7e0b5e8294196972 review guidance,pr,186501,issue,119191,high,pr.reviews[0].body,historically we only introduce non copy_ mutating ops after a certain point in post grad. so i would be worried about introducing this and not auditing other passes etc. or just do this optimization in post grad mutating phase.,https://github.com/pytorch/pytorch/pull/186501,ae289e5eda9cf73a4ea83bb93f5102e15818746e9c9d16c3a7d44faa138afd07 review guidance,pr,186501,pr,186501,high,pr.reviews[0].body,historically we only introduce non copy_ mutating ops after a certain point in post grad. so i would be worried about introducing this and not auditing other passes etc. or just do this optimization in post grad mutating phase.,https://github.com/pytorch/pytorch/pull/186501,bb12789fbbd605296ad8773941e12ebde59733acd51706f62eedfc29c65f8d6a review guidance,pr,186501,issue,119191,high,pr.reviews[1].body,"My agent says I moved the foreach-copy fold out of AOTAutograd and into Inductor's late post-grad mutating section in the latest push. AOT validation, stream assignment, and partitioning now keep seeing plain copy_ epilogues, and the new tests cover both that invariant and the late Inductor fold.",https://github.com/pytorch/pytorch/pull/186501,69e92d082414fbf13f25d4bb50d501b35f365d69f81cbbc2df4099a85253f2fa review guidance,pr,186501,pr,186501,high,pr.reviews[1].body,"My agent says I moved the foreach-copy fold out of AOTAutograd and into Inductor's late post-grad mutating section in the latest push. AOT validation, stream assignment, and partitioning now keep seeing plain copy_ epilogues, and the new tests cover both that invariant and the late Inductor fold.",https://github.com/pytorch/pytorch/pull/186501,888860025cdf0d58012a781747b9366284a52fc679f088e2e5f6c2afe8de31c6 review guidance,pr,186501,issue,119191,high,pr.reviews[2].body,historically we only introduce non `copy_` mutating ops after a certain point in post grad. so i would be worried about introducing this and not auditing other passes etc. or just do this optimization in post grad mutating phase.,https://github.com/pytorch/pytorch/pull/186501#pullrequestreview-4545882196,6f51dde0f9ddde33561282cfb9d3ae1728eb37cef37227a94794a75f6718911a review guidance,pr,186501,pr,186501,high,pr.reviews[2].body,historically we only introduce non `copy_` mutating ops after a certain point in post grad. so i would be worried about introducing this and not auditing other passes etc. or just do this optimization in post grad mutating phase.,https://github.com/pytorch/pytorch/pull/186501#pullrequestreview-4545882196,ba28c78e8e7fc080d0129108b540199179eb26bf043e59ac178eaf2c621222e7 review guidance,pr,186501,issue,119191,high,pr.reviews[3].body,"My agent says I moved the foreach-copy fold out of AOTAutograd and into Inductor's late post-grad mutating section in the latest push. AOT validation, stream assignment, and partitioning now keep seeing plain copy_ epilogues, and the new tests cover both that invariant and the late Inductor fold.",https://github.com/pytorch/pytorch/pull/186501#pullrequestreview-4552418358,eca0663401e7ba85552c4512d0fceec74d457dc1743c53e00c6d20447fd76386 review guidance,pr,186501,pr,186501,high,pr.reviews[3].body,"My agent says I moved the foreach-copy fold out of AOTAutograd and into Inductor's late post-grad mutating section in the latest push. AOT validation, stream assignment, and partitioning now keep seeing plain copy_ epilogues, and the new tests cover both that invariant and the late Inductor fold.",https://github.com/pytorch/pytorch/pull/186501#pullrequestreview-4552418358,19e5a3804e5f404ed072ab83625ef88b6e8a4adfb93df305c3d657ca04bec8cc closes,pr,186210,issue,138141,high,pr.body,"eeds the existing true-branch clone workaround for cond output aliasing, but the branch-local shape assertion now compiles correctly. Fixes #138141 Generated by my agent Test Plan: python test/inductor/test_control_flow.py -k cond_shape_assert python test/test_dynamic_shapes.p...",https://github.com/pytorch/pytorch/pull/186210,4d5d9af1a38ba610a0a5bc05019e6a0e72fe9b349a2a0700a7f0b55f733e256d closes,pr,186214,issue,137996,high,pr.body,"aring the existing helper keeps forward and backward profile naming, logging, SVG generation, and fb-only upload behavior consistent. Fixes #137996 Generated by my agent Test Plan: TMPDIR=$(mktemp -d) TORCH_COMPILE_CPROFILE=1 python - <<'PY' ... torch.compile(fn, backend='indu...",https://github.com/pytorch/pytorch/pull/186214,1f1c2b577664f6debd5e6f678c78bf692a31baceac4d13732932dd1c120fc58c closes,pr,186217,issue,137875,high,pr.body,cache metadata and tlparse artifact path. This keeps cache miss debugging available without flooding the regular Inductor debug log. Fixes #137875 Generated by my agent Test Plan: python test/inductor/test_codecache.py TestFxGraphCacheHashing.test_compiled_fx_graph_hash_detail...,https://github.com/pytorch/pytorch/pull/186217,7428bec132aeba80f851d8dbe635336c1bdd01776ea502af46d0bc1ee932719d closes,pr,187604,issue,185185,high,pr.body,ng slice/select semantics and avoids the broader alternative of introducing Piecewise/sym_ite expressions for ambiguous slice bounds. Fixes #185185 Generated by my agent Test Plan: python test/test_dynamic_shapes.py TestUbackedOps.test_tensor_derived_multidim_slice_bounds -v p...,https://github.com/pytorch/pytorch/pull/187604,66cab054dfd0b1e8d1c6e9e1bdbaf51a3613c04961ce98dd40817db905b86663 closes,pr,186222,issue,137388,high,pr.body,population traversal rather than changing epilogue guard storage because the latter would affect guard ordering and runtime behavior. Fixes #137388 Generated by my agent Benchmark Results: Command compared 5 repeats of 30 torch.compile calls for the issue's dynamic-shape repro...,https://github.com/pytorch/pytorch/pull/186222,0d64925cbbfe2f52b7c587f077fc4c75cca8f6c85dc331fa210b1fe9afff1add closes,pr,186225,issue,137275,high,pr.body,"in the generic AOT subclass output path, so threading the nested-int metadata there fixes the root cause for the reconstruction path. Fixes #137275 Generated by my agent Test Plan: python - <<'PY' import torch def f(nt): return nt.to(device=""cpu"") compiled_f = torch.compile(f)...",https://github.com/pytorch/pytorch/pull/186225,1ad0053fb531d58817bd5ceac8cd668b5fe483094259c3db038db5fd04ac9d9d closes,pr,184621,issue,182649,high,pr.body,m raw FakeTensor scalar values. This covers tensor-backed reshape shapes from buffers and equivalent ONNX-converted reshape patterns. Fixes #182649 Fixes #182651 Fixes #182652 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zh...,https://github.com/pytorch/pytorch/pull/184621,6d2a9357f5807fd93ce9a0dc285d69cc10687e948c343e92bed62830dddf529c closes,pr,184621,issue,182651,high,pr.body,or scalar values. This covers tensor-backed reshape shapes from buffers and equivalent ONNX-converted reshape patterns. Fixes #182649 Fixes #182651 Fixes #182652 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzhe...,https://github.com/pytorch/pytorch/pull/184621,30a2d2b755d673170eedc7a8b59b1ce51fb2521309e05bcfa597f4d056f10a53 closes,pr,184621,issue,182652,high,pr.body,es. This covers tensor-backed reshape shapes from buffers and equivalent ONNX-converted reshape patterns. Fixes #182649 Fixes #182651 Fixes #182652 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/184621,9c1e5df84bdf265a9427289d8f2e86643f7d63f1f1ea60ab6d7d3451c028c883 review guidance,pr,184621,issue,182649,high,pr.reviews[0].body,"Missing test coverage. According to claude, these tests should be added: - A -1 dimension test (and decide: fix it for fullgraph, or make it a clean graph break — don't leave the u >= 0 assertion crash). - A test built on the onnx2torch-style structure (fx GraphModule with get_attr shape + torch....",https://github.com/pytorch/pytorch/pull/184621,67bcfa73529aa8762a80178260b1e35aacda9796d72598cf52c5ba2d046ca8e1 review guidance,pr,184621,issue,182651,high,pr.reviews[0].body,"Missing test coverage. According to claude, these tests should be added: - A -1 dimension test (and decide: fix it for fullgraph, or make it a clean graph break — don't leave the u >= 0 assertion crash). - A test built on the onnx2torch-style structure (fx GraphModule with get_attr shape + torch....",https://github.com/pytorch/pytorch/pull/184621,049f3d8c4a43e901cf40142a90925e9e8720781df37baeb488e60a79714485f5 review guidance,pr,184621,issue,182652,high,pr.reviews[0].body,"Missing test coverage. According to claude, these tests should be added: - A -1 dimension test (and decide: fix it for fullgraph, or make it a clean graph break — don't leave the u >= 0 assertion crash). - A test built on the onnx2torch-style structure (fx GraphModule with get_attr shape + torch....",https://github.com/pytorch/pytorch/pull/184621,611d419c853dc3f7492c7f21ed60e3ca32b9908e311a60dfaa50ea4c94ddfeb8 review guidance,pr,184621,pr,184621,high,pr.reviews[0].body,"Missing test coverage. According to claude, these tests should be added: - A -1 dimension test (and decide: fix it for fullgraph, or make it a clean graph break — don't leave the u >= 0 assertion crash). - A test built on the onnx2torch-style structure (fx GraphModule with get_attr shape + torch....",https://github.com/pytorch/pytorch/pull/184621,35b33e1a5c5e7251b7f1c55da6cd6f9a9f8f3d2d1549cbd390aa570ffaab4678 review guidance,pr,184621,issue,182649,high,pr.reviews[1].body,letting william take this,https://github.com/pytorch/pytorch/pull/184621,6d0913c3fb13c22f156087e51dd3297a5ae0f7acbed1896f83397356ad4aaade review guidance,pr,184621,issue,182651,high,pr.reviews[1].body,letting william take this,https://github.com/pytorch/pytorch/pull/184621,e719284bd5f464606675d2b275881dd1b67e558f1a6d5d1f1940b1df5987cbe2 review guidance,pr,184621,issue,182652,high,pr.reviews[1].body,letting william take this,https://github.com/pytorch/pytorch/pull/184621,a1db51ad53964f8e6981765e8b2d61861a983e8959b90c0e4f3a1e5d468d07d5 review guidance,pr,184621,pr,184621,high,pr.reviews[1].body,letting william take this,https://github.com/pytorch/pytorch/pull/184621,a3577bc5f399673c8b3474e4ee365185e409992caa2e8129dab27ab560c9c4c6 review guidance,pr,184621,issue,182649,high,pr.reviews[2].body,"Missing test coverage. According to claude, these tests should be added: ``` - A -1 dimension test (and decide: fix it for fullgraph, or make it a clean graph break — don't leave the u >= 0 assertion crash). - A test built on the onnx2torch-style structure (fx GraphModule with get_attr shape + to...",https://github.com/pytorch/pytorch/pull/184621#pullrequestreview-4453591312,41ea35576f2d5f2e315a02d1a488f856c3e8c896c0df376242ff386fe171a5f9 review guidance,pr,184621,issue,182651,high,pr.reviews[2].body,"Missing test coverage. According to claude, these tests should be added: ``` - A -1 dimension test (and decide: fix it for fullgraph, or make it a clean graph break — don't leave the u >= 0 assertion crash). - A test built on the onnx2torch-style structure (fx GraphModule with get_attr shape + to...",https://github.com/pytorch/pytorch/pull/184621#pullrequestreview-4453591312,9f4c7f5ec3050b6d1994841e6288dc24f6786e5561b1076efabc2feca4f727b8 review guidance,pr,184621,issue,182652,high,pr.reviews[2].body,"Missing test coverage. According to claude, these tests should be added: ``` - A -1 dimension test (and decide: fix it for fullgraph, or make it a clean graph break — don't leave the u >= 0 assertion crash). - A test built on the onnx2torch-style structure (fx GraphModule with get_attr shape + to...",https://github.com/pytorch/pytorch/pull/184621#pullrequestreview-4453591312,b2c6243031474804702b9f8cbd35fbcd801180cd40aefa91a366b0fe528f5f2c review guidance,pr,184621,pr,184621,high,pr.reviews[2].body,"Missing test coverage. According to claude, these tests should be added: ``` - A -1 dimension test (and decide: fix it for fullgraph, or make it a clean graph break — don't leave the u >= 0 assertion crash). - A test built on the onnx2torch-style structure (fx GraphModule with get_attr shape + to...",https://github.com/pytorch/pytorch/pull/184621#pullrequestreview-4453591312,f1ff4c103fdf4035757e568158f62c68d6a9244563b5e703d2659f0dee7e3812 review guidance,pr,184621,issue,182649,high,pr.reviews[3].body,letting william take this,https://github.com/pytorch/pytorch/pull/184621#pullrequestreview-4472664088,a953a0e291ccb86c037d1dccff0acc720aa612fe11d0c83b2ef20192442b76bf review guidance,pr,184621,issue,182651,high,pr.reviews[3].body,letting william take this,https://github.com/pytorch/pytorch/pull/184621#pullrequestreview-4472664088,1c70e11c7b16c6a65f9625233a69e4f45be3b25e1e6831400a1999cf326f5bd8 review guidance,pr,184621,issue,182652,high,pr.reviews[3].body,letting william take this,https://github.com/pytorch/pytorch/pull/184621#pullrequestreview-4472664088,828e1da96e3798d0bb7abb1b3963d4a59dd6732c091dd006a43714477bec13ec review guidance,pr,184621,pr,184621,high,pr.reviews[3].body,letting william take this,https://github.com/pytorch/pytorch/pull/184621#pullrequestreview-4472664088,4e6358206c61674cf34c879027282db571ede56949f1d6ee18c8a34bd2cf91ed closes,pr,186249,issue,137068,high,pr.body,"a backend-visible op because the value is nondeterministic Python state, not a tensor operation that should require lowering support. Fixes #137068 Generated by my agent Test Plan: python test/dynamo/test_unspec.py -k time_time python test/dynamo/test_unspec.py -k random pytho...",https://github.com/pytorch/pytorch/pull/186249,2f7681b1a92efaf28dfd9d6404fd26b550dedaa62739cc177700b6ff3e79d8cd closes,pr,186261,issue,137009,high,pr.body,"sion coverage for saved-variable detach tracing, no-handler enable_pre_dispatch detach, and pre-dispatch priority over a normal mode. Fixes #137009 Generated by my agent Test Plan: ninja -C build c10 torch_python cp build/lib/libc10.so build/lib/libtorch_python.so build/lib/li...",https://github.com/pytorch/pytorch/pull/186261,12b5251f9d64648da73d27189f160ad9f520992d66a872e26446c0ce7d07ff5f closes,pr,187890,issue,185888,high,pr.body,"d internal-tensor root cause in the generic inplace metadata sync path instead of adding an as_strided_-specific graph-input handler. Fixes #185888 Generated by my agent Benchmark Results: A tiny compile-time micro-benchmark compiled a function that creates a tensor, runs x.ad...",https://github.com/pytorch/pytorch/pull/187890,ce812a50e642f15b5e3d33c6041eef77956ea21c9b358af1475af91ffa841a6a references,pr,187890,pr,186205,medium,pr.body,be specialized; otherwise stale static values can survive and hide the symbolic fake metadata. This is intentionally narrower than stale PR #186205: it does not change the existing graph-input inplace-view graph-break policy. It fixes the reported internal-tensor root cause in...,https://github.com/pytorch/pytorch/pull/187890,c454d8e890641520e8d14dd66b263952dec151be422accf6f47b15f7f293989a closes,pr,186375,issue,186374,high,pr.closingIssuesReferences,pr #186375 declares a closing reference to issue #186374.,https://github.com/pytorch/pytorch/pull/186375,63092ba5224e03b5310d29efd6fec73ebb1b96bef15bdc27ec4ed3f923043a82 supersedes,pr,186375,pr,185946,medium,pr.body,"fabric_access wraps it in an assertion, causing a crash on Orin and other pre-Hopper devices instead of falling back gracefully. Similar to #185946, this PR replaces the assertion with a warning and allows control flow to continue. Fixes #186374 cc @eqy",https://github.com/pytorch/pytorch/pull/186375,79bfce3e6bee3e7d5d7a06d66018b829bf851915f9f5cbbb8b425bfb9592c094 closes,pr,186375,issue,186374,high,pr.body,"ad of falling back gracefully. Similar to #185946, this PR replaces the assertion with a warning and allows control flow to continue. Fixes #186374 cc @eqy",https://github.com/pytorch/pytorch/pull/186375,796bec6cad2f1bb153fda6fec29af36eea9a57ffc90a0c1d9273a7a1086449cb closes,pr,186263,issue,136642,high,pr.body,ners pointed at more aggressive constant propagation and the same lost constant payload affects multiple downstream scalar consumers. Fixes #136642 Generated by my agent Test Plan: python test/export/test_export.py TestExport.test_strict_export_constant_props_small_shape_tenso...,https://github.com/pytorch/pytorch/pull/186263,df004148820c65b6372317c0588212127e29d567c8bcb187fd9ccf5d1650f906 references,pr,186263,issue,136642,medium,pr.comments[2].body,"is is a clean, well-motivated fix. Raising `CONSTANT_NUMEL_LIMIT` from 1 to 8 is a minimal change that directly addresses the root cause of #136642 — shape tensors losing their constant payload and triggering `GuardOnDataDependentSymNode` in strict export. ### Feedback **1. Du...",https://github.com/pytorch/pytorch/pull/186263,c3d0c35884c41d3e7a5ee812a9459390b53f797612d599c4532a8877ef2cae9b references,pr,186263,issue,136642,medium,pr.comments[5].body,- [x] Review the test additions - [x] Post review feedback --- ### Review Summary This is a well-motivated fix addressing the root cause of #136642. The change from limit 1 → 8 is conservative and directly solves shape tensors (built via `torch.as_tensor(...)`) losing their co...,https://github.com/pytorch/pytorch/pull/186263,de1aa0763e5f0fcc46c04acbda5310b394bd3997bc646785ff0c230d07dcf100 references,pr,186263,issue,136642,medium,pr.comments[9].body,tent with fewer breaks folding more into each graph. - **`test_strict_export_constant_props_small_shape_tensor`**: good end-to-end repro of #136642. ### Other - Constant dedup via the `fake_tensor` import is clean — resolves the earlier sync-hazard feedback. - The ~1.87x micro...,https://github.com/pytorch/pytorch/pull/186263,25b66814217e34c6f99c47c4a434e1d0b7126778bd7c627e4590a27cb09da2eb closes,pr,186280,issue,136404,high,pr.body,ecause users still need runtime and compiled-region events in the trace; only compiler Python stack tracing is the pathological part. Fixes #136404 Generated by my agent Test Plan: ninja -C build torch_python timeout 600s python test/dynamo/test_profiler.py DynamoProfilerTests...,https://github.com/pytorch/pytorch/pull/186280,c4a0f643c9bc2d99517eb540081622c69d9d1be695b8a1a175fc19d3334d05e3 closes,pr,186287,issue,135759,high,pr.body,"CIA ops remain handled by the export CIA override path, while non-functional post-autograd entries now fail early with a clear error. Fixes #135759 Generated by my agent Benchmark Results Measured table construction with timeit.repeat(number=20, repeat=5), comparing main again...",https://github.com/pytorch/pytorch/pull/186287,b3259b76af3156fcf40563c0c373d70572c69a9dbac11efab72bfe73bb35671c closes,pr,184294,issue,135492,high,pr.body,r dependencies. This avoids recompilation when a large tensor storage is present but the generated kernel only indexes a small slice. Fixes #135492 Generated by my agent cc @voznesenskym @penguinwu @EikanWang @jgong5 @Guobing-Chen @XiaobingSuper @zhuhaozhe @blzheng @wenzhe-nrv...,https://github.com/pytorch/pytorch/pull/184294,9eb7c0307cf302167fcaf1dd414f54ae72c63bf9969e562e7e59bc85b2187c5a closes,pr,186340,issue,132923,high,pr.body,rst. Keep Python-side tracking as a fallback for environments where the local extension has not yet been rebuilt with the new getter. Fixes #132923 Generated by my agent Test Plan: python test/dynamo/test_recompile_ux.py IsolateRecompilesTests.test_compile_skip_code_function_a...,https://github.com/pytorch/pytorch/pull/186340,24095c9611c9c39903e3f94419be76a96dd9fb04378bbea61f96e721d5eaf7d8 closes,pr,186291,issue,135061,high,pr.body,"t cause is common Distribution validation, so the fix belongs in the base class paths that perform the tensor-mask-to-bool reduction. Fixes #135061 Generated by my agent Test Plan: python test/export/test_export.py -k test_non_strict_export_distribution_validation lintrunner -a",https://github.com/pytorch/pytorch/pull/186291,753262bebb934ca9d58a88abcf61f56918ce339084cf7037db7db3b17ab311a2 references,pr,189070,pr,189068,medium,pr.body,ISTRIBUTED_USE_TORCHCOMMS=1. Stack created with Sapling. Best reviewed with ReviewStack. #189074 #189073 #189072 #189071 -> #189070 #189069 #189068,https://github.com/pytorch/pytorch/pull/189070,655f0e2ac77c3d9d734fff06d731b3299dc22dfaa384d9ee131eb5f7e8edac6b references,pr,189070,pr,189069,medium,pr.body,TORCH_DISTRIBUTED_USE_TORCHCOMMS=1. Stack created with Sapling. Best reviewed with ReviewStack. #189074 #189073 #189072 #189071 -> #189070 #189069 #189068,https://github.com/pytorch/pytorch/pull/189070,34c7ce496ac57b16a04081b978a8782b11118de97df4e907e5d9ac2a63593a79 references,pr,189070,pr,189071,medium,pr.body,a gloo group under TORCH_DISTRIBUTED_USE_TORCHCOMMS=1. Stack created with Sapling. Best reviewed with ReviewStack. #189074 #189073 #189072 #189071 -> #189070 #189069 #189068,https://github.com/pytorch/pytorch/pull/189070,c27a9f8faad28edf11f4e35f5809ad36ae39591113bf7e7c27ca8b8372fdf9be references,pr,189070,pr,189072,medium,pr.body,rrier on a gloo group under TORCH_DISTRIBUTED_USE_TORCHCOMMS=1. Stack created with Sapling. Best reviewed with ReviewStack. #189074 #189073 #189072 #189071 -> #189070 #189069 #189068,https://github.com/pytorch/pytorch/pull/189070,be2255c2c8e4bdc49f2b8fb5851fa4ee87e8a1d3c04dfdcf2f56319a7ee33361 references,pr,189070,pr,189073,medium,pr.body,tored_barrier on a gloo group under TORCH_DISTRIBUTED_USE_TORCHCOMMS=1. Stack created with Sapling. Best reviewed with ReviewStack. #189074 #189073 #189072 #189071 -> #189070 #189069 #189068,https://github.com/pytorch/pytorch/pull/189070,aefe53f14301d42b0a1c500ab9f901f157b4030a7bd61d1065489e0bc4283fa2 references,pr,189070,pr,189074,medium,pr.body,an: monitored_barrier on a gloo group under TORCH_DISTRIBUTED_USE_TORCHCOMMS=1. Stack created with Sapling. Best reviewed with ReviewStack. #189074 #189073 #189072 #189071 -> #189070 #189069 #189068,https://github.com/pytorch/pytorch/pull/189070,61c93fc9dc7f23436a746acb64c2981a76a1bdbcb4393cfffd7e54043adb6e75 references,pr,189073,pr,189068,medium,pr.body,ame and the split completes. Stack created with Sapling. Best reviewed with ReviewStack. #189074 -> #189073 #189072 #189071 #189070 #189069 #189068,https://github.com/pytorch/pytorch/pull/189073,477616eff12c43ac3cc48216478a8c8b36f3292594b4eb5e86d3d043a345564c references,pr,189073,pr,189069,medium,pr.body,e same name and the split completes. Stack created with Sapling. Best reviewed with ReviewStack. #189074 -> #189073 #189072 #189071 #189070 #189069 #189068,https://github.com/pytorch/pytorch/pull/189073,5f252cf518ff3925fad09279689a7eb16cf232d4c35d3f0c6538618183b9eab7 references,pr,189073,pr,189070,medium,pr.body,mpute the same name and the split completes. Stack created with Sapling. Best reviewed with ReviewStack. #189074 -> #189073 #189072 #189071 #189070 #189069 #189068,https://github.com/pytorch/pytorch/pull/189073,797ab0f7456877b67c8a9487dd37d20e7c65e280b6edb8aef98447b010449e1b references,pr,189073,pr,189071,medium,pr.body,ranks compute the same name and the split completes. Stack created with Sapling. Best reviewed with ReviewStack. #189074 -> #189073 #189072 #189071 #189070 #189069 #189068,https://github.com/pytorch/pytorch/pull/189073,2444a7e1c04ad99d09b9543de6d499679fb4a430359f50fdc94dcdc9e4b61cd3 references,pr,189073,pr,189072,medium,pr.body,er: all ranks compute the same name and the split completes. Stack created with Sapling. Best reviewed with ReviewStack. #189074 -> #189073 #189072 #189071 #189070 #189069 #189068,https://github.com/pytorch/pytorch/pull/189073,a7298ba4bbccc53d050089b17351ca4cbe1cb580551ed29b5f500c6804182afe references,pr,189073,pr,189074,medium,pr.body,onnectFullMesh. After: all ranks compute the same name and the split completes. Stack created with Sapling. Best reviewed with ReviewStack. #189074 -> #189073 #189072 #189071 #189070 #189069 #189068,https://github.com/pytorch/pytorch/pull/189073,650a9fd14aa11b5c68a38baf661ebac63105e6aa37b99ea9afd23b2151b6cc14 references,pr,189074,pr,189068,medium,pr.body,meError: already finalized). Stack created with Sapling. Best reviewed with ReviewStack. -> #189074 #189073 #189072 #189071 #189070 #189069 #189068,https://github.com/pytorch/pytorch/pull/189074,1b7a5dd8a3a0e5fd40e347286a7fbd990c8a592c13c34efc06bef91837828d6e references,pr,189074,pr,189069,medium,pr.body,ly RuntimeError: already finalized). Stack created with Sapling. Best reviewed with ReviewStack. -> #189074 #189073 #189072 #189071 #189070 #189069 #189068,https://github.com/pytorch/pytorch/pull/189074,32b8011837dce4f76588a37bdb9e3b87137a79ceb14c81b21d1e712ccc9dee2e references,pr,189074,pr,189070,medium,pr.body,previously RuntimeError: already finalized). Stack created with Sapling. Best reviewed with ReviewStack. -> #189074 #189073 #189072 #189071 #189070 #189069 #189068,https://github.com/pytorch/pytorch/pull/189074,a7264d25d886305acc8375b61933cf0bf406dda708331060ec145c084e954a52 references,pr,189074,pr,189071,medium,pr.body,leanly (previously RuntimeError: already finalized). Stack created with Sapling. Best reviewed with ReviewStack. -> #189074 #189073 #189072 #189071 #189070 #189069 #189068,https://github.com/pytorch/pytorch/pull/189074,6768bb5c9614c8a340101d16affabd79b6744b661a2b5f85eaae1ceef7f71b50 references,pr,189074,pr,189072,medium,pr.body,exits cleanly (previously RuntimeError: already finalized). Stack created with Sapling. Best reviewed with ReviewStack. -> #189074 #189073 #189072 #189071 #189070 #189069 #189068,https://github.com/pytorch/pytorch/pull/189074,1e3ed37bd86b0f5c6de2f0c02681d719c812726ad3ea308b449ed97b8de174c9 references,pr,189074,pr,189073,medium,pr.body,omms now exits cleanly (previously RuntimeError: already finalized). Stack created with Sapling. Best reviewed with ReviewStack. -> #189074 #189073 #189072 #189071 #189070 #189069 #189068,https://github.com/pytorch/pytorch/pull/189074,b81c38cd7d8b0c1af47e5de8e5c862b98443d46a9077c7cf390cd9139cd53712 references,pr,189072,pr,189068,medium,pr.body,it (TP/PP) under TorchComms. Stack created with Sapling. Best reviewed with ReviewStack. #189074 #189073 -> #189072 #189071 #189070 #189069 #189068,https://github.com/pytorch/pytorch/pull/189072,640dc956803d3de99a7d5772153fdbe57357139cde3b6c0c7b2132374c286e45 references,pr,189072,pr,189069,medium,pr.body,atron init (TP/PP) under TorchComms. Stack created with Sapling. Best reviewed with ReviewStack. #189074 #189073 -> #189072 #189071 #189070 #189069 #189068,https://github.com/pytorch/pytorch/pull/189072,936200022afedefd386cd18eb097bd3405a9347dd692daea24e254d429ad0e82 references,pr,189072,pr,189070,medium,pr.body,full megatron init (TP/PP) under TorchComms. Stack created with Sapling. Best reviewed with ReviewStack. #189074 #189073 -> #189072 #189071 #189070 #189069 #189068,https://github.com/pytorch/pytorch/pull/189072,42ae2f2e7579b443e9edb66ff6ccc59a718d6abcf3b65f6fd0187699ab5fa792 references,pr,189072,pr,189071,medium,pr.body,group + full megatron init (TP/PP) under TorchComms. Stack created with Sapling. Best reviewed with ReviewStack. #189074 #189073 -> #189072 #189071 #189070 #189069 #189068,https://github.com/pytorch/pytorch/pull/189072,66285af9693acb297c5e3a2f15e2044659a813628059f922ac4f246a65aa8cfc references,pr,189072,pr,189073,medium,pr.body,k nccl-lazy P2P subgroup + full megatron init (TP/PP) under TorchComms. Stack created with Sapling. Best reviewed with ReviewStack. #189074 #189073 -> #189072 #189071 #189070 #189069 #189068,https://github.com/pytorch/pytorch/pull/189072,ba797088f941657fa43e1f79e69c6d0540132010a31cf1ce4b5cc6e1c645b9c2 references,pr,189072,pr,189074,medium,pr.body,n: 2-rank nccl-lazy P2P subgroup + full megatron init (TP/PP) under TorchComms. Stack created with Sapling. Best reviewed with ReviewStack. #189074 #189073 -> #189072 #189071 #189070 #189069 #189068,https://github.com/pytorch/pytorch/pull/189072,3bdfe5b685fd6d9a40fb079aae8ff6244d53094b711ae46938ed80239a7023e5 references,pr,189071,pr,189068,medium,pr.body,ectly; no split delegation). Stack created with Sapling. Best reviewed with ReviewStack. #189074 #189073 #189072 -> #189071 #189070 #189069 #189068,https://github.com/pytorch/pytorch/pull/189071,a9257c3c3f61d55b18725f33fbb74c3abf8d2fe5e0884e07069cb956c9426f85 references,pr,189071,pr,189069,medium,pr.body,roup directly; no split delegation). Stack created with Sapling. Best reviewed with ReviewStack. #189074 #189073 #189072 -> #189071 #189070 #189069 #189068,https://github.com/pytorch/pytorch/pull/189071,8814c6da7c06b25bf9cb47d786d6f2fa9ed8664f1e32135c359206b5b0bd5bb2 references,pr,189071,pr,189070,medium,pr.body,s-only group directly; no split delegation). Stack created with Sapling. Best reviewed with ReviewStack. #189074 #189073 #189072 -> #189071 #189070 #189069 #189068,https://github.com/pytorch/pytorch/pull/189071,fa681ddd5f9c1c37a31fc6abc0545517c2a8f0eed460226ff06996d1f5092571 references,pr,189071,pr,189072,medium,pr.body,oup builds a members-only group directly; no split delegation). Stack created with Sapling. Best reviewed with ReviewStack. #189074 #189073 #189072 -> #189071 #189070 #189069 #189068,https://github.com/pytorch/pytorch/pull/189071,8df705fe9a9ec6d9c740f18d0b2cdcc658a3fc1c120865d89e16beb299a5b032 references,pr,189071,pr,189073,medium,pr.body,(new_group builds a members-only group directly; no split delegation). Stack created with Sapling. Best reviewed with ReviewStack. #189074 #189073 #189072 -> #189071 #189070 #189069 #189068,https://github.com/pytorch/pytorch/pull/189071,264f4db0d6905ad5c9b6fbfccef6d0fda3e23b47e1615eb8acb85d2941101574 references,pr,189071,pr,189074,medium,pr.body,comms.py (new_group builds a members-only group directly; no split delegation). Stack created with Sapling. Best reviewed with ReviewStack. #189074 #189073 #189072 -> #189071 #189070 #189069 #189068,https://github.com/pytorch/pytorch/pull/189071,e4a09969be6eb7907f8a5339b51d3071a7bb43dbf03c58a645e5772cdd88fcb3 closes,pr,184997,issue,169188,high,pr.body,"nfirmed the root cause: the pre-fix backend returned a wrong scalar result, while loading this PR's backend produced the eager value. Fixes #169188 Generated by my agent Test Plan: PYTHONPATH=$PWD python test/dynamo/test_backends.py TestOptimizationsCPU.test_tvm_sets_scalar_te...",https://github.com/pytorch/pytorch/pull/184997,08cf628f4bbf13be1ffc1ea9b39621a3e9789145c39978755d3daf678c5dc58d review guidance,pr,184997,issue,169188,high,pr.reviews[0].body,"My biggest problem with this PR is that it is untested by your agent: ""Real TVM integration was not run because tvm is not installed in this environment."" Before landing, can you confirm that at least the CI tests this? What Claude suggests here seems to be a good idea, especially since your agen...",https://github.com/pytorch/pytorch/pull/184997,31efd092a510104583fa5021991eb2964f9fd9e9019f13181fa9e82f68af59c0 review guidance,pr,184997,pr,184997,high,pr.reviews[0].body,"My biggest problem with this PR is that it is untested by your agent: ""Real TVM integration was not run because tvm is not installed in this environment."" Before landing, can you confirm that at least the CI tests this? What Claude suggests here seems to be a good idea, especially since your agen...",https://github.com/pytorch/pytorch/pull/184997,f2a758a8cb6c476176bd66b0aa6271eaeefd1698f69e31a86171c2091c814c4a review guidance,pr,184997,issue,169188,high,pr.reviews[1].body,"Someone should run real TVM integration here. If noone does or is able to, noone cares about this change.",https://github.com/pytorch/pytorch/pull/184997,ffa822af7e7aa1926502c253a32ecc1408693472453cc85ec44682df1ef7afdb review guidance,pr,184997,pr,184997,high,pr.reviews[1].body,"Someone should run real TVM integration here. If noone does or is able to, noone cares about this change.",https://github.com/pytorch/pytorch/pull/184997,d0c0c5eb2cf39012c7028518f345b3e0d10dfbcbb6b4eb003467b0fbf7157d6c review guidance,pr,188095,pr,188095,high,pr.reviews[0].body,"Testing (blocker) The new auto-enable logic in flex_attention is untested. test_prescale_qk_default_gating only exercises _apply_kernel_options (which always returns False) plus explicit pinning -- its own comment admits the real opt-in is ""at the flex_attention call site."" Two untested branches:...",https://github.com/pytorch/pytorch/pull/188095,cf8a5daa6fb2a2b0fc57b685f92f2388f832e25bd012bcd80a5b04d2abae656a review guidance,pr,188095,pr,188095,high,pr.reviews[1].body,"presacle_qk should not be set to default, it has numeric implementations",https://github.com/pytorch/pytorch/pull/188095,70395678136866ac957a838472d2fce6cf800d5bcc6eae0d36bf3b79b42d9d42 closes,pr,188546,issue,188541,high,pr.closingIssuesReferences,pr #188546 declares a closing reference to issue #188541.,https://github.com/pytorch/pytorch/pull/188546,e16883d87005ce375a5ba04019e2e6bb8487fb66b84eefa4a05c416c22b5315e closes,pr,188546,issue,188541,high,pr.body,Summary Fixes #188541 Root cause: Triton kernels run with FTZ enabled by default (disable_ftz: False in triton_meta). NVIDIA libdevice does NOT call log1pf as an,https://github.com/pytorch/pytorch/pull/188546,4158770f7f379fdad705a73885a15e4cc963a0fafc64a79b54be127886d40f15