Skip to content

test(distributed): compare DTensor basic ops against eager, not torch.compile [14/n] - #2851

Open
bhimrazy wants to merge 1 commit into
bhimrazy/cross-entropy-nvfuser-timeoutfrom
bhimrazy/dtensor-eager-reference
Open

bhimrazy wants to merge 1 commit into
bhimrazy/cross-entropy-nvfuser-timeoutfrom
bhimrazy/dtensor-eager-reference

Conversation

@bhimrazy

@bhimrazy bhimrazy commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

What does this PR do?

Compares test_dtensor_basic_op against eager instead of torch.compile, as the opinfo tests in the same file already do.

Problem: the test built its reference with torch.compile, so an inductor Triton launch ran before thunder was reached. On torch nightly that launch fails with CUDA driver error: invalid argument, and the other rank then waits for the 300s process timeout. So one cause shows up as a crash in one variant and as timeouts in two others. The test is about thunder's lowering of the op over a DTensor, not inductor's.

Checked on 2× L4 with the pt_main-dev image: 3 of the 6 variants fail before (11 min, with the timeouts), and all 6 pass after (55s). Binding each rank to its device first doesn't change the result.

Related CI:

  • all-tests.yaml distributed on pt_main-dev: the three test_dtensor_basic_op_* failures.
CI log (distributed, pt_main-dev, on #2847)
  File ".../torch/_inductor/runtime/static_triton_launcher.py", line 291, in run
    self.C_impl._launch_kernel(
RuntimeError: CUDA driver error: invalid argument

FAILED thunder/tests/distributed/test_dtensor.py::DTensorTest::test_dtensor_basic_op_executor_torch_fn_key_x * w - RuntimeError: Process 1 exited with error code 10 and exception:
FAILED thunder/tests/distributed/test_dtensor.py::DTensorTest::test_dtensor_basic_op_executor_nvfuser_fn_key_x_mul(w) - RuntimeError: Process 0 terminated or timed out after 300.01899886131287 seconds
FAILED thunder/tests/distributed/test_dtensor.py::DTensorTest::test_dtensor_basic_op_executor_torch_fn_key_torch_mul - RuntimeError: Process 0 terminated or timed out after 300.0474462509155 seconds

Part 14 of the breakdown of #2832.

@bhimrazy
bhimrazy added this pull request to stack #2849 September 24, 2026 13:31
test_dtensor_basic_op built its reference with torch.compile, which put an
inductor Triton launch in front of a test that is about thunder's lowering of
the op over a DTensor. On the torch nightly image that launch fails:

  torch/_inductor/runtime/static_triton_launcher.py:291
  RuntimeError: CUDA driver error: invalid argument

and the other rank then waits until the 300s process timeout, so the same
cause shows up as a crash in one variant and as timeouts in two others.
Binding each rank to its device first does not change this; an eager
reference fixes all three, as the opinfo tests already use.
@bhimrazy
bhimrazy force-pushed the bhimrazy/dtensor-eager-reference branch from 4d8ab36 to 741efbf Compare September 24, 2026 13:44
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants