Phartakos MLX

a fart machine for the modern age

Initial release: 0.1.0.

An inference-only, unconditional waveform fart generator for Apple Silicon. It emits 32 kHz mono audio with a frozen multiscale waveform U-Net and the exact 200-step ancestral DDPM sampler used in evaluation. This is a small research model, not a claim of realism or parity with recorded audio.

Run locally

Requires Apple Silicon, macOS 14 or newer, Python 3.12 and uv.

uv sync
uv run python app.py

Or use the CLI (one second is the default):

uv run phartakos-mlx --seconds 1 --seed 0 --output fart.wav

Use --random-seed to choose a new unsigned 32-bit seed. The CLI prints the selected seed, and the Gradio app writes it back into the Seed field, so generated audio remains exactly replayable by disabling random mode and reusing that value. Gradio downloads use phartakos_mlx_010_<seed>_<length>s.wav filenames.

The CLI and Gradio app call the same generate() function. Choices are exactly:

  • 1 second: default;
  • 2 seconds: experimental creative option;
  • 3 seconds: experimental and pending validation.

These are output waveform lengths, not guaranteed single-event durations.

Duration evidence

In the closed, fresh-seed 1s-versus-2s zero-shot study, two seconds received 8 exclusive preferences versus 5 for one second, with 3 ties. Mean naturalness was 3.875 versus 3.625. It nevertheless missed the preregistered adoption gate: two-second continuity was slightly lower (4.1875 versus 4.3125), and promotion required at least 9 exclusive wins. One second therefore remains the default. Three-second generation is architecture-compatible but has not had a comparable listening validation and may use materially more time and memory.

Exact inference contract

  • Model: unconditional multiscale waveform U-Net, epsilon prediction.
  • Audio: mono float32, 32,000 Hz; trained on one-second canvases.
  • Process: discrete variance-preserving diffusion, 200 linearly spaced betas from 0.0001 to 0.02.
  • Sampler: full 200-step ancestral DDPM; no shortened or deterministic schedule.
  • Output: per-sample attenuation only, capped at -1 dBFS; it never amplifies or clips the raw result.
  • Seed: unsigned 32-bit integer and deterministic under the pinned runtime on the same supported platform.

config.json is the machine-readable contract. MANIFEST.json records hashes for every exported file and binds this bundle to the source checkpoint, run identity and recipe. model.safetensors contains only model.* tensors renamed to inference keys: optimizer state, training metrics, RNG continuation state, datasets and samples are excluded.

Provenance and limitations

Thank you to Alec Ledoux for creating and sharing the Fart Recordings Dataset on Kaggle, which provided the source recordings used for this model's training lineage.

The selected terminal checkpoint SHA-256 is 052764c7eeea3c6781cbab6e83885ef2d40e29e601925c9033296178227fdc6d. Its run identity SHA-256 is 25e7c585e4abf974683a65b6597fd74956caf2ca7318ebe4cd496d19892c4391. It was trained for 25,000 updates from the RMS-centroid, 2,048-record lineage, ending with frequency-weighted and multiresolution spectral reconstructed-x0 auxiliary losses. Those losses affect training; inference remains epsilon prediction with ancestral DDPM.

Generated audio can be implausible, repetitive, fragmented or resemble training material. The underlying recordings are not distributed here. This package is inference-only and has no training or upload path. The MLX backend is validated for local Apple-Silicon use. A Linux-hosted Hugging Face Space would require a separate backend port, numerical parity checks and listening validation.

License

MIT for code and weights. See LICENSE and NOTICE for implementation ancestry.

Downloads last month
-
Safetensors
Model size
24M params
Tensor type
F32
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support