Auditius
▸ audio editor in your browser
Cool Edit Pro is back,
in your browser.
A waveform editor and a multitrack studio with loops, buses and automation, 38 effects, noise reduction and optional AI for transcription and voices. All of it runs on your device: your audio is never uploaded.
Auditius is an audio editor and multitrack studio that runs in the browser, in the spirit of Cool Edit Pro, the editor I used in 2002 until it stopped working and that Adobe bought in 2003 to turn into Audition. Plain menus that do what they say, a waveform you edit directly and a multitrack view next to it. It is free and there is no account.
The audio never leaves the computer
That is the rule everything else follows. Opening a file, applying an effect, reducing noise, stretching the tempo and encoding the result all happen on the device. There is no upload step because there is no server to upload to: the site is a set of static files.
The work is saved in the browser as you go. Close the tab and come back the next day, and the files, the cues and the multitrack session are where you left them.
Editing a waveform
The edit view is the classic destructive editor:
- Waveform and spectral display, with selections that can take a single stereo channel, the way Cool Edit did it.
- Cut, copy, paste, mix paste and trim, snapped to zero crossings, with undo as deep as memory allows.
- 38 effects with presets and live preview: EQ, compression, reverb, delay, chorus, flanger, phaser, distortion, pitch and tempo.
- Restoration: noise reduction from a captured profile, click and pop removal, hiss reduction and clip restoration.
- Cues, scrubbing, a pencil to redraw samples, recording from the microphone, and scripts that replay a chain of effects over many files.

The multitrack
Blocks loop by dragging their right edge, fade with handles and crossfade where they overlap. Each block has volume and pan envelopes, each track has a real-time effects rack and automation lanes, and buses with sends let several tracks share one reverb. A built-in loop library of drums, bass and synths follows the session tempo, and a tempo map moves every block that follows it when the tempo changes. The mix goes out as WAV, FLAC, MP3 or Ogg, with the reverb tails allowed to ring out.

AI that runs on your machine
The AI is optional and local too. Every model is downloaded only when you click Download, kept in the browser cache and loaded from there, with any other network request blocked while it runs.
- Transcription with Whisper, with a time for every word: clicking a word plays from it, and the text becomes SRT or VTT subtitles or one cue per sentence.
- Text to speech with Kokoro, Piper and Chatterbox, in Spanish and English, with pause tags in the text.
- Voice cloning with Chatterbox, gated by a consent checkbox. Saved voices stay in that browser, and generated audio is marked as synthetic in the file's metadata.
How it is built
Svelte 5 and TypeScript on the Web Audio API, with no backend at all. The landing pages are rendered at build time, so search engines read plain HTML, and the editor is only fetched when someone opens it.
Heavy work runs off the main thread. Effects, time stretching and the encoders run in a Web Worker with progress and cancel, so the interface never freezes on a long file. MP3 and Ogg go through LAME and Vorbis compiled to WebAssembly; the FLAC encoder and the ZIP writer behind the session files are my own.
Undo stores only what an edit changed. Each step keeps the affected range of samples rather than a copy of the file, under a shared memory budget, which is what lets an hour-long recording be edited in a browser tab.
The models fit the browser's rules. They run on ONNX Runtime, on the GPU through WebGPU when there is one and on WebAssembly otherwise. Even the runtime's own 26 MB binary is downloaded like a model instead of shipping with the site, which keeps every deployed file under the 25 MiB limit of static hosts. Two fixes were needed on the way: Whisper's word timings arrived one token late from the library, and the Spanish voices needed phonemes, which come from a rule-based converter of my own that matches eSpeak NG on 99.9% of words.
It was built with Claude, as a test of fast AI-assisted development that keeps the codebase coherent. The layers can only import downwards and a test enforces it, and the project ships with 365 unit tests and 128 end-to-end tests.
Status
Online since October 2026, in English and Spanish.