How it works

What actually happens to your audio

The short version: upload, preview free, download the master. This page is for when you want to know what the engine does in between, at whatever depth you like.

1

Upload

Audio or video, dragged anywhere. Files are scanned, stored temporarily, and analyzed within seconds.

2

Preview free

The first 30 seconds, fully processed, marked with the FixMark. Iterate as often as you like.

3

Master

The full file, credit-priced by duration, verified by measurement before you get it.

The pipeline, stage by stage

1. Analysis first, always

Before anything touches your audio, the engine measures it: loudness, noise floor, spectral balance, stereo behavior, and whether it is speech, music, or a mix. Every later decision follows from this analysis rather than from fixed settings, which is why a whispered voice memo and a hot club mix get different treatment from the same click.

For the curious: how content detection works

Speech detection is an ensemble of five signals: energy in the voice band, spectral contrast, zero-crossing patterns, syllable-rate onsets, and harmonic structure via pitch tracking. A neural voice activity detector cross-checks the result; when the two disagree on noisy material, the VAD wins, because heavy noise fools spectral heuristics before it fools a trained model.

2. Noise reduction

A deep-learning model separates voice from broadband noise (hiss, fans, street rumble), with a spectral fallback for material where the AI model is the wrong tool. Strength is chosen from your measured noise floor and content type: speech gets treated harder than music, because on music the “noise” is often air and texture you want to keep.

For the curious: which model, and its limits

The AI path is DeepFilterNet, a real-time speech-enhancement network. It is excellent down to roughly the point where noise approaches the level of quiet speech itself; below that, reconstruction starts costing syllables, which is a physics problem more than a model problem. Our extreme rescue demo deliberately sits at that edge so you can hear where it is.

3. Corrective EQ and dynamics

Rumble gets high-passed, muddiness or harshness flagged by the analysis gets shelved down, thin recordings get warmth back. A gentle compressor evens out level differences without flattening the life out of transients.

4. Loudness, verified

The master lands at streaming loudness (-14 LUFS by default, adjustable). We do not trust the math; the engine measures its own output and corrects until it is right, then checks true peak with 4x oversampling and keeps it at or under -1 dBTP so your file survives MP3 and AAC encoding without clipping.

For the curious: why verification beats trusting a limiter

Compression and limiting raise perceived loudness beyond what any gain calculation predicts, and standard limiters only see sample peaks, not the higher inter-sample peaks that appear after lossy encoding. We wrote up both failure modes with numbers in why audio clips after mastering.

5. Optional: dereverberation (beta)

Room echo reduction using weighted prediction error filtering. Honest status: it helps on light rooms, costs real processing time, and strong reverb remains beyond it. It is off by default and clearly labeled beta in the tuning panel.

6. The FixMark, on previews only

Free previews carry the FixMark, a short spoken tag placed a few times across the 30 seconds. It never touches your original upload and never appears in paid masters. It exists so the free preview cannot silently replace the product; that is the entire story.

The preview is the honest demo: your file, the full pipeline, free, about half a minute.

Try it on your audio