MusicGen Audio Quality: Fix 32kHz Roll-Off and Narrow Stereo
A practical guide to diagnosing Meta MusicGen EnCodec roll-off, narrow stereo, harmonic grit, and low-end instability before mastering.
Clean my MusicGen render
A MusicGen render can have a strong musical idea and still feel boxed in when you prepare it for release. Cymbals may stop opening up, bright synths can turn sandy, the center may feel crowded, and added loudness can make the texture more tiring rather than more detailed. Those are useful observations. Hearing one does not prove that MusicGen caused it.
Start with the exact render you generated. Keep the original, note the checkpoint and interface, and mark the second where the sound bothers you. Then compare any change at the same perceived loudness. A louder version often appears wider and brighter even when it is only louder.
Checked September 21, 2026. Meta documents MusicGen as a single-stage autoregressive Transformer trained over a 32 kHz EnCodec tokenizer. The published family includes 300M, 1.5B, and 3.3B parameter sizes, text-to-music and melody-conditioned variants, mono base checkpoints, and later stereophonic checkpoints. Those facts explain what to inspect. They do not replace listening to your file.
The short answer: clean the MusicGen render before mastering
A 32 kHz source has a 16 kHz Nyquist limit: it cannot represent normal sampled frequencies above 16 kHz. That ceiling is not a defect you can reverse with an upsampler. If the mix also carries metallic grain, narrow resonances, unstable low end, or edit clicks, clean those audible problems before mastering. A final limiter can pull low-level grit forward and make a narrow mix feel even more congested.
Use a restrained sequence. Confirm the file and checkpoint, listen in stereo and mono, inspect a spectrogram, repair isolated clicks locally, control only resonances that actually become harsh, and leave headroom. If you add stereo width or harmonics, treat them as new production decisions rather than recovered information.
I usually begin with the busiest ten seconds and one exposed tail. The busy section reveals whether grit builds with the arrangement; the tail shows whether the problem is a musical decay, a moving codec-like texture, or steady background noise. That distinction keeps me from processing the whole song for one short event.
For a broader explanation of shimmer, warble, smeared tails, and other symptoms, use AI music artifacts explained. Name what you hear before choosing a processor.
Diagnose the MusicGen render you actually have
Loop two or three seconds around each problem and describe it in plain language. Do not start with a preset called “AI repair.” Start with a symptom and a stop condition.
- High-frequency grit or sandy highs: Cymbals, noise layers, or bright synths break into a fine spray instead of holding a stable tone. Check whether the texture moves with the instrument or remains constant during gaps.
- Midrange resonance: A note, vocal-like lead, or snare rings harder than nearby material. Lower your monitor level; a real resonance often remains annoying when the excitement of loud playback falls away.
- Narrow or crowded center: Several parts sit on top of each other and the sides contribute little space. Compare stereo and mono. A centered arrangement is not automatically broken, and aggressive widening can create a hollow mono fold-down.
- Low-end instability: The sub changes weight from hit to hit or triggers a limiter without sounding stronger. Check a small speaker as well as headphones. Inaudible sub energy can consume headroom.
- Edit discontinuity: A continuation or stitched passage clicks at a join. Zoom into the waveform. A short crossfade can repair a real discontinuity, but it cannot correct a wrong note or an awkward rhythmic transition.
- Flattened contrast: Attacks and sustained sections feel equally dense. Measure peaks and short-term loudness, then decide whether the issue came from the render, later normalization, or your own processing.
A spectrum analyzer shows where energy exists. A spectrogram shows how it changes over time. Neither tool can decide whether the sound is musical, and neither proves an internal cause. A clean-looking graph can accompany a dull mix; a dense graph can belong to an intentional distorted arrangement.
The harsh-highs guide gives a safe dynamic-EQ starting point. Use it only when a specific band becomes aggressive. If the cymbals lose shape or the mix becomes darker everywhere, back off.
Source preflight: checkpoint, channels, format, and rights
Record the exact source before processing: repository or interface, checkpoint name, prompt, seed when available, duration, channel count, and every conversion after generation. MusicGen can run through AudioCraft, a local Gradio demo, a notebook, Hugging Face tooling, or a third-party wrapper. Those paths can normalize and save audio differently.
Meta’s AudioCraft documentation lists small, medium, melody, large, melody-large, and stereo variants. The original family used mono checkpoints; later stereophonic checkpoints interleave two EnCodec token streams. A two-channel file is not proof that you used a stereo checkpoint. It may contain the same mono signal in both channels, so check a correlation meter and listen to the side signal.
The official AudioCraft example writes WAV at the model sample rate. Its demo also returns temporary WAV files. The MusicGen documentation does not document built-in MP3, FLAC, or separated-stem output as a stable MusicGen generation contract. A wrapper may add those exports, but then its settings and terms become part of your source record. Converting a WAV to MP3 and back to WAV does not improve it, and an internal conditioning stem operation is not the same as downloadable production stems.
Licensing needs equal care. The AudioCraft code is MIT-licensed, while the published model weights are CC BY-NC 4.0. That non-commercial weights license matters when you plan a release or client use. A third-party service may offer different terms for its own workflow, but you must verify them directly. Cleanup changes sound; it does not change the license that applied to the model, interface, inputs, or output.
MusicGen EnCodec spectral cutoff at 16kHz
MusicGen EnCodec Spectral Cutoff at 16kHz
Because the documented tokenizer runs at 32 kHz, its Nyquist limit is 16 kHz. In a spectrogram, a first-generation MusicGen file may show energy ending close to that ceiling. The edge can look steep, but measure your own file: a wrapper may resample, encode, filter, or process it after generation.
| Spectral check | What you may observe | Useful next action | What it cannot prove |
|---|---|---|---|
| 0-12 kHz | Most musical fundamentals, harmonics, percussion, and brightness | Diagnose resonances and moving grain here before chasing “air” | That all energy is clean or natural |
| 12-16 kHz | Upper harmonics and texture approaching the source ceiling | Compare the busiest section with an exposed tail | That a fixed artifact belongs to every MusicGen checkpoint |
| At about 16 kHz | A steep ceiling in a first-generation 32 kHz source | Preserve the original and document later resampling | That upsampling can recover content above it |
| Above 16 kHz after upsampling | Usually no recovered original detail; processing may add new energy | Treat new harmonics as synthesis and level-match the result | That the added energy was present in the source |
Upsampling a 32 kHz render to 44.1 or 48 kHz creates a file with a higher sample rate, but it cannot restore source information above the 16 kHz Nyquist limit. A harmonic exciter can generate related upper harmonics. That may make a synth or cymbal feel brighter, yet it is an aesthetic addition, not restoration.
Use a harmonic exciter only as a creative choice. Start with a filtered parallel signal, keep the amount low, and bypass it at matched loudness. Stop if consonant-like attacks become brittle, cymbals turn to spray, or the mix seems brighter only at high playback level. Listeners do not need a graph filled to 20 kHz; they need a track that stays comfortable and clear.
Restrained manual cleanup in your DAW
Preserve the original generated WAV before changing the sample rate, channels, gain, or metadata. Work on a copy and make one change at a time.
First lower clip gain if the file reaches the ceiling. This creates working headroom without pretending to repair clipped samples. High-pass only sub-rumble that steals headroom, starting below the musical bass and moving slowly. Stop when the kick or bass loses weight.
For a narrow harsh band, use dynamic EQ so the reduction appears only when the band becomes aggressive. Sweep briefly to find the area, return the gain to zero, and then reduce the smallest useful amount in context. Do not leave a narrow boost running while you judge the song; it makes almost any frequency sound guilty.
For a genuinely mono source, decide whether width is necessary. Short, frequency-limited ambience or a subtle decorrelated layer can add space, but keep bass and important transients stable in the center. Check the side signal and fold the result to mono. If the lead becomes hollow or the snare loses impact, the widening is too strong.
Repair a single click with a short crossfade around the waveform break. If the texture, chord, or rhythm changes badly across the join, regenerate or edit a wider musical region. Broadband cleanup is the wrong tool for a structural transition.
Use level-matched before-and-after playback after every step. The cleanup-versus-mastering guide explains the handoff: cleanup reduces distracting texture; mastering handles final tonal balance, loudness, sequencing, and delivery. Do not ask a limiter to hide grit. It usually makes the grit louder.
The Sunofix cleanup path for MusicGen audio
I founded Sunofix for the point where the song already works but the exported mix still carries a synthetic edge. A local click or one ringing note may be faster to fix manually. Sunofix is more relevant when haze, metallic movement, or rough upper texture travels through the full mix and a broad EQ cut would remove useful music with it.
Upload a lawful WAV or MP3, choose a restrained cleanup, and compare the result with the original at matched loudness. Listen to the busiest passage, one exposed tail, and the center of the mix. Keep the version that reduces distraction while preserving melody, timing, energy, and the character of the arrangement.
Frequency diagnostics and before-and-after views are supporting evidence. They can show that energy changed, but your listening decision remains primary. If the cleaned file looks smoother and sounds duller, keep the original or use a lighter pass.
Sunofix works on the mixed audio you provide. It does not turn dual mono into independently generated stereo, separate native stems, rewrite the prompt, replace a wrong note, or recreate missing top-end information. It cannot restore source information above the 16 kHz Nyquist limit. Regenerate when the musical material is wrong; clean when the musical material works and an unwanted texture sits on top of it.
Technical and legal boundaries
MusicGen is documented as an EnCodec-tokenized autoregressive model, not a diffusion model. Avoid blaming “diffusion steps” for a MusicGen symptom. Neural tokenization and decoding set real technical boundaries, but one rough cymbal does not prove which stage caused it. The audible result may also reflect prompt choice, seed, checkpoint, duration extension, normalization, resampling, lossy export, or later processing.
The model card describes important limitations: realistic vocals are not a primary strength, English descriptions fit the training setup better than other languages, and performance varies across musical styles and cultures. Those published limits are more useful than invented frequency profiles. They help you decide whether another prompt, checkpoint, or render is the cleaner solution.
I do not try to make the spectrum look like a 48 kHz recording when the source is 32 kHz. I listen for whether the song feels closed, harsh, or tiring, then solve that audible problem with the smallest intervention. Often the honest result still shows the original ceiling and sounds better because nothing unnecessary was added.
Audio cleanup does not grant legal clearance or guarantee distributor approval. Review the current model-weights license, the terms of the interface you used, and the rights attached to prompts, melodies, samples, and intended use. Keep your original files and source notes. Mastering, metadata, artwork, copyright analysis, and distributor review remain separate tasks.
MusicGen audio release checklist
- Confirm the exact MusicGen checkpoint, interface, prompt, seed, duration, and channel count.
- Preserve the original generated WAV or the best first-generation file available.
- Verify whether the source is mono, dual mono, or generated by a stereophonic checkpoint.
- Mark each audible problem by timestamp before opening processors.
- Compare stereo and mono on headphones, small speakers, and a quiet listening level.
- Inspect the spectrum and spectrogram without treating the graph as a quality score.
- Repair isolated clicks locally and regenerate structural musical problems.
- Filter sub-rumble only when it steals headroom or triggers processing.
- Use dynamic EQ only when a measured resonance becomes audibly harsh.
- Treat widening and harmonic excitation as creative additions, not recovered source detail.
- Run restrained cleanup before mastering when unwanted texture moves through the full mix.
- Compare every version at matched loudness and stop when the song starts losing energy or identity.
- Leave headroom, audition a delivery codec, and check peaks after final mastering.
- Recheck the model-weights license, interface terms, input rights, metadata, and destination rules.
A 16 kHz ceiling does not automatically make a track unusable, and a filled-in spectrum does not automatically make it better. Preserve the source, solve audible problems, and keep production choices honest. The goal is not to disguise where the music came from. It is to prepare the strongest lawful version of a song that is already worth keeping.
Continue listening
Related reading
FAQ
MusicGen Audio Quality: Fix 32kHz Roll-Off and Narrow Stereo FAQ
Why does MusicGen audio stop near 16 kHz?
Meta documents MusicGen as using a 32 kHz EnCodec tokenizer. A 32 kHz sample rate has a 16 kHz Nyquist limit, so the generated source cannot contain ordinary sampled audio above that ceiling. Upsampling changes the container rate but does not restore the missing source information.
Can an exciter restore the missing MusicGen high end?
No. An exciter can synthesize new harmonics below or above the original ceiling, which may add a sense of brightness, but it cannot recover detail that MusicGen never generated. Treat excitation as a creative choice and compare it quietly at matched loudness.
Is every MusicGen output mono?
No. Meta released mono base checkpoints and later stereophonic checkpoints. Check the exact model and the file channels instead of judging from the MusicGen name alone. Duplicating a mono file into two channels creates dual mono, not a real stereo image.
Can Sunofix clean Meta MusicGen audio?
Sunofix can process a lawful local WAV or MP3 and may reduce full-mix haze, harsh resonances, or moving digital grain. It cannot recreate missing high-frequency source information, invent true stereo separation, repair the composition, or grant commercial rights.
