AI Music GeneratorsPublished

SongGen Audio Quality: Clean Codec Limits and Low-End Mud

A practical guide to checking SongGen codec bandwidth, bass buildup, vocal grit, and mixed or dual-track WAV output before cleanup and mastering.

Clean my SongGen render
Producer adjusting faders on analog desk in intimate studio with warm lighting

A SongGen clip can have a convincing lyric, voice, and instrumental idea while the file still feels cramped. The kick and bass may merge into one soft weight. A vocal edge can turn grainy. Cymbals may stop sounding like individual hits. None of those observations proves that the model has one fixed defect, and none is a reason to master harder.

Keep the first output untouched. Listen to the same phrase on headphones, studio monitors, and one ordinary speaker at a steady level. Compare mixed and dual-track outputs when both are available. The goal is to locate an audible problem before you interpret tokens, bandwidth, or a spectrum display.

Checked September 21, 2026. The official paper and repository describe SongGen as a fully open-source, single-stage auto-regressive transformer for text-to-song generation. It accepts lyrics and descriptive text, can use an optional three-second reference voice, and offers Mixed Pro and dual-track modes. Its research system predicts discrete X-Codec audio tokens at a 50 Hz frame rate and decodes audio at 16 kHz. The current model is limited to English songs up to 30 seconds.

The repository’s inference example writes songgen_out.wav at the model sampling rate. It does not document MP3, FLAC, or 24-bit output. That is a much narrower and more useful claim than assuming every interface or fork exports the same formats.

The short answer: inspect the SongGen file before mastering

Start with the musical decision. If the lyric is wrong, the vocal phrasing feels disconnected, or the kick pattern fights the melody, regenerate or revise the input. Cleanup cannot turn an arrangement problem into the intended performance.

If the song works but a repeatable texture distracts you, isolate it. Compare the first verse, the densest section, and the final tail. A broad low-frequency buildup can make the mix feel slow and hide the kick. A narrow midrange resonance can make the voice tiring. Grain around consonants and cymbals can become more obvious after limiting. Treat each as a listening test, not as a model diagnosis.

I start with the unprocessed WAV and lower the monitor level before touching an EQ. At a quieter level, it is easier to tell whether the kick has a clear front edge or only seems impressive because the bass is loud. I then level-match every bypass. The louder version nearly always wins a casual comparison, even when its artifacts are worse.

If you need names for what you hear, use AI music artifacts explained. Label one sound and one timestamp. That is enough to choose the next action.

Diagnose the SongGen render you actually have

SongGen’s architecture explains how the research model works, but it does not publish a universal artifact profile. Diagnose the file in front of you.

  • Low-end mud: Kick and bass occupy the same broad area, so neither has a clean outline. Check the same passage in mono. If the problem disappears when one part stops, it may be arrangement masking rather than codec damage.
  • Sub-rumble: Energy below the useful bass range consumes headroom without adding a stable note. Bypass a high-pass filter at matched loudness before deciding that removal helps.
  • Vocal grit: Sustained vowels or consonants sound granular, especially against sparse accompaniment. Compare Mixed Pro with a dual-track vocal when possible; do not assume token quantization is the sole cause.
  • Midrange congestion: Voice, guitar, keys, and snare seem to occupy one narrow foreground. A dynamic cut may reveal separation, but a broad static scoop can make the entire song distant.
  • Restricted top end: The file can sound closed because the research codec operates at 16 kHz, which sets a much lower theoretical bandwidth than a 44.1 or 48 kHz production file. Upsampling does not recreate information that was never decoded.
  • Edit clicks or ambience jumps: A join between generated sections produces a discontinuity. Fix a local boundary locally; do not process the whole mix for one click.
  • Limiter stress: A louder test exaggerates grain, bass pumping, or vocal edge. Compare it with the untouched file at equal perceived loudness.

A spectrum analyzer can show where energy gathers. It cannot tell you whether the source was the model, the codec, a reference clip, a later edit, or your monitoring room. A visible ridge does not prove that SongGen caused it. If you cannot hear the problem at the same timestamp, leave the file alone.

Source preflight: model, mode, format, and rights

Record the repository revision, checkpoint, generation mode, prompt, lyrics, seed, sampling settings, and reference-voice choice. SongGen supports text and lyric control plus an optional three-second voice reference. That reference must be lawful to use; a technical voice-cloning feature is not permission to imitate a person.

The paper describes two output strategies. Mixed mode predicts the combined vocal-and-accompaniment signal. Mixed Pro adds an auxiliary vocal-token learning objective during training to improve vocal attention, but those extra heads are not used during inference. Dual-track mode generates synchronized vocal and accompaniment token streams using parallel or interleaving patterns. These are architecture details, not promises that one mode will always sound cleaner in your song.

Compare mixed and dual-track outputs when you generated both from the same conditions. In the README example, the dual-track arrays are trimmed to the same length, summed, and written as songgen_out.wav. If your workflow keeps the vocal and accompaniment arrays, archive them before summing. They give you a more precise place to investigate a vocal resonance or accompaniment haze. Do not run a third-party separator and call those files native SongGen tracks.

The paper specifies X-Codec with eight codebooks, 1,024 entries per codebook, a 50 Hz token frame rate, and 16 kHz audio. This confirms a finite codec representation and restricted decoded bandwidth. It does not establish a guaranteed low-end harmonic distortion curve, a fixed defect below 150 Hz, or the cause of every audible blur.

The official repository uses the Apache 2.0 license. That covers the code under its terms; it does not automatically clear lyrics, a reference voice, training-derived concerns, samples, trademarks, or a generated recording for your intended use. Keep a dated copy of the terms you relied on and review every input right before distribution.

SongGen token discretization vs low-end harmonic distortion

SongGen Token Discretization vs Low-End Harmonic Distortion

Listening area What you might hear What the source confirms Smallest useful check
Below 30 Hz Unstable rumble or headroom loss No universal SongGen defect is documented here Bypass a gentle high-pass filter and stop if kick or bass loses weight
30–80 Hz Kick fundamental and bass note blur together X-Codec represents audio as discrete tokens; this does not prove the cause Solo nothing at first; compare stereo, mono, and a small speaker
80–150 Hz Boxy bass tail masks the kick body The official sources do not publish a fixed distortion curve below 150 Hz Try a narrow dynamic reduction only during the buildup
150–500 Hz Warmth becomes cloudy across several instruments Arrangement density, room monitoring, or decoding may contribute Compare mixed output with native dual-track arrays when available
Upper range The render feels closed or grainy The documented X-Codec path operates at 16 kHz Avoid synthetic air boosts until harshness and level are under control

This matrix is a diagnostic guide, not a measured specification for every SongGen file. Token discretization can introduce reconstruction limits, but a bass conflict can also be musical: two sustained parts may simply share the same notes and envelope. A codec explanation is useful only after the listening comparison supports it.

I check the 80–150 Hz area on a modest speaker before reaching for a surgical tool. If the kick becomes easier to follow when the bass note ends, the arrangement is telling me more than the analyzer. If a short, narrow reduction helps only at the crowded moment, I keep it dynamic. If the whole song becomes thinner, I undo it.

Restrained manual cleanup in your DAW

Duplicate the WAV and mark four places: an exposed vocal, the strongest kick-and-bass overlap, the brightest percussion moment, and every edit boundary. Set a consistent monitoring level. Make one reversible move at a time.

For low-end buildup, begin in mono and watch a spectrum analyzer with slow averaging. Add a dynamic EQ band only while the buildup appears. Dynamic means the cut reacts to the problem instead of removing body from every note. Start with a small reduction, then bypass. Use dynamic EQ only when the buildup appears.

For energy below roughly 30 Hz, a gentle high-pass filter can recover headroom, but only if the range is not carrying an intentional bass fundamental. Raise the cutoff slowly. Stop as soon as the kick or bass becomes smaller. A neat-looking spectrum is not the goal.

For vocal or cymbal grit, loop the exact syllable or hit. A narrow dynamic band may soften a resonance. If the whole top end feels hard, follow the harsh-highs guide and avoid a broad treble cut that removes clarity from everything.

For a click, zoom into the edit boundary and use the shortest fade or crossfade that removes it. A crossfade cannot repair a wrong chord, missing consonant, or change in performance. Send those problems back to generation or editing.

Do not upsample a 16 kHz file and call it restored bandwidth. A higher container sample rate can be useful for a DAW session, but it does not recreate absent source detail. Keep the original file, document the conversion, and compare the audible result rather than the new number.

The Sunofix cleanup path for SongGen audio

Manual editing is the better tool for one click, one bass collision, or one vocal resonance. Sunofix is more relevant when an artificial texture moves through the stereo mix and a static correction damages too much music.

I founded Sunofix for this cleanup-before-mastering step. The song idea should already be worth keeping. Upload a lawful WAV or MP3 working copy, choose a conservative pass, and compare the same timestamps at matched loudness. Sunofix aims to reduce metallic texture, harsh highs, and artificial grain while preserving melody, lyrics, arrangement, dynamics, and emotional shape.

Sunofix does not inspect SongGen tokens, prompts, checkpoint state, or reference-voice settings. It does not separate stems, rewrite a bass line, restore information outside the decoded bandwidth, or regenerate a vocal. It prepares a cleaner source for the next decision.

Use level-matched before-and-after playback. If the vocal edge recedes while the kick retains its attack and the chorus keeps its energy, the pass may be useful. If the mix becomes smaller, darker, or less expressive, return to the original and reduce the processing.

Cleanup and mastering are separate stages. Cleanup reduces an unwanted source texture. Mastering sets final tonal balance, dynamics, sequencing, and delivery level. The cleanup-versus-mastering guide explains why aggressive limiting can magnify an existing problem and make diagnosis harder.

Sunofix cannot repair lyrics, melody, arrangement, or performance. It cannot reconstruct samples already lost to clipping, codec reconstruction, or lossy conversion. It cannot create authentic stems from a finished mix, remove watermarks, help evade detection systems, certify provenance, or promise that a distributor will accept a track.

Apache 2.0 is a permissive software license, not a blanket music-clearance certificate. Your lyric, descriptive prompt, three-second reference voice, imported audio, samples, and intended distribution remain separate rights questions. Cleanup does not grant legal clearance or guarantee distributor approval.

The research model’s 30-second, English-language limitation also matters. Extending clips with another tool, stitching several generations, or converting formats creates a new workflow with its own failure points. Preserve each source and document each transition so you can trace a click, tone change, or rights issue back to the correct stage.

SongGen audio release checklist

  1. Preserve the original generated WAV and keep the complete output folder unchanged.
  2. Record the repository revision and checkpoint, prompt, lyrics, seed, generation mode, sampling settings, and reference-voice choice.
  3. Verify the real file rather than assuming MP3, FLAC, 24-bit, 44.1 kHz, or 48 kHz output.
  4. Compare mixed and dual-track outputs under the same conditions when both are available.
  5. Keep native vocal and accompaniment arrays separate before making a summed working copy.
  6. Check rights to the reference voice, lyrics, samples, and every other input.
  7. Listen to an exposed vocal, dense section, bass-heavy moment, and ending at one monitoring level.
  8. Check stereo, mono, headphones, and one ordinary speaker before calling a bass problem a codec artifact.
  9. Name one audible problem and timestamp before adding an EQ, filter, or cleanup pass.
  10. Use dynamic EQ only when the buildup appears and bypass it before adding another processor.
  11. Repair a local click locally with the shortest useful fade or crossfade.
  12. Test a conservative Sunofix pass only when the unwanted texture moves through the full mix.
  13. Use level-matched before-and-after playback and stop if the song loses punch, air, or expression.
  14. Master only after cleanup decisions are stable, keeping source, cleaned file, and master separate.
  15. Review the current license and distribution rules for the exact workflow and release date.

If the distraction becomes quieter while the vocal, rhythm, bass shape, and musical intent remain intact, you have a better source for mastering. If the result is merely louder, smoother-looking, or more processed, return to the untouched WAV. The finish line is not a perfect graph. It is a song that remains itself without the texture that was getting in the way.

FAQ

SongGen Audio Quality: Clean Codec Limits and Low-End Mud FAQ

What causes digital artifacts in SongGen audio?

SongGen predicts discrete audio tokens and decodes them with X-Codec, but that architecture alone does not identify the cause of a sound in your file. Grit, bass blur, or a click can also enter through the prompt, reference voice, generation mode, edit, conversion, or later limiting. Compare the untouched WAV at the same timestamp before choosing a fix.

Does SongGen always create low-end distortion below 150 Hz?

No. The official paper does not define a universal defect below 150 Hz. Use that range as a listening and measurement area only when kick and bass audibly mask each other, the sub range consumes headroom, or mono playback changes the balance.

Can SongGen provide separate vocals and accompaniment?

The official dual-track mode generates vocal and accompaniment arrays in sync, while the README example sums them into songgen_out.wav. Preserve the arrays separately if your lawful local workflow exposes them, and do not describe third-party separation as a native SongGen output.

Can Sunofix clean a SongGen WAV?

Sunofix can process a lawful WAV or MP3 copy when an unwanted artificial texture moves through the full mix. It cannot rewrite lyrics, melody, arrangement, timing, or performance, restore clipped samples, grant rights, or guarantee distributor approval.