AI Music GeneratorsPublished

YuE AI Music Audio Quality: Clean xCodec Haze and Check the 12 kHz Edge

A practical guide to diagnosing original YuE vocal hiss, xCodec haze, dual-track balance, and apparent high-frequency roll-off before mastering.

Clean my YuE render
Music producer in studio headphones critically comparing YuE vocal, instrumental, and full-mix tracks beside a spectrogram

An original YuE song can be musically convincing and still arrive with a soft veil around the vocal, a sandy tail on held syllables, or a top end that seems to stop opening beyond a certain point. Clean the symptom you can hear. Do not boost everything above 12 kHz just because a graph looks empty, and do not assume every rough edge came from xCodec.

Keep the untouched render, inspect vocal and accompaniment separately when you actually have them, and compare every change at the same loudness. A short local repair may be enough for one hiss or edit. Broader cleanup makes more sense only when the unwanted texture follows several elements through the song.

Checked September 21, 2026. The original YuE release is now preserved on the official YuE-v1 branch. Its paper describes songs up to five minutes, track-decoupled vocal and accompaniment token prediction, a two-stage language-model pipeline, X-Codec operating on 16 kHz material, and a lightweight vocoder that upsamples 16 kHz audio to 44.1 kHz output. The official sources do not describe a universal 12 kHz brickwall in every render.

The short answer: clean the render you have, not a rumored cutoff

Begin with three loops: an exposed vocal phrase, the busiest chorus, and a final reverb decay. Listen before opening a spectrum. If the voice carries a steady hiss between syllables, note whether it also appears in the instrumental. If cymbals turn into a soft spray only in the chorus, check whether the problem is tonal, dynamic, or part of the arrangement.

The original YuE system models vocals and accompaniment as paired token streams, then reconstructs audio through X-Codec and an upsampling stage. That architecture matters because it gives you useful files and checkpoints to record. It does not turn a visual boundary at 12 kHz into a diagnosis. The paper reports fair reconstruction quality for X-Codec and says semantic clustering can compromise acoustic dynamics, but it does not publish “xCodec blur” or “vocal hiss” as fixed defects with a guaranteed frequency.

I usually hide the analyzer for the first pass. If I cannot point to the distracting sound in a blind loop, I do not want a bright display talking me into a repair the song never needed.

Use AI music artifacts explained if you need help separating hiss, shimmer, clipping, resonance, and edit clicks. Each one calls for a different next move.

Diagnose the original YuE render you actually have

Name a timestamp and a symptom before reaching for a plugin.

  • Vocal hiss: Pause on the space after a sustained word. A constant broadband bed may respond to a conservative denoise pass. A texture that moves with pitch or reverb is less likely to behave like ordinary room noise.
  • Codec haze: Compare consonants, cymbals, and reverb tails. If their fine detail collapses into the same soft grain, mark it as a repeatable texture rather than assuming its internal cause.
  • Midrange spikes: Lower the monitor level. A vowel, snare, or synth that still jumps forward may need a narrow dynamic correction. If the whole vocal disappears only when the arrangement gets dense, masking may be the real problem.
  • Low-end instability: Switch the mix to mono. A bass note that becomes smaller or moves sideways suggests phase or width trouble. A spectrum alone will not tell you whether the change is musically harmful.
  • Hard edit or segment join: Zoom into the waveform. One click at one boundary needs a small fade or source edit, not processing across a five-minute song.
  • Missing air: Compare the actual high-frequency energy with the audible result. A darker file is not automatically damaged, and a bright shelf cannot recreate detail that is absent from the source.

Check the original file and any later conversion at the same timestamp. If the hiss appears only after MP3 encoding, saturation, or limiting, fix that stage first. If the vocal is rough in one section but clean in another, regeneration or a local edit may protect more of the performance than full-mix processing.

Source preflight: YuE v1, dual tracks, format, and rights

Record the exact branch, checkpoint, and settings before cleanup. Note the stage-one and stage-two models, seed, genre tags, lyric structure, number of generated segments, repetition penalty, reference-audio mode, and any third-party interface. The official default uses two roughly 30-second sessions to reduce memory pressure; longer songs require more sessions and substantially more compute.

The phrase “dual-track” needs care. YuE’s core Dual-NTP design predicts vocal and accompaniment tokens separately. Separately, dual-track ICL uses separate vocal and instrumental reference tracks to guide style. That is not the same as promising that every interface gives you pristine final stems. Work with the files your run actually produced, label references and generated outputs correctly, and do not call a post-separated mix an original stem.

The preserved inference code reconstructs vocal and instrumental paths, mixes them, and writes a 44.1 kHz result after upsampling. It also contains intermediate 16 kHz reconstruction paths. The official documentation does not document WAV or FLAC as the default final export, so inspect the container, codec, sample rate, channels, and bit depth instead of copying a specification from a prompt or third-party UI. Converting a lossy file to WAV prevents another lossy save during editing; it does not restore discarded information.

Original YuE v1 is preserved under Apache 2.0, and its README encourages commercial use of outputs with credit to “YuE by HKUST/M-A-P.” Keep a license snapshot, but treat it as only one part of clearance. You still need rights to lyrics, reference audio, voices, samples, and any material the output may resemble.

YuE original 12 kHz edge versus reconstructed air

YuE Original 12 kHz Edge vs Reconstructed Air

What you observe What it may mean Smallest useful action Stop when
Energy falls sharply near 12 kHz A feature of this render, decoder, conversion, or analyzer setup Confirm the original file and compare another render from the same workflow The boundary is verified, even if no processing is needed
Top end is dark but comfortable A tonal choice rather than an artifact Leave it alone; compare at matched loudness Brightening makes cymbals or consonants brittle
Vocal hiss moves with pitch Generated or reconstructed texture, not steady background noise Try narrow dynamic cleanup or a conservative full-mix test Words remain natural and breath does not vanish
Cymbals become a broad spray Harshness, limiting, encoding, or reconstruction texture Use a small dynamic cut only when the spray appears Attack stays clear without a papery top end
Added “air” sounds detached Harmonic enhancement is inventing brightness rather than recovering detail Reduce or bypass the enhancer The vocal and ambience sound connected again

A frequency plot can show where energy changes. It cannot show that the missing band once contained useful musical information. The official YuE report explains that X-Codec was trained on 16 kHz audio, uses a 50 Hz frame rate and multiple residual vector-quantizer layers, and feeds a later upsampling vocoder. It does not document a universal hard cutoff at 12 kHz.

Treat “reconstructed air” as a listening goal, not a promise of restoration. A gentle exciter may create harmonics that feel more open, but those harmonics are new content. Keep the untouched source and decide by level-matched listening, not by whether the spectrum looks fuller.

Restrained manual cleanup in your DAW

Duplicate the working file and disable any final limiter. Put markers at the exposed vocal, loudest chorus, lowest bass note, longest tail, and each segment boundary. Make one reversible change at a time.

For steady hiss, capture a short noise-only moment only if one exists. Lower the reduction until the vocal stops sounding watery or gated. If the texture follows the voice, try a dynamic EQ that moves only when the rough band becomes obvious. The harsh-highs guide explains why a moving problem usually responds better to a moving correction than to a permanent shelf.

For one resonant vowel or instrument note, sweep at low gain, find the smallest repeatable area, then automate a narrow cut around the event. Bypass often. If several cuts make the singer sound distant, return to the original and solve only the worst moment.

High-pass sub-rumble only when it consumes headroom or triggers compression without adding musical weight. Raise the filter slowly while listening in stereo and mono, then back down as soon as kick impact, bass sustain, or warmth changes. Do not use “below 30 Hz” as an automatic rule.

For a click at a segment join, use the shortest fade that hides the discontinuity. If ambience, tempo, pitch, or performance changes across the boundary, a long crossfade may simply smear the mismatch. Edit or regenerate the section instead.

I finish this pass by turning every processor off, then enabling each one alone. That simple check catches the point where three tasteful moves combine into a smaller, duller song.

The Sunofix cleanup path for original YuE audio

Manual repair is the right tool for one hiss, one click, one resonance, or one weak edit. Broader cleanup becomes useful when the same artificial grain travels through the vocal, cymbals, synths, and ambience across several sections. Repeating static EQ cuts in that situation can remove more music than artifact.

I built Sunofix for this stage: the lyrics, melody, performance, and arrangement already work, but the exported mix still has a synthetic edge before mastering. Upload a lawful WAV or MP3, keep the pass conservative, and compare the same vocal phrase, chorus, and tail with the original.

Sunofix aims to reduce audible artifact texture without rewriting the song. It does not know your YuE seed or token stream, create true stems, change lyrics, repair timing, or recover missing samples. It cannot restore information that is absent from the source.

Level-match every before-and-after comparison. Check headphones, an ordinary speaker, and mono. If the hiss recedes while breath, consonants, cymbal attack, bass weight, and chorus lift remain intact, you have a cleaner source for mastering. If the processed version wins only because it is louder, or it sounds smoother but smaller, go back and use a narrower move.

Cleanup and mastering solve different problems. Cleanup should come before final mastering, because a limiter can make hiss, codec grain, and harsh resonance more obvious while reducing your room to repair them.

The official YuE paper supports specific claims: generation up to five minutes, track-decoupled vocal and accompaniment prediction, X-Codec tokenization, 16 kHz reconstruction, and learned upsampling to 44.1 kHz. It does not prove that every hiss, blur, resonance, phase problem, or high-frequency boundary in your file came from the model.

Do not market harmonic enhancement as recovered source detail. Do not call separated audio “native stems” unless your exact workflow produced and preserved those tracks. Do not claim that cleanup changes the model, removes provenance, bypasses detection, or makes an unauthorized reference lawful.

Sunofix cannot correct lyrics, melody, composition, arrangement, pronunciation, or a poor generation choice. It does not grant legal clearance or guarantee distributor approval. Software and model terms do not replace permission for lyrics, voices, recordings, samples, or other protected material.

Preserve the original output, prompts, lyrics, references, settings, branch and checkpoint revisions, license snapshot, working copy, cleaned source, and final master as separate files. That record makes listening comparisons repeatable and keeps technical and rights decisions from getting mixed together.

Original YuE audio release checklist

  1. Preserve the untouched output before conversion, cleanup, editing, or mastering.
  2. Record the exact branch, checkpoint, and settings, including both model stages, seed, segments, tags, lyrics, and reference mode.
  3. Inspect the actual file for container, codec, sample rate, channels, clipping, and previous lossy encoding.
  4. Inspect vocal and accompaniment separately when available, but label generated tracks, references, and post-separated files accurately.
  5. Mark an exposed vocal, busiest chorus, lowest bass note, longest tail, and every segment join.
  6. Name one audible symptom before opening an EQ or denoiser.
  7. Verify any apparent 12 kHz edge against the untouched source and another controlled render.
  8. Repair isolated clicks, hiss, or resonances locally before processing the whole mix.
  9. Check low end in stereo and mono before filtering or narrowing it.
  10. Test broader cleanup only when unwanted texture moves through several sources.
  11. Level-match every before-and-after comparison on headphones and an ordinary speaker.
  12. Stop when breath, attack, air, bass weight, or emotion begins to shrink.
  13. Master only after the cleanup decision is stable, then archive each stage separately.

If the unwanted layer becomes less distracting while the song keeps its voice, punch, space, and emotional shape, the source is ready for the next stage. If the graph looks cleaner but the music feels smaller, the graph has won the wrong argument.

FAQ

YuE AI Music Audio Quality: Clean xCodec Haze and Check the 12 kHz Edge FAQ

What causes digital artifacts in original YuE songs?

Original YuE uses a semantic-acoustic X-Codec tokenizer, two language-model stages, and a lightweight upsampling vocoder. Those stages can help explain why codec texture is worth checking, but they cannot identify the cause of a hiss, dull edge, or click in one file. Generation settings, edits, encoding, later processing, and playback can create similar symptoms.

Does original YuE have a hard 12 kHz cutoff?

The official YuE v1 paper and repository do not document a universal hard cutoff at 12 kHz. The paper describes reconstruction at 16 kHz followed by learned upsampling to 44.1 kHz. If your spectrogram changes sharply near 12 kHz, treat it as a property of that render and workflow until a controlled comparison proves more.

Can Sunofix clean YuE vocal hiss and xCodec haze?

Sunofix can process a lawful WAV or MP3 when an unwanted synthetic texture moves through a finished mix. It cannot recreate missing high-frequency information, correct lyrics or melody, rebuild true stems, or repair audio already destroyed by clipping or lossy encoding.

Can I use original YuE output commercially?

The preserved YuE v1 branch states that the model and weights use Apache 2.0 and encourages creators to incorporate outputs into commercial projects with attribution. That does not clear lyrics, reference recordings, voices, samples, or possible resemblance to existing work. Check the current license and every input before release.