Spectrogram Frequency Axes Explained: Hz, kHz, Time, and Color
Learn how to read time, Hz, kHz, and relative color in an audio spectrogram, why resolution changes the picture, and how to confirm every visual clue by listening.
Read the axes, then listen
A spectrogram has three practical coordinates: time runs left to right, frequency runs bottom to top, and color shows relative signal energy under the current display settings. Read those coordinates to find a short passage, then play that passage and decide with your ears. The image can tell you where to listen. It cannot decide whether the sound is musical, damaged, generated, or worth changing.
Start with one visible feature rather than the whole picture. Note its time, estimate its frequency area, check the color legend, and replay the matching seconds at a comfortable level. That simple sequence keeps a useful measurement from turning into an automatic diagnosis.
Time axis
The horizontal axis represents elapsed time. The left edge is the beginning of the file, and the right edge is the end. A mark at 0:30 belongs thirty seconds after playback starts; a mark at 2:15 belongs two minutes and fifteen seconds in. Some viewers label the axis only in seconds, so 135 seconds and 2:15 refer to the same point.
Use time as a navigation tool. If a narrow vertical streak appears at 1:08, start playback just before it. The short lead-in gives you musical context and helps you distinguish a click from a drum transient, a hard consonant, or an intentional edit. A picture of one instant rarely carries enough context on its own.
The width of a visual event is also measured against time. A one-frame click may look like a thin vertical line. A reverb tail occupies a wider area because its energy continues after the source sound. A steady noise floor can stretch through verses, pauses, and choruses. These shapes suggest different listening tests, but none supplies a cause by itself.
Zoom changes the amount of time shown on screen. A ten-second selection spreads small events across more pixels than a four-minute overview. The audio has not changed. Before comparing two screenshots, match the start time, end time, and image width or accept that one view may make the same event look thicker.
I usually check the busiest chorus first. That is where a synthetic edge tends to stop hiding behind the arrangement, and it is also where broad processing can take away the most wanted detail. I use the time axis to mark the passage, then I switch my attention back to playback.
The free browser spectrogram generator can create a fixed PNG from a local MP3 or WAV. Keep the audio open beside the image and write down the time of each feature you want to test. The tool does not send the selected file to an upload API or pass it to the commercial app.
Frequency axis in Hz and kHz
The vertical axis represents frequency: how quickly air pressure oscillates, expressed in hertz. One hertz means one cycle per second. A kilohertz is 1,000 hertz, so a label of 8 kHz means 8,000 Hz. Spectrograms use both units because audio covers a broad range and labels such as 12 kHz are easier to scan than 12,000 Hz.
Low frequencies appear near the bottom. Bass weight, kick fundamentals, and low hum often occupy this area, though every instrument can spread beyond a single band. The middle contains much of vocal intelligibility, guitars, keys, snare character, and dense harmonic structure. Higher areas can contain cymbal detail, vocal air, hiss, distortion, and brittle edges. These are listening landmarks, not fixed boxes for individual instruments.
A vocal does not appear as one frequency. Its fundamental pitch may sit low or in the middle, while harmonics extend above it. Consonants and breath add noise-like energy higher up. A cymbal spreads across a wide range, and a distorted synth can fill much of the display. Avoid pointing at one stripe and naming an instrument without listening.
Check whether the vertical scale is linear or logarithmic. On a linear scale, every equal number of hertz receives equal space: the distance from 1 to 2 kHz matches the distance from 10 to 11 kHz. A logarithmic scale allocates space by ratios, which often makes octaves and lower musical regions easier to inspect. Audacity’s official manual shows that the same audio can look quite different under linear, logarithmic, Mel, and other frequency scales.
Neither scale reveals more truth about the source. Each organizes the same analysis for a different task. A logarithmic view may make bass harmonics easier to separate visually, while a linear view can give more room to upper-frequency detail. Record the scale before comparing renders.
The top of the axis is also a display boundary, not a quality score. A dark area near the top can reflect the source, the sample rate, an earlier export, or the viewer’s chosen maximum frequency. It does not establish which encoder, generator, or internal process produced the file. If the upper boundary concerns you, compare known versions and listen before making any change.
Color and relative energy
Color represents the measured magnitude or power at each time-frequency position. Many palettes map low energy to dark colors and stronger energy to warm or bright colors. That convention is common, but it is not universal. Read the legend instead of assuming that yellow, red, blue, or white always means the same value.
Audio spectrograms often convert their measurements to decibels. The librosa documentation makes the reference explicit: its amplitude-to-decibel conversion compares values with a selected reference and can hide everything below a chosen range. Audacity likewise exposes gain and range controls that change which levels receive each color. A faint line can become vivid when you raise display gain even though the file remains untouched.
Treat color as relative evidence under fixed settings. Within one view, you can compare a quiet verse with a dense chorus or see whether a narrow band continues through a pause. Between two files, use the same palette, reference, gain, dynamic range, frequency scale, window, and image dimensions. Otherwise, a color difference may belong to the display rather than the audio.
Brighter does not mean cleaner, louder to a listener, or more important to the song. A wanted snare hit can create a strong vertical flash. A soft but irritating whistle may be less bright. A high-frequency layer can look dramatic while sitting at a level you cannot hear in context. The useful question is not “Which image has more color?” It is “What do I hear at this time, and does the display help me locate it again?”
Color also cannot identify a source. Similar bright bands can come from a harmonic, a sustained synth, electrical whine, feedback, or resonant ringing. Similar high-frequency clouds can come from cymbals, breath, ambience, hiss, or deliberate distortion. Use plain listening language before you choose a repair.
Time/frequency resolution trade-off
A spectrogram is built from short overlapping slices of audio. The short-time Fourier transform, or STFT, applies a window to each slice and estimates its frequency content. SciPy’s ShortTimeFFT documentation describes the result as a grid whose columns correspond to time positions and whose rows correspond to frequency bins.
The slice length creates a trade-off. A longer window separates nearby frequencies more clearly, which can help when you are looking for a steady tone. The same long window blurs quick events across time. A shorter window places clicks and drum attacks more precisely, but nearby frequencies blend into wider shapes. Audacity demonstrates this directly: a larger window narrows a tone while two close clicks smear together.
This matters because a fuzzy edge is not automatically a fuzzy sound. It may reflect the chosen window. A thin horizontal line can widen under one window type, and a sharp transient can stretch under another. Overlap and hop size affect how often the analysis samples a new time position. Zero padding can make the color interpolation look smoother without escaping the underlying time-versus-frequency limit.
Choose settings for the question. Use finer time resolution when locating clicks, edits, consonants, or short drum events. Use finer frequency resolution when examining a steady whistle, hum, or narrow resonance. For a before-and-after comparison, keep every setting fixed. Changing the window between renders removes your common visual reference.
You do not need to optimize an FFT to make a useful first pass. Note the current settings, find a repeatable feature, and listen. If a repair decision depends on whether two close tones are separate, then a more detailed analyzer view may help. If the problem is already obvious in playback, spending ten minutes tuning colors adds little.
Reading workflow
Turn one visual clue into one controlled listening question:
- Keep the original file and analyze the exact version you intend to review.
- Check the time span, frequency scale, maximum frequency, color legend, gain, range, and window settings.
- Pick one broad feature and note its time rather than scanning for every unusual shape.
- Describe what the shape might help you hear: click, steady tone, hiss, glassy vocal edge, splashy cymbal tail, or wanted bright synth.
- Loop ten to twenty seconds around the marked time at a comfortable level.
- Lower the playback level once and check a second familiar playback system.
- Compare a nearby passage where the feature is absent.
- If the sound remains distracting, make the smallest relevant change on a copy.
- Match perceived loudness before switching between original and processed versions.
- Keep the change only when the named problem recedes without shrinking the song around it.
For a stable narrow resonance, a modest targeted EQ reduction can be a useful manual test. For broad upper-frequency fatigue, follow the restrained workflow to fix harsh highs in AI music. Stop when the treatment pushes the vocal backward, dulls consonants, thins cymbals, breaks reverb tails, or becomes easier to hear than the original problem.
If the artifact moves through a finished mix, repeated broad cuts may remove more music than defect. I built Sunofix for that stage: the song already works, but the exported audio still needs cleanup before mastering. Its before-and-after playback and frequency diagnostics can support the same process, provided you compare the same passage at matched loudness and keep the original close.
Read the axes, then listen if you want to compare a track-aware cleanup pass with the source.
Sunofix can reduce some broad metallic and harsh textures in an existing MP3 or WAV and return a cleaner WAV for review. It does not isolate every instrument in a stereo mix, repair arrangement or performance choices, reconstruct missing source detail, decode a picture into sound, or finish mastering. It also does not change lyrics, melody, or composition. The full artifact-removal guide can help you decide whether the next step should be cleanup, stem work, regeneration, mastering, or no change.
No automatic diagnosis
A spectrogram measures a file under chosen settings. It does not know the musical intention, the production history, or which detail bothers you. It cannot determine that a track came from a particular generator. It cannot reveal a closed model’s internal process. It cannot tell whether a bright band is an artifact or the exact cymbal texture the song needs.
Avoid three common shortcuts. First, do not treat every horizontal line as hum; sustained notes and harmonics also form lines. Second, do not treat every bright upper area as harshness; wanted air and percussion can be strong there. Third, do not treat a dark upper boundary as a forensic codec result; the same appearance can follow several export and display choices.
Listening supplies the missing context. Replay the marked passage, identify the sound in ordinary words, and decide whether it remains distracting at moderate and low levels. Compare on another familiar system. If you cannot hear a repeatable problem, leave the track alone. Processing a silent visual suspicion can make a good song smaller.
When you do test a change, judge the result by preservation as well as reduction. Check vocal diction, cymbal decay, bass weight, transient shape, ambience, stereo movement, and emotional lift. A darker or smoother image is not a win if the chorus loses life. The best use of the axes is modest: they help you return to the right few seconds and repeat a fair comparison.
Read time, frequency, and color as coordinates, not conclusions. Once they have pointed you toward a useful listening decision, their job is done.
Continue listening
Related reading
FAQ
Spectrogram Frequency Axes Explained: Hz, kHz, Time, and Color FAQ
What do the axes of a spectrogram show?
The horizontal axis shows time, while the vertical axis shows frequency in hertz or kilohertz. Each colored position represents measured energy near one frequency during one short slice of time.
What is the difference between Hz and kHz?
Hz means cycles per second. One kilohertz equals 1,000 hertz, so 5 kHz is the same as 5,000 Hz. Spectrogram viewers use both units to keep labels readable across a wide frequency range.
Does a brighter color mean better sound?
No. Brighter color usually means more measured energy relative to the current viewer reference, gain, and range. Wanted cymbals, vocal air, noise, and an unwanted metallic edge can all appear bright.
Can a spectrogram diagnose an AI music artifact automatically?
No. It can help you locate a time and frequency area worth replaying, but the picture cannot decide whether the sound is unwanted or reveal the generator, model, or hidden production cause.
