A minute of music or video contains a great deal of information. Store every audio sample or every pixel of every frame without any cleverness and files become enormous. Compression is the collection of tricks that lets a device describe much of that information more economically.
First, stop repeating yourself
Some patterns can be represented in fewer bits without losing anything. If a stretch of data repeats, a compressor can describe the pattern and how often it occurs instead of writing it out afresh. Other mathematical methods assign shorter codes to common values and longer ones to rarer values. A lossless decoder reverses the process and recovers the original data exactly.¹
That is valuable for archiving and editing, where every detail may matter. But there is a limit to how much ordinary recordings can be shrunk by finding repetition alone. To make streaming practical at lower bitrates, many formats go further.
What lossy audio leaves behind
A lossy audio codec models the signal and spends its limited bits where they will make the most perceptual difference. Human hearing does not treat every frequency and sound equally. A loud sound can make a quieter nearby sound harder to notice, a phenomenon codecs can exploit. The encoder may keep a more precise description of prominent elements and a rougher one of elements likely to be masked.²
This is not simply “removing the high notes”. Modern codecs use more elaborate transforms and bit allocation across time and frequency. At a generous bitrate, the changes may be hard to hear in ordinary listening. Push the bitrate too low and the trade-off becomes obvious: cymbals may swish, transients may smear, or a complex passage may sound oddly watery. Different codecs and settings fail in different ways.
Crucially, once information has been discarded, saving the file in a lossless format does not bring it back. Converting a heavily compressed recording into a larger file changes its packaging, not its missing detail.
Video has another dimension: time
A video is a sequence of pictures, but neighbouring frames often resemble one another. A codec can save a complete reference frame, then describe how blocks of the image move or change in later frames. It need not store the entire background afresh because a person has shifted a few pixels across it.³,⁴
Within a frame, video compression also uses spatial patterns: neighbouring pixels and broad smooth regions can be described more compactly than a list of unrelated values. Lossy coding then allocates detail according to the chosen quality target. At low bitrates, you may see blockiness, blurred texture, banding in a sky or mosquito-like noise around edges. Fast motion, confetti and foliage are harder to compress than a static presenter against a plain wall.
An encoder balances file size, quality and the work needed to compress and play the result. A newer codec may produce a smaller file at similar perceived quality but require more processing or lack support on older devices. A container such as MP4 is the wrapper that can hold audio, video and timing information; it is not itself the method that decides which detail to keep.³
Why settings matter
Bitrate is the amount of data used per second. More bits generally give a codec more freedom, though quality also depends on the source material, encoder and resolution. Repeatedly encoding a lossy file can compound damage, because each pass works from an already altered signal. For important originals, keep a high-quality master and make smaller delivery copies from it.
Why some scenes are harder than others
Compression is strongest when there is structure to describe. A steady camera pointed at a speaker has long stretches of similar background. The encoder can devote attention to a moving face while reusing much of the surrounding picture. A sudden cut to falling snow offers less repeated structure. At the same bitrate, a complex scene may show artefacts that were invisible a moment earlier.³
Audio has parallel challenges. A simple voice recording may be easy to keep intelligible at a modest bitrate. Dense percussion, reverberation and overlapping instruments give an encoder more competing detail to represent. The word “quality” therefore cannot be reduced to one file-size rule. It depends on the content and on what the listener notices.
Compression also has an energy cost. More sophisticated encoding may take longer and use more computing power, especially when a service must prepare many versions for different screen sizes and connection speeds. Decoding has to be practical on the audience’s devices. A technically efficient codec is useful only if the people receiving the file can play it.²,³
One further distinction helps when comparing formats: resolution is the number of picture samples, while bitrate is the amount of data devoted to describing them. A high-resolution stream starved of bits can look worse than a lower-resolution stream encoded well. The right question is not “How many pixels?” alone, but “How much meaningful detail survived the whole encoding process?”
Compression succeeds when it removes information the audience does not need for the purpose at hand. A podcast on a train, a film on a phone and an archival recording each ask for a different compromise. The tiny file is not magic: it is an agreed shorthand, and sometimes a carefully chosen omission.
