Profile    Mohammed Shiroz Status   Loading  
Logo
Share This
Back to blog
Filter by:
Tags
//Article title

How Compression Works on the Web: gzip, Brotli and Why Repetition Is Free

About Post

Open your browser's network tab and look at a big JavaScript file. Two sizes: the one transferred, and the real one. The transferred number is often a small fraction of the real one.

Nothing was deleted. Your browser rebuilt the exact same file, byte for byte, from a much smaller package. That's lossless compression, and it's quietly saving bandwidth on almost every text file on the web.

The trick behind it is surprisingly human: don't say the same thing twice.

The big idea: repetition is free

Look at a typical chunk of HTML:

<div class="card"><h3 class="card-title">Unit A</h3></div>
<div class="card"><h3 class="card-title">Unit B</h3></div>
<div class="card"><h3 class="card-title">Unit C</h3></div>

The second and third lines are almost identical to the first. A compressor doesn't store them again. It stores something like "go back 59 characters and copy 46 of them", then the one letter that's different, then another short "copy from 59 back" for the closing tags.

That's the heart of LZ77, the algorithm family behind both gzip and Brotli: replace repeated sequences with short back-references to where they appeared before. Code, HTML, CSS and JSON are wonderfully repetitive (the same tags, keywords, property names and indentation over and over) which is why they compress so well.

Step two: short codes for common things

After the repetition is squeezed out, there's a second trick. In normal text, some symbols are far more common than others. Spaces and the letter e show up constantly; Q rarely does.

Huffman coding gives common symbols shorter bit patterns and rare ones longer patterns, like Morse code giving E a single dot. Each file gets codes built from its own symbol frequencies.

Put those two together, LZ77 then Huffman, and you have DEFLATE. gzip is DEFLATE plus a small header and a checksum. It has been the web's default compression for decades, and every browser and server supports it.

What Brotli does better

Brotli, developed at Google and standardised as RFC 7932, uses the same two core ideas and adds a few clever ones:

  • A built-in dictionary. Brotli ships with a static dictionary of common strings from real web content: HTML tags, common English words, CSS and JavaScript fragments. Even the first occurrence of something like </div> or function can be a short reference, because both sides already know it. That helps a lot with small files, where gzip has no earlier repetition to refer back to.
  • A bigger memory. DEFLATE can only refer back 32 KB. Brotli can refer back much further, so repetition across a large bundle is still caught.
  • Context modelling. It picks its coding tables based on what came just before, which squeezes text a bit tighter.

The result: for text assets, Brotli files are typically noticeably smaller than gzip ones. The catch is speed at the highest settings. Brotli's maximum level (11) is slow to compress, so it's meant for files you compress once, not on every request. Decompression is fast in both cases.

How the browser and server agree

Compression on the web is a negotiation, done with headers:

  1. The browser says what it understands: Accept-Encoding: gzip, deflate, br, zstd.
  2. The server picks one and labels the response: Content-Encoding: br.
  3. The server also sends Vary: Accept-Encoding, so caches and CDNs know to store separate versions for clients that accept different encodings.

Two details worth knowing: browsers only advertise Brotli over HTTPS, and newer encodings like Zstandard (zstd) are showing up in browsers too. The negotiation means you can add them without breaking older clients. Everyone gets the best format they understand.

You can check a site from the terminal:

curl -s -o /dev/null -D - -H "Accept-Encoding: br, gzip" https://example.com \
  | grep -i -E "content-encoding|vary"

What not to compress

Compression only works on redundancy, and some files have already had theirs removed.

  • Images, video and audio (JPEG, PNG, WebP, AVIF, MP4) are already compressed with formats designed for that media. Running gzip over them wastes CPU for little or no gain, and can even make them slightly bigger. Optimise images with image tools instead.
  • Fonts in WOFF2 are already compressed (WOFF2 uses Brotli internally).
  • ZIP files and PDFs are usually compressed already.
  • Tiny responses gain almost nothing, because the format's overhead eats the savings. Setting a minimum size of around a kilobyte is a common choice.

The rule of thumb: compress text (HTML, CSS, JavaScript, JSON, SVG, XML). Don't compress what's already compressed. Pre-compress static assets at the highest level at build time; compress dynamic responses at a moderate level on the fly.

Turning it on

On Nginx, gzip is built in. A sensible starting point:

gzip on;
gzip_comp_level 5;
gzip_min_length 1024;
gzip_vary on;
gzip_types text/css application/javascript application/json image/svg+xml;
# text/html is always included

Brotli on Nginx needs the separate ngx_brotli module. On Apache, mod_deflate handles gzip and mod_brotli handles Brotli. Many CDNs (CloudFront, Cloudflare and others) can compress for you at the edge, which is often the simplest option.

For static assets, compress once during your build, save app.js.br and app.js.gz next to app.js, and let the server pick the right file without compressing anything per request (gzip_static, and brotli_static with the module, do exactly this in Nginx).

A small security footnote

Compression has one well-known sharp edge: an attack called BREACH. If a page compresses a secret (like a token) together with text an attacker can influence, response sizes can leak information about the secret over many requests. It's not a reason to turn compression off; it's a reason to avoid reflecting user input next to secrets in compressed responses, and to use per-request masked tokens where they matter.

The short version

  • Compression replaces repetition with references, then gives common symbols shorter codes.
  • gzip is universal; Brotli is usually smaller for text, especially with pre-compression.
  • Headers negotiate the format; always send Vary: Accept-Encoding.
  • Compress text, skip media, ignore tiny responses.

Have you checked what your own site actually sends? Open the network tab, compare the two sizes on your biggest files, and tell me what you find.

Comments (0)
Leave your review

Thanks for your valuable comments. Your comments has been updated and appreciate your getting in touch...

01. About Shiroz

Mohammed Shiroz

Hi, I'm Mohammed Shiroz, a software engineer and AI enthusiast from Sri Lanka who turns ideas into intelligent, real-world solutions. With over 9 years of hands-on experience, I currently lead real estate ERP development at Kate Group, a...

03.My Projects

04. Categories

Ready To order Your Project ?

Get in Touch
Close