What Is a Hash? (And Why MD5 Had Its Day)
September 29, 2026 · 6 min read
Download a Linux ISO and the site lists a long hex string next to it — sha256: 9f86d081884c7d659... — with instructions to "verify the checksum." Git commit IDs are hashes. When you log in somewhere, the server compares hashes, not your actual password. Hashing is one of the most-used and least-understood ideas in computing. The core concept, though, is simple: a hash function takes any input and produces a fixed-size fingerprint of it.
The three promises of a hash function
A cryptographic hash like SHA-256 makes three guarantees:
- Deterministic. Same input, same output, every time.
helloalways hashes to2cf24dba5fb0a30e26e83b2ac5b9e29e1b161e5c1fa7425e73043362938b9824. No randomness, no exceptions. - One-way. Given the hash, you can't reconstruct the input. There's no "unhash" button. The only way to find an input matching a hash is to guess inputs until one matches — and with 2²⁵⁶ possibilities, you'll be guessing for longer than the universe has existed.
- Avalanche. Change one character of the input and the output changes completely and unpredictably.
helloandhellpproduce hashes with nothing in common. There's no "close" — either every bit matches or it doesn't.
That combination — deterministic, irreversible, chaotic — is what makes hashes useful as fingerprints. Two files with the same SHA-256 hash are, for all practical purposes, the same file.
What people actually use hashes for
- Verifying downloads. That ISO checksum: hash the file you downloaded, compare with the published hash. Match means the download wasn't corrupted (or tampered with). One flipped bit anywhere in a 4 GB file changes the hash entirely.
- Storing passwords. Good systems never store your password — they store
hash(password + salt). When you log in, they hash what you typed and compare. A database breach then leaks hashes, not passwords. (Salting — adding a unique random value per user before hashing — defeats precomputed "rainbow table" attacks. And for passwords specifically, slow hashes like bcrypt or Argon2 are preferred over SHA-256, precisely because SHA-256 is fast and attackers love fast.) - Git. Every commit ID is a hash of the commit's contents. Change history and the IDs change — that's how Git detects tampering and how content-addressing works.
- Deduplication. Cloud storage services hash file chunks; identical hashes mean identical data, stored once.
- Data structures. Hash tables (the thing behind Python dicts and JS objects) use non-cryptographic hashes for speed — same idea, weaker guarantees.
So what happened to MD5?
MD5 was the hash function of the 90s and 2000s. Then researchers found collisions — two different inputs producing the same MD5 hash. At first it took serious computing power; by the late 2000s, researchers were generating colliding files on a laptop, and someone famously created a rogue CA certificate using an MD5 collision. A hash function where attackers can craft collisions is broken for security purposes. Full stop.
Does that mean MD5 is useless? Not quite. For non-security uses — checksums where nobody is attacking you, hash tables, quick file identification — MD5 is fast and fine. The rule is simple: if an adversary could benefit from faking a hash, don't use MD5 (or SHA-1, which fell the same way). For anything security-related, SHA-256 is the current sane default, with SHA-3 and BLAKE3 as modern alternatives.
One more misconception to kill
Hashing is not encryption. Encrypted data can be decrypted with the key; hashed data can't be "decrypted" at all. When someone says "the passwords are encrypted with MD5," they're confused twice over — it's hashing, not encryption, and MD5 shouldn't be near passwords anyway. Words matter here because they imply very different security properties.