Huffman Coding is a lossless data compression algorithm that reduces the size of data by assigning shorter binary codes to more frequent symbols. It builds an optimal prefix tree, ensuring efficient encoding and decoding. Commonly used in file compression formats like ZIP and JPEG, Huffman Coding minimizes storage space without losing information.
The algorithm starts by treating every symbol as an independent node, each weighted by its frequency. It repeatedly merges the two minimum nodes into a new parent node whose weight is their sum. This merging continues until there is only a single tree. Traversing from the root to a leaf produces a binary code, where each left or right move adds a bit.
(5 credits)
Counts character frequencies, pushes symbol leaf nodes to a Min-Priority Queue. Repeatedly extracts the two smallest frequency nodes, combines them into a parent node with sum frequency, and re-inserts it. Leaves form prefix codes where frequent characters receive shorter bit sequences.
Building the tree takes O(N log N) time using a Min-Priority Queue, as there are N - 1 merge steps, each requiring queue insertion and extraction taking O(log N) time. If frequencies are pre-sorted, it can be built in O(N) using two linear queues.
Sign in to join the discussion
Hand-picked resources to deepen your understanding
© 2025 See Algorithms. Code licensed under MIT, content under CC BY-NC 4.0