In digital media, lossy compression is the trade-off we make for convenience - we accept the loss of some data from a file to make it small enough to send. I think that we do the same thing to our thoughts every time we open our mouths. We send out carefully crafted sentences, only to find they arrive at their destination scrambled, with half the meaning lost. This isn’t a personal failure but a fundamental limitation of language itself.
That is, language is a lossy compression algorithm, and the person listening to you speak is interpreting your words with an entirely different context than you.
What Do You Mean, Lossy?
Think of your thoughts as a high-quality, complex and multi-sensory file sitting in your brain. It’s a 4K, 120fps video with all your reasoning behind it.
But you can’t transmit this thought in all its glory to the person you’re speaking to. That would take hours, and you probably don’t want them knowing every step you took to finalise the thought anyway. So, you choose your words carefully to find the best way to express yourself - to transmit the high-quality video.
You choose your vocabulary, your grammar, and which information to drop for the sake of conciseness or convenience. But what information is lost in the compression? Emotional nuance, sensory detail, most of the context, and very often the tone.
When the person listening hears what you say, they have to ‘decompress’ it to interpret what it all means. The key lies here - they have their own algorithm. They don’t share your experiences, your mood, and your intention. They will deconstruct what you have said based on their own worldview. The final image they recreate in their head will never be identical to your original.
This is where cognitive bias comes in. People will be actively looking for data that confirms their existing beliefs and discarding the rest, leading to a completely different recreated image than the one the speaker intended. Two people can even listen to the same speech and come to two entirely different conclusions - because of their pre-existing conceptions and biases, and because of how they deconstructed the language of the speech in their mind.
All of this is only worsened if the conversation is taking place by text message, or email. The algorithm becomes aggressively lossy. Facial expressions, tonal shifts in your voice, pacing of words - these are all lost when you are writing messages out on your phone’s keyboard.
Think about the last time you tried to make a joke in a group chat. Assuming you didn’t use tone indicators, how did you convey that it was a joke? Or sarcastic? Or meant to be taken lightly? Presumably, you assumed the recipients would infer that from context. This works with friends, but when cultures clash, or generations have different expectations of humour, this can result in awkward misunderstandings. Modern communication has attempted to re-introduce emotions and tone to language via emojis. Sending a smiley face :) or a laughing emoji 😂 is a tiny packet of information that replaces the expected body language or facial expressions that were intended to go along with the message sent.
Take the commonly-sent text message ‘OK.’ Lots of people won’t give it a second thought, but many more people will take it to heart as a personal attack, an expression of disinterest or muted anger. This is the subtle nuance we lose without the metadata of tone and body language that is carried when communicating in real life.
How about describing a dream? Dreams are an excellent example of an internal experience that inevitably becomes absurd and nonsensical when you try to compress it into communicable language. You can’t convey dreams properly via speech, you can’t express how you physically felt at the time - we will never have all the words to describe the human experience.
Different Languages, Different Codecs?
This idea, that the tools our language gives us can shape what we are able to express, relates to a fascinating linguistic concept - the Sapir-Whorf hypothesis, commonly known as linguistic relativity.
In short, it suggests that the ‘compression algorithm’ we are given (our native language), doesn’t simply convey our thoughts - it can actually shape the thoughts we’re able to have in the first place. This guarantees a certain amount of ‘loss’ when communicating between different linguistic systems - if English doesn’t have a word for a concept, do English speakers simply think of that concept less often?
The commonly cited claim in pop linguistics that Inuit indigenous people have hundreds of words for snow1 is actually a persistent myth that has been thoroughly debunked by linguists. First and foremost, English already has powder, slush, sleet, frost, flurries, blizzards, etc. The myth points to a feature of polysynthetic languages - they can create very long and complex words by adding different suffixes to a root word to describe different versions of a concept. English is not polysynthetic, in that we rely on adding helper words instead - but I think this is largely a meaningless comparison, because we can produce essentially the same result, just with more spaces.
For example: imagine your workplace invents a theme day: wear a hat that you inherited from your grandparent day. Nobody hears that and thinks the day belongs to your grandparent, or that your grandparent is named Day - every word before day is acting the role of a suffix in a polysynthetic language - attaching extra meaning to the root word. We are still building one complex idea by stacking parts. The difference is largely orthographic (how we write it) rather than conceptual.
I’m going to temporarily dive back into pop linguistics, give you two examples of “untranslateable” words, and then proceed to translate them.
Mamihlapinatapai(from the South American language Yaghan) means the situation when multiple people both want something to happen, but neither want to start it.Sumak Kawsay(a neologism from Quechua) is “good living”, but encompasses a broader idea of a balanced life in relation to community and nature.
The reason I bring these up is not to say that nobody can translate them - I just did. It is to bring attention to the circular fact that since English has not needed a single word for these, we do not express this feeling as often in our daily lives, which leads us to think of it less frequently. Packaging the whole feeling above into the single word of Mamihlapinatapai may mean that it becomes easier to notice or discuss. And sometimes the stakes are bigger - Sumak Kawsay is important enough in parts of the Andes that it has been written into the constitutions of Ecuador and Bolivia 2.
There also exists the concept of high-context and low-context language cultures3. Low-context cultures expect direct and explicit communication, while high-context cultures rely on shared backgrounds. A businessman in a high-context culture might find it desirable to create a working relationship with a customer before attempting to sell them something. A customer in a low-context culture might find this intimidating or unusual, in an environment where privacy and individuality are more valued. This is of course a simplification, but it is nonetheless interesting - when you communicate with your family, you will switch to a ‘high context’ environment, where you can relax and speak in shorter sentences, with the comfort that your conversation partners share lots of context with you. However, when with strangers, you may switch to low-context - intentionally over-communicate, or spell things out, to prevent miscommunication between someone that you don’t expect to understand your nuance or subtleties.
When Is It Useful?
This incredibly advanced capability of human language is not all bad. We are very good at efficiently utilising ‘codecs’ - shortcuts to convert data streams into simpler, more understandable ones. If you were to write down the tone of voice, the facial expressions, the body language - and deliver this packaged with the actual content of the message - this would likely overwhelm us. But when we are talking to someone in person, we all do this daily and subconsciously, adapting our behaviour and our interpretation of what we are being told, because of all those factors.
When you have a deep connection or a shared history with someone, like a partner, best friend, or sibling, you develop a highly efficient, personalised, and partially shared compression system. An inside joke, for example, is a tiny packet of information that can decompress into a huge, shared memory - while not making sense to an outsider, purely because of your shared contexts.
As we grow, so does our compression algorithm - a trained lawyer will be much more proficient at summarising legal briefs (without losing key data!) than a person on the street. We tweak our algorithms to filter out all unnecessary information, and at the same time keep the important information. This is exactly how lossy algorithms work in digital media, too.
We need to realise that communication is inherently flawed, and that’s okay. Improving our algorithm will help us be more patient listeners, and prevent all sorts of misunderstandings. Active listening, where you constantly ask clarifying questions, is a form of ‘error-correction’ - by intentionally trying to understand the original intent of the person, you can simulate a confirmation handshake - you are checking that what you received is the same as what they intended to send.
Ultimately, I believe seeing language this way is a step towards a more empathetic society. We all need to try our best to understand the other person - and we all need to try to communicate cooperatively, hopefully without losing too much of our original intent.