Article
Disaster! Dirty Data, Data Misinterpretation, and Nonsensical Actions
I have a peculiar relationship with messy data. On one hand, it can drive me mad; there are times when I explode like a volcano after realising that decisions were made incorrectly because of it. Especially when it has been going on for a long time and has had a significant negative impact. And particularly when I realise that I could have identified the issue much earlier. On the other hand, resolving such situations pulls me into a state of flow, a state where I immerse myself deeply into the problem and shut out the outside world.
I have realised that I would probably be bored in a place where everything is perfectly organised, all information is accurate, everyone knows precisely what the data tells us, and everyone can place it into the broader context of the organisation.
One example of such a place is an overly simple organism. In the past, when I was approached with a job offer from one of the most renowned Slovak e-commerce projects, I was not interested in changing jobs. But at the same time, I asked myself: what could I significantly contribute there? It is just too simple a business: they buy and sell, buy and sell. It did not excite me at all. I saw no intellectual adventure in it, no opportunity to dive into entirely new situations and learn something new while solving them.
My place is elsewhere, in the jungle. Somewhere that at first glance seems chaotic and impossible to navigate. A place where most people only know their specific area of expertise. And that is when the work becomes enjoyable for me. That is when half a day flies by like half an hour.
This “jungle” can look very different depending on the situation. I do not want to go into specifics; I will keep it general, though I realise that might make it less clear for the reader. It starts with it being one of those more complex organisms. And in such an organism, you might encounter four types of challenges related to working with information and the disasters that can arise from them. Many of these disasters begin with one confusion: mistaking correlation vs causation, reading a cause into data that only shows two things moving together.
Join the Library
Full access to my findings, personal stories, and what I learn from the people I meet.
Join the Library · €29.99 per year How to Start With AI in a Company: Automating… Fable 5 vs Opus: I Ran the Same Audit… I Ran Object Detection on My Laptop, and Saw… Bot or Human? What My… Dependent on AI: Are We… Which Work Will AI Not Replace? Not Yet…Get the full article by email and feel free to reply if you want to discuss it further.
Summary
Common questions on this article's topic
What are the main types of data quality problems in organisations?
Why is data misinterpretation more dangerous than missing data?
What does it mean to act on data without context?
How does flow state relate to solving complex data problems?
Why do some people thrive in messy data environments?
How can organisations reduce the risk of data-driven disasters?
What is garbage in, garbage out (GIGO)?
What is dirty data?
What is data literacy and why does it matter?
What is the difference between data quality and data integrity?
What is data governance?
What is data interpretation and why does it go wrong?
Why do data driven decisions still go wrong?
Related articles

This project changed the way I think about generative AI.
He used words like “certainty” as if statistics were part of Newtonian physics.
This is a serious issue, and it is high time we start acting responsibly.
More articles
One of my family members described a firm that has its data in several different tools, and its employees still join it by hand, in spreadsheets. That is the ordinary state of firms. When somebody tells such people not to do it by hand and to use AI, they will not understand. Most firms do not even have the basic connectors between the tools they use. These firms need help with AI transformation.
The same task, two models. Fable 5 against Opus 4.8. On paper Fable is the better model, with a larger context and stronger specs. And still it lost. Opus handled the task with a single round of checking for 721,000 tokens, while Fable needed nine rounds and burnt through 2.78 million tokens. The difference was not in the model, but in how I set the task. And I know it, because I measured it.
A few weeks ago I installed a small local AI model on my laptop that watches a live camera feed. I turned the webcam on in the dark, and in near total darkness it recognised me and the objects in the room. That such things exist, I have known for a long time. What opened my eyes was the accessibility. I installed it in one prompt, free, and it runs entirely on my machine, sending data nowhere.

I once wrote about building my own privacy-friendly analytics tool. It had bot detection from the first version, yet it was not enough. Direct visits took a strangely high share of my traffic. When someone claims that 20% of their visits are bots and 80% are humans, I used to think the same. Today I would say the opposite ratio is closer to the truth. This is how I got there.

I have Heidegger and my notebook beside me. I am asking where all of this is heading, where artificial intelligence is taking us.
Seventy per cent. That is where the first AI output begins, even when you give it the full company context and the best examples from the past. We are talking about the kind of output that cannot be defined programmatically. It is more complex. Often it is creative work. On one repeated type of output I reached eighty per cent within a week. Every further percentage point is harder than the one before.
For a long time we treated the internet as the main road. The place where work and relationships happen. Yet most of what we see on it today is, or soon will be, AI-generated: text, images, profiles and comments. The internet is turning into an online game full of bots, where you cannot be sure that a human is on the other side of anything. So I ask: was the online world the main road, or only a temporary detour that part of us will return from, back offline?
A few days ago I interviewed a senior marketer. An experienced man, years of practice. I asked him about AI. He said he barely uses it. He had one bad experience with the output and decided he was too senior for it to add value when it is not perfect. I know the other side too: professionals who automate everything that can be automated.
Europe does not have the capacity to face a full-scale, mass drone war of the kind we see in Ukraine. Three dependencies weaken it: China supplies the physical material for defence systems, the United States supplies capabilities Europe does not have, and twenty-seven states cannot agree how fast, or who pays. Rearmament plans exist, but they are being carried out slowly.
AI produces the graphic, the newsletter and the product page faster than a person. What is left for the one who used to do it is the judgement, knowing whether the output is good. But most people have worse judgement than AI. And whoever cannot judge quality cannot delegate either. How do you tell whether yours is the judgement a company relies on, or the kind it can replace?

Four days in Catalonia. No computer, no AI, almost no social media. I bought this notebook so that I could write down what I would think about, and what I would come across and learn on the trip.

