Article
Disaster! Dirty Data, Data Misinterpretation, and Nonsensical Actions
I have a peculiar relationship with messy data. On one hand, it can drive me mad – there are times when I explode like a volcano after realising that decisions were made incorrectly because of it. Especially when it’s been going on for a long time and has had a significant negative impact. And particularly when I realise that I could have identified the issue much earlier. On the other hand, resolving such situations pulls me into a state of flow – a state where I immerse myself deeply into the problem and shut out the outside world.
I’ve realised that I’d probably be bored in a place where everything is perfectly organised, all information is accurate, everyone knows precisely what the data tells us, and everyone can place it into the broader context of the organisation.
One example of such a place is an overly simple organism. In the past, when I was approached with a job offer from one of the most renowned Slovak e-commerce projects, I wasn’t interested in changing jobs. But at the same time, I asked myself: what could I significantly contribute there? It’s just too simple a business – they buy and sell, buy and sell. It didn’t excite me at all. I saw no intellectual adventure in it, no opportunity to dive into entirely new situations and learn something new while solving them.
My place is elsewhere – in the jungle. Somewhere that at first glance seems chaotic and impossible to navigate. A place where most people only know their specific area of expertise. And that’s when the work becomes enjoyable for me. That’s when half a day flies by like half an hour.
This “jungle” can look very different depending on the situation. I don’t want to go into specifics; I’ll keep it general, though I realise that might make it less clear for the reader. It starts with it being one of those more complex organisms. And in such an organism, you might encounter four types of challenges related to working with information and the disasters that can arise from them. Many of these disasters begin with one confusion: mistaking correlation vs causation, reading a cause into data that only shows two things moving together.
Join the Library
Full access to my thoughts, personal stories, findings, and what I learn from the people I meet.
Join the Library · €29.99 per yearGet the full article by email and feel free to reply if you want to discuss it further.
Summary
Common questions on this article's topic
What are the main types of data quality problems in organisations?
Why is data misinterpretation more dangerous than missing data?
What does it mean to act on data without context?
How does flow state relate to solving complex data problems?
Why do some people thrive in messy data environments?
How can organisations reduce the risk of data-driven disasters?
What is garbage in, garbage out (GIGO)?
What is dirty data?
What is data literacy and why does it matter?
What is the difference between data quality and data integrity?
What is data governance?
What is data interpretation and why does it go wrong?
Why do data driven decisions still go wrong?
Related articles
This project changed the way I think about generative AI.
He used words like “certainty” as if statistics were part of Newtonian physics.
This is a serious issue, and it is high time we start acting responsibly.
More articles
A few weeks ago I installed a small local AI model on my laptop that watches a live camera feed. I turned the webcam on in the dark, and in near total darkness it recognised me and the objects in the room. That such things exist, I have known for a long time. What opened my eyes was the accessibility. I installed it in one prompt, free, and it runs entirely on my machine, sending data nowhere.
I have Heidegger and my notebook beside me. I am asking where all of this is heading, where artificial intelligence is taking us.
Seventy per cent. That is where the first AI output begins, even when you give it the full company context and the best examples from the past. We are talking about the kind of output that cannot be defined programmatically. It is more complex. Often it is creative work. On one repeated type of output I reached eighty per cent within a week. Every further percentage point is harder than the one before.
For a long time we treated the internet as the main road. The place where work and relationships happen. Yet most of what we see on it today is, or soon will be, AI-generated: text, images, profiles and comments. The internet is turning into an online game full of bots, where you cannot be sure that a human is on the other side of anything. So I ask: was the online world the main road, or only a temporary detour that part of us will return from, back offline?
A few days ago I interviewed a senior marketer. An experienced man, years of practice. I asked him about AI. He said he barely uses it. He had one bad experience with the output and decided he was too senior for it to add value when it is not perfect. I know the other side too: professionals who automate everything that can be automated.
Europe does not have the capacity to face a full-scale, mass drone war of the kind we see in Ukraine. Three dependencies weaken it: China supplies the physical material for defence systems, the United States supplies capabilities Europe does not have, and twenty-seven states cannot agree how fast, or who pays. Rearmament plans exist, but they are being carried out slowly.
AI produces the graphic, the newsletter and the product page faster than a person. What is left for the one who used to do it is the judgement, knowing whether the output is good. But most people have worse judgement than AI. And whoever cannot judge quality cannot delegate either. How do you tell whether yours is the judgement a company relies on, or the kind it can replace?
In April, in the first part of this series, I wrote about an AI prediction system I had started building on my own machine. At the time the software was a few hours old and the prediction record was empty. The record since then has shown one thing: the system does not yet understand the market it is being asked to forecast. It can pull macro context, book value, earnings. But it cannot put those together into something that helps it understand the price.
Prague, 13 May 2026. On my way to work I started thinking about something that stayed with me for days. If most routine work on a computer disappears in the next ten years, and a large share of repetitive manual work disappears with it, what happens to the flow of money? Who pays whom for what? Which economic layers will exist, how large will they be, and what relationships will run between them? This is the six-layer map I sketched as an answer.
I am building an AI system to predict the S&P 500. It runs on my own machine, uses free public data (yfinance, FRED, the Shiller dataset), and grades every forecast against reality. This series documents the build itself: the decisions, the methodology, the mistakes. What I will eventually share from the running system is a separate question, and an honest one.
Four days in Catalonia. No computer, no AI, almost no social media. I bought this notebook so that I could write down what I would think about, and what I would come across and learn on the trip.
