There's a version of you that exists somewhere you've never actually visited. It doesn't know your name. It's never heard your voice. It has no idea what you were thinking on any particular Tuesday, or why you were awake at 2am searching for something you'd rather not explain.
But it has a record of what you do. And in some surprisingly specific ways, that turns out to be enough.
That's the interesting part. Not so much how much platforms can record about us, a conversation we're already fairly familiar with, but the question underneath it: what can a system actually know about a person simply by watching what they do? And how close does that get to knowing who they are?
The answer turns out to be both more impressive and more limited than you might expect.
The Raw Material: Actions, Not Thoughts
hink about how many small decisions we leave behind every day. A search. A Like. A song we play again. A video we abandon halfway through. An app we tend to open at the same time every morning. Individually, they can seem almost meaningless. Together, they start to form something much more interesting.
In 2013, Michal Kosinski, David Stillwell and Thore Graepel analysed the Facebook Likes of 58,466 volunteers who had also provided demographic information and completed psychometric tests [1]. They wanted to see how much could be inferred about a person from something as simple as the pages they'd chosen to Like.
Quite a lot, as it turned out.
Using those Likes, the models could distinguish between some personal attributes with considerable accuracy. Among the published results, they distinguished between homosexual and heterosexual men in 88% of cases, between Democrats and Republicans in 85%, and also found predictive signal for personality traits and other attributes [1].
That doesn't mean a Like contains a tiny psychological explanation of the person who clicked it. That's precisely what makes the result interesting. The model didn't need to ask why someone had pressed that button. It looked for regularities across many actions and many people.
The raw material was still behaviour.

When the Pattern Actually Works
t would be easy to dismiss Facebook Likes as a particularly strange case, but the same principle appears when we look at much more ordinary behaviour.
In 2020, Clemens Stachl and his colleagues followed 624 volunteers for 30 days and analysed the traces they left while using their smartphones [2]. They examined patterns related to communication, mobility, app use, music consumption, overall phone activity, and behaviour at different times of the day.
They then compared those patterns with participants' scores on the Big Five, a widely used model for describing dimensions of personality.
The phone data contained enough signal to predict those personality dimensions above chance [2].
And that changes the question slightly. We're no longer talking about a platform knowing that you pressed Like on a particular page. We're talking about patterns that emerge simply because you carry a phone around and use it as you normally would.
Those inferences don't even have to be perfect to be useful.
Matz, Kosinski, Nave and Stillwell took that idea into a different setting: digital advertising [3]. Across three field experiments reaching more than 3.5 million people, they compared advertising messages designed to match particular psychological profiles. When the message matched the inferred profile, they observed up to 40% more clicks and up to 50% more purchases in some of their comparisons [3].
The result matters here for a fairly simple reason. For those differences to appear, the system didn't need to reconstruct each person's entire psychological story or know why they responded to the ad. The inference only needed to be useful enough to improve the choice of message.
Which raises a slightly strange possibility: a system can be right about us without necessarily understanding us.

The Context the System Never Had
Now imagine spending a few weeks planning a trip to Japan. You search for Tokyo neighbourhoods, trains, hotels, restaurants, basic Japanese phrases, and probably end up watching more videos about where to eat ramen than you realised existed.
Looking only at your behaviour, the signal seems fairly clear. Japan occupies a significant part of what you're searching, reading, and watching during those weeks.
But perhaps the trip isn't even yours. You're helping your best friend organise her honeymoon.
Nothing about what you did changes. The searches are still yours. The clicks happened. You spent time comparing hotels and learning how the Shinkansen works. The recorded behaviour is completely real.
What isn't necessarily contained in that trail is the situation that produced it.
This matters because our behaviour isn't simply a direct translation of stable preferences. Icek Ajzen's Theory of Planned Behavior distinguishes between wanting to perform an action and having the conditions to actually carry it out. Intention is related to our attitude towards the behaviour, the social norms we perceive, and the control we believe we have over the situation. Real-world factors can also make it easier or harder for an intention to become an action [4].
Warshaw and Davis made a related distinction between behavioural intention, what someone intends to do, and behavioural expectation, what that person thinks they will actually end up doing given their circumstances [5]. In their study, expectations were, overall, better predictors of subsequent behaviour than intentions.
We don't need to turn either of those theories into an explanation of how algorithms work. They weren't developed for that. What they do show is something more basic: what we do isn't a transparent window into what we want. Circumstances are part of the path between the two.
A record can capture the action perfectly without containing the full situation that made the action make sense.
When a Signal Works Without Making Sense
his is where the Facebook Likes study becomes even more interesting.
Some of the associations Kosinski and his colleagues found made fairly intuitive sense. Others seemed to make almost none.
One of the more curious examples was Curly Fries, a Facebook Page devoted, quite literally, to curly fries. At the time, Liking a page was one of the small signals recorded on a user's profile. And surprisingly, Liking the Curly Fries page appeared in the data as being associated with higher intelligence scores [1].
Obviously, liking curly fries doesn't make anyone more intelligent. Nor is there an obvious psychological reason why the two things should be connected.
The explanation proposed by the authors is much more interesting for our purposes. Some signals can become predictive because of the way they spread socially. A page might initially become popular among a group of people who share certain characteristics and then spread through their connections. With enough data, that Like can become statistically useful even though the page itself has absolutely nothing to do with the thing it helps predict.
That creates a rather peculiar kind of information.
If I tell you someone reads a lot of philosophy books because they're interested in philosophy, I've given you an explanation. If I tell you that Liking a page about curly fries slightly improves my chances of predicting another characteristic about that person, I've given you something different.
A correlation that works.
For us, the absence of an explanation can feel unsatisfying. For a system whose job is to improve a prediction, it may not be a problem.
And that's where the distance between finding a signal and understanding what it means starts to become much easier to see.
We Live in Episodes. Data Records Actions.
There's another complication when we add time to the picture.
Imagine spending a month organising a wedding. Suddenly your searches are full of venues, dresses, flowers, photographers, invitations, and hotels. Or perhaps you spend three weeks helping someone move abroad and start searching for visas, rentals, neighbourhoods, taxes, and flights.
To you, those actions belong to a particular episode. They have a beginning, a reason, and eventually an end.
Months later, you can still place them inside that story: that was when we were organising the wedding. That was when I helped my sister move. That was when Dad was changing jobs.
A behavioural record can contain every search from those periods without necessarily containing the story that groups them together.
That doesn't mean every system will treat a three-week obsession as a permanent preference. Different systems can weight recent, older, or repeated behaviour in different ways. The point is simpler: the sequence of actions is available in the data in a way that their biographical meaning may not be.
We organise our behaviour into episodes because we know the circumstances that connected those actions.
The data has to reconstruct that structure from the outside, if it needs it at all.
Prediction and Understanding Are Different Things
By this point, the distinction starts to become clearer.
A prediction asks something like: given what this person has done so far, what are they likely to do next?
Understanding asks something different: why did they do it, what did it mean to them, and what does it really tell us about who they are?
The studies we've looked at are impressive precisely because they show how much can be achieved without resolving that second question.
Kosinski and his colleagues found associations between Likes and particular attributes. Stachl and his team found patterns between everyday smartphone use and dimensions of personality. Matz and his colleagues showed that some psychological inferences could be used to select messages that produced measurable differences in clicks and purchases.
None of those results requires the model to recover the explanation the person themselves would give for their behaviour.
Curly Fries makes that especially clear. A signal can contribute to a prediction even when there is no intuitive causal story connecting it to the thing it predicts.
That doesn't make the prediction false. It simply means that predictive accuracy and understanding aren't the same achievement.
Why Getting It Right Feels Like Being Understood
From our side, though, that distinction is much harder to notice.
We don't experience statistical models or probabilities. We experience the moment Spotify gives us a song we love, a platform surfaces a video about something we've been thinking about for days, or an article appears that answers a question we hadn't quite managed to formulate yet.
When the match is good enough, it feels personal.
That makes sense. Over days or weeks, we may have been leaving small signals around a subject without ever thinking of them as a group. A search here, a video watched to the end, something we opened again two days later. To us, they were separate moments. To a system, they can form a pattern.
Then something appears that fits.
And we do something the model doesn't need to do: we recognise why it matters. We connect it to what's happening in our life, to something we want, fear, or are trying to work out.
Maybe that's why being predicted can feel so much like being known.
The system finds a pattern.
We know what that pattern means to us.
From the inside, those two things can feel remarkably similar.
The Version of You That Exists in Your Data
So what version of you actually exists in your search history?
It contains things that are true. The searches happened. The clicks happened. The songs you replayed, the apps you opened, and the videos you abandoned are all pieces of your real behaviour. Given enough of those traces, patterns can emerge that predict things you may never have explicitly said.
That's impressive in its own right.
But the data doesn't necessarily contain the story you would use to explain each of those actions. It can record weeks of searches about Japan without knowing you were organising someone else's trip. It can record a month of wedding searches without knowing where that episode begins and ends in your life. It can even find a useful correlation between two things for which neither of us would have an intuitive explanation.
The version of you that emerges isn't false. Nor does it need to be a complete psychological portrait to be useful.
It's built from the parts of you that can be observed.
Maybe that's why the interesting question was never simply how much an algorithm knows about us. It's how much it can infer about what we think when it can only observe what we do.
Your search history isn't a version of you that never lies.
It's a version of you that never gets the chance to explain what it meant.
References
- Kosinski, M., Stillwell, D., & Graepel, T. (2013). Private traits and attributes are predictable from digital records of human behavior. Proceedings of the National Academy of Sciences, 110(15), 5802–5805. https://doi.org/10.1073/pnas.1218772110
- Stachl, C., et al. (2020). Predicting personality from patterns of behavior collected with smartphones. Proceedings of the National Academy of Sciences, 117(30), 17680–17687. https://doi.org/10.1073/pnas.1920484117
- Matz, S. C., Kosinski, M., Nave, G., & Stillwell, D. J. (2017). Psychological targeting as an effective approach to digital mass persuasion. Proceedings of the National Academy of Sciences, 114(48), 12714–12719. https://doi.org/10.1073/pnas.1710966114
- Ajzen, I. (2020). The theory of planned behavior: Frequently asked questions. Human Behavior and Emerging Technologies, 2(4), 314–324. https://doi.org/10.1002/hbe2.195
- Warshaw, P. R., & Davis, F. D. (1985). Disentangling behavioral intention and behavioral expectation. Journal of Experimental Social Psychology, 21(3), 213–228. https://doi.org/10.1016/0022-1031(85)90017-4