AI Without Context Sucks

Ask an AI to recommend a restaurant. You will get the same five places everyone else gets, because the AI knows almost everything about the internet and almost nothing about you.
The models themselves are extraordinary. The problem is what they have to work with. An AI can only be as useful as the information it can see, and the information that would make it useful to you personally, what you like, where you have been, what you are actually trying to do this week, is exactly the information it has never seen. AI without context sucks. Real context from real humans is what makes AI work.
The industry knows this. Over the past year, the most sought-after skill in AI has shifted from writing clever prompts to something called context engineering: the craft of getting the right information in front of a model at the right moment. Anthropic's engineering team describes context as a critical and finite resource. Put simply, the people building AI have concluded that the model is only half the product. What it knows about you is the other half.
Which raises the obvious question: where is that knowledge supposed to come from?
How AI learned everything, and why that stopped working
Nearly everything an AI knows, it learned by reading the public internet. Companies send out automated programs called crawlers that visit billions of pages and copy what they find, and that copied text becomes the raw material models learn from. For years this was treated as a free and effectively infinite resource.
It is now neither. Researchers at MIT's Data Provenance Initiative spent a year measuring how much of the web remains open to AI crawlers and published the result under a blunt title: Consent in Crisis. In twelve months, 5 percent of the material in the most widely used training collection became fully off limits. Among the most actively maintained websites, the sites with the freshest and best-kept content, 28 percent shut crawlers out entirely. Count the sites whose terms of service forbid AI use and 45 percent of that collection is restricted.
The plumbing of the internet moved the same direction. In July 2025, Cloudflare, a company whose infrastructure sits in front of roughly one in five websites, began blocking AI crawlers by default and introduced a system for charging them per visit. Website owners now decide whether AI companies get their content at all, and increasingly the answer is no, or not for free.
But there was always a deeper problem with scraping, and it would exist even if the web had stayed wide open. The context that makes AI useful to a person was never on the public web in the first place. Your listening history sits inside a streaming app. Your purchases sit inside a shopping account. The gap between what you planned to cook this week and what you actually ordered is not published anywhere a crawler could reach. That information exists, but it belongs to someone. There is no crawler for it. There is an owner, and the owner decides.
Real context starts with real consent
If the best context can't be scraped and has to be contributed, then getting it is no longer a technical race. It is a trust question. People will bring their real data to a system when they get something real back and when they stay in control the whole time.
Most of the industry handles that control as paperwork: a checkbox, a settings page, a privacy policy nobody reads. Vana, as open data infrastructure for human-grounded AI, handles it as a working mechanism. Through Vana's data portability API, you grant an app access to a specific slice of your data. The grant is recorded, and from then on the network serves that data to that app and nothing else. Change your mind and revoke the grant, and the serving stops. Consent is not a promise in a document. It is how the system physically behaves.
For the people building apps, this flips the usual starting position. Instead of guessing at who a user is from scattered clues, the app is handed the real thing, directly, because it gave the person a reason to share it.
If we want AI that works for real people, it has to be built on real people. Not statistical sketches assembled from the public web, but actual humans, contributing their actual context, on purpose.
Forty apps are testing this right now
The Vana Cup is this idea run as a live competition. Builders ship apps on the network, and the apps score points: a goal each time someone's data is connected through the app, an assist each time that data proves useful to another app on the network. Assists count double, deliberately. Data moving between apps, with the owner's consent recorded at every step, is the whole point of open infrastructure. A scoreboard that rewarded hoarding would contradict the system underneath it.
Midway through, with one week to the final whistle, 40 apps are on the leaderboard. Every point on that table is the mechanic working: a grant made, data served, and in the case of assists, context proving useful beyond the app that first earned it. If consent-based data were something nobody would build on, the table would be quiet. It is not quiet.
What this means
The scraped web is closing, and the context that matters most was never there anyway. The builders who learn to earn context, rather than take it, will make the products everyone else wonders how to compete with.
The Cup closes 18 August at 23:59 UTC. Ship an app. Give people a reason to bring their data. See what it does to your product. Enter at builders.vana.org.