Everyone in San Francisco is predicting an “era of abundance.” This is wrong.
We will solve hard problems in medicine, economic development, and all the arenas we battle in today. That doesn’t mean we will escape scarcity. Scarcity is not a technical problem, it is a fact of the world we live in. We will always have to grapple with having limited resources to satisfy our infinite wants. The game is changing though. We are entering a post-labor regime.
Today, labor is the binding constraint on economic growth. The only reason your dollar is worth anything is because you can use it to tell someone to do something for you, because they know they could use it to tell someone to do something for them. That is changing. In a world where AI systems deliver the bulk of economically valuable goods and services, labor alone ceases to be the binding constraint of economic growth.
Consider the function of labor. The reason I pay my cab driver or doctor is because they can fill in the gaps to figure out how to solve a problem I care about. The difference in what I pay my doctor and Uber driver is a function of how much less accessible the knowledge required to fill in the gaps is. Whenever I pay for something, I’m fundamentally paying for people to fill in the gaps, resolve the ambiguities, to figure out what needs to be done so that what I want can happen. In a world where “agents do all the work,” then the question is how do these gaps get filled.
One bet here is that these systems really do it all. Humans become somewhat passive observers of machines that fulfill their wants and needs. Compute is the binding constraint. This point of view seems eerily similar to bets that have been tried before and didn’t work out so well.
The truth is that omniscience is a myth. Everyone is uniquely positioned somewhere and can access information no one else can because of where they are positioned. That’s why we work together to solve problems. You can take the smartest person in the world and they’ll still ask you where the cups are in your kitchen. The knowledge we need to solve problems is situational and labor is an arduous way of getting it out of the people who have it to get things done.
We have not even scratched the surface of the data we need to train AI systems to give adequate background knowledge, yet alone have we established the pipelines to allow for the real-time contextual knowledge these systems will need to solve real problems. Frontier models today are trained on a small sliver of the knowledge we’ve accumulated.
The public, accessible web is a small subset of the open internet. The data from human expert pipelines isn’t even a rounding error of the knowledge people have. For the past decade, the data required to produce step improvements in model capability was easily tractable. The web was relatively easy to scrape. Experts were relatively easy to source and incentivize. Data quality was relatively easy to adjudicate. That’s not true anymore.
Data will need to come from everywhere. This is more than what’s been publicly articulated or what could be articulated if you asked someone. It is knowledge that is fundamentally hard to source. That is often the case because revealing the knowledge compromises someone’s competitive advantage.
This is how the labor market works as well. If you hire someone to do something, then that person has to spend their time doing that for you and can’t spend it doing something for someone else or themself. The reason the labor market can function, however, is because of pricing. You can tell me how much you’ll pay me for my time, and if it’s better than my alternatives (the amount others are offering me or the amount I’d be willing to forego for the freedom) we have a deal.
Human data has been able to be priced this way as well because the data market has functionally operated through labor to capital arbitrage. A model company wants to train their model and can contract experts to produce training data. The experts get paid the market rate for the labor and the company gets the product of their labor to train their model.
We are entering a regime where this exchange shifts away from labor. The data needed to generate further improvements in model capability will need to come from the wild not just through annotation platforms. Every company will have a data business and they will have to decide, just as we do with our labor, how much their institutional knowledge is worth to them across all the ways it can be spliced. A private practice has everything from phone calls, to imaging records, to HL7 code, to appointment logs, to insurance claims, to internal documents, and more. Each of these can be anonymized and made useful by someone else.
In order for the private practice to make sense of whether they would prefer to keep their data internal, versus sell some splice of it to another company, whether it be a foundation model company, a competitor training a specialized AI system, or another company in a different market, they will need prices. They will need to be told how much they will be paid for their data, and if it’s better than their alternatives (the amount others are offering them or the amount they’d be willing to forego for the freedom) then they’ll have a deal. The rate of growth will be bottlenecked by the rate at which these decisions are made.
These prices require reconciling hard questions. Where is the data? What are the potential downstream use cases of the data? Can this data be found elsewhere or is there other data that can enable the same downstream use cases? How likely is the data to enable these downstream use cases? The technologies we build to resolve these questions will differ across modalities and domains, but they will be needed to settle prices so that the market can function and, ultimately, so that the gaps can be filled in so that what we want can happen.