I'm calling it now. My prediction is that, 5-10 years from now(ish), once training efficiency has plateaued, and we have a better idea of how to do more with less, curated datasets will be the next big thing.
Investors will throw money at startups claiming to make their own training data by consulting experts, finetuning as it is now will be obsolete, pre-ChatGPT internet scrapes will be worth their weight in gold. Once a block is hit on what we can do with data, the data itself is the next target.
Funny you should say that. There was a push to have more officially collected DIET data for exactly this reason. Unfortunately such efforts were recently terminated.