Scaling laws are likely to keep improving model quality while reducing the cost of reaching a given level of quality.
2
Agents have strong potential, but reliable applications at scale are still difficult, so early deployments will focus on narrow domains and controlled tasks.
3
Companies will use portfolios of proprietary, open-source, and fine-tuned models, with data preparation becoming a major part of fine-tuning costs.
Summary
Euro Beinat describes the forces he expects to shape Generative AI over the next few years. He starts with scaling laws: more compute, data, and parameters have continued to improve model capabilities, while the cost of reaching a given quality level has fallen. He then discusses the uneven abilities of current models, which makes deployment and evaluation difficult. Agents are promising because language models can break work into tasks, use tools, retain memory, and iterate, but reliable production systems remain rare. Open-source models are improving and can outperform general models on narrow tasks after fine-tuning, while proprietary models still lead on many benchmarks. Euro expects applications to combine several model types. He also describes a progression from prompting through chaining, retrieval, plugins, and fine-tuning to training a model from scratch. The limiting factor in fine-tuning is often expert data preparation rather than compute. GPU supply, regulation, reliability, security, and ethics also affect how quickly these ideas reach production.
Scaling laws are likely to keep improving models while lowering quality costs
Euro says model capabilities have continued to increase with compute, data, and model size. He connects this to research associated with OpenAI and Anthropic, while acknowledging that the trend may eventually weaken. He also says the unit cost of achieving a given quality has tended to fall by about 50% every 16 months, with the exact period varying. Together, these trends mean models should keep becoming better and cheaper to use. That makes investment difficult because a tool that is useful today may become unnecessary when a stronger model arrives.
Model capabilities form a jagged frontier that is hard to predict
Euro compares human abilities with the uneven performance of language models. A model can be far better than an average person on some tasks while being much less capable on others. The boundary changes from model to model and is difficult to predict before deployment. He says teams compensate by combining multiple models or by using trial and error. This affects how organizations choose applications because benchmark scores do not fully reveal where a model will succeed or fail in real work.
Agents are promising because language models can plan, use tools, and iterate
Euro describes agents as systems that break a task into smaller steps, decide how to perform those steps, use external tools, access memory, and reflect on their results. He says language models make this approach more practical than it was in earlier research. Agent research has grown rapidly, and he sees demand across the companies in his group and across the industry. He expects early production systems to focus on narrow domains, such as a data analyst that writes code, runs it, and returns the result. Reliability, security, and ethical concerns remain unresolved.
Open-source models will coexist with proprietary models
Proprietary models such as GPT-4 and Anthropic models still perform better on most benchmarks, according to Euro. Open-source models are improving, and focused training with suitable data can make them exceed a general model on a narrow task. The tradeoff depends on the application. A use case that needs the highest possible performance may choose a proprietary model, while a task with many acceptable answers may favor cost, control, or the ability to run the model on premises. Euro expects model portfolios to combine proprietary, open-source, and fine-tuned systems.
Application-specific models can be built through a progression of interventions
Euro lays out a path from an existing model to one tailored for an application. Teams can begin with prompting, then chain models, add plugins, use augmented generation, fine-tune, align, or train a model from the ground up. Better general models may make prompting sufficient for more tasks. Other applications will still require deeper adaptation. Training from scratch brings major compute costs, while fine-tuning is likely to become more common.
Euro says teams often underestimate the data cost of fine-tuning. In his experience, compute can be a small part of the total cost. Teams have to create and curate datasets in several ways, especially when a model needs to be fine-tuned for multiple purposes. He argues that this work cannot simply be outsourced because it requires people who understand the data, the domain, and how models use the data. Better tools for preparing fine-tuning data are still needed.
Compute supply matters, but Euro has not seen GPU shortages stop his work
When asked about the GPU shortage and NVIDIA's position, Euro says the shortage affects what organizations can do, but it has not limited his group's work so far. He expects multiple suppliers to emerge because demand is high and says it is difficult to imagine NVIDIA remaining the only important supplier over the next two or three years. He also gives a brief positive answer about AMD's possible role.
Generative AI is being treated as a new computing paradigm
Euro explains that Prosus already depends on machine learning to operate its businesses at scale. He sees language-enabled use cases that were previously impossible and describes Generative AI as a new computing paradigm that could change technology companies. Prosus pays attention because it wants to benefit from that shift and stay ahead of it. This perspective explains why the company studies both immediate production use cases and longer-term drivers.