AI systems become useful in organizations when their outputs are placed inside existing workflows and applications.
2
Fine-tuning needs continuing feedback from trusted subject matter experts, including corrections, context, and rankings of possible answers.
3
AI Squared uses LLM Link, application context, other models, and templates to create personalized experiences outside chat applications.
Summary
Benjamin Harvey argues that organizations need to move generative AI beyond standalone chat applications. Users often need an answer inside the tool they already use, with the content and context required to make that answer relevant and actionable. AI Squared combines information from organizational repositories, predictive models, other generative models, and the user's current workflow. Harvey also stresses that a vector database and an open-source model are not enough. Teams need a continuing feedback process in which subject matter experts correct answers, add context, rank alternatives, and identify missing sources. AI Squared calls its connection and discovery technology LLM Link. The platform can use templates to generate dashboards or other interface components inside existing applications, while collecting feedback for later tuning. Harvey is candid that this requires sustained organizational involvement, focused initial use cases, and product owners who care about how business users actually use the technology.
AI adoption requires integration into the tools people already use
Harvey says his work at the National Security Agency exposed a major gap: organizations could deploy AI and machine learning models, but struggled to put them inside existing tools and applications. AI Squared was founded around this problem. Its goal is to place generative and predictive AI directly into workflows so analysts and business users can make decisions without leaving their applications for a separate chatbot.
A vector database alone does not produce a dependable organizational model
Harvey distinguishes pre-training from fine-tuning. A model can begin with a large vectorized source of information, many parameters, and many tokens, but that is only the first stage. Organizations then need examples, rewards, rankings, and corrections that reflect their own domain. He says teams often underestimate the work required after bringing an open-source model into the organization.
Human feedback needs corrections and context, not only a thumbs-up signal
In AI Squared's feedback process, a subject matter expert can mark an answer as good or bad. When it is wrong, the expert supplies a better answer and can explain why the original response failed. The system can use those question-answer pairs and the added context to tune the model. Multiple candidate answers can also be ranked for reinforcement learning.
Fine-tuning has to continue throughout the model's deployment
Harvey describes feedback as a lifecycle activity rather than a one-time training step. Teams should track accuracy, performance, timeliness, relevance, actionability, and context. They should also ask whether another source could have produced a better answer. AI Squared uses observational studies and repeated feedback to monitor how the model improves over time.
An LLM can combine many organizational sources and other models
In a cybersecurity example, the system can draw on IP information, hash information, domain information, help documents, vulnerability intelligence, threat intelligence, and other language or predictive models. The LLM decides which sources to reach for when it needs more information. The aim is to connect the model to useful documented and undocumented sources rather than relying only on one repository.
Application context changes the answer an LLM should provide
Harvey says useful integration needs both content and context. Content includes the sources available to the system. Context includes the page a user is viewing, the application they are in, the lead they are examining, and where they are in a workflow. AI Squared can also collect implicit behavior, such as the page a user is on and how long they have been there, along with explicit feedback such as likes or shares.
LLM Link connects sources while templates turn model output into interfaces
AI Squared's LLM Link technology discovers documented and undocumented data sources and connects the LLM to other models and data. Generative AI then creates code from templates for experiences inside existing applications. Harvey gives examples such as circle charts, bar charts, and pie charts. The platform can also collect feedback on those generated experiences and use it for later tuning.
A financial-services workflow can combine predictions, recommendations, and feedback
Harvey describes a sales use case in which account managers work with potential clients. AI Squared combines a propensity score, product recommendation models, web and event signals, and application context. It generates visualizations inside the sales application so the user can assess a lead without switching tools. The user can then give feedback on the result and request a different presentation or answer.
Production use needs focused scope and people who own the business workflow
Harvey says organizations should not try to answer every possible question in the first model. They can start with a focused capability, then add other sources, predictive models, and LLMs. He also says production adoption needs a connection between data science teams and product owners who are responsible for embedding technology into business workflows. Spending money on models does not by itself create something users will adopt.
"The fine-tuning is what we're seeing as the most important piece of the puzzle over time to provide these organizations with the ability to get a high fidelity of quality in the answers that they're getting from these large language models."13:21
Who should watch
You are integrating an LLM into a business application and need to decide how much workflow context to provide.
Your team has a retrieval prototype but lacks a process for collecting corrections and using them to improve the model.
You sell AI to regulated or government organizations and want Harvey's view on product ownership, observational studies, and adoption.