Personalization Needs to Be Trained
Frontier models have made remarkable gains in reasoning, instruction following, tool use, full computer-use pipelines, and deep domain knowledge in code, math, and other areas. Each of those gains came the same way: define the capability, build the training data that exercises it at scale, and train the model on it until the behavior holds.
What about personalization?
Today’s models are eager to comply, yet genuine personalization is rarely captured by benchmark leaderboards. It often requires post-training because it does not simply emerge from general competence, and neither a system prompt nor retrieved memory can supply it on its own.
An agent that adapts to the person it works for needs to learn from long-horizon, personalized tasks at scale. Why is that, and how was Innodata’s platform designed to produce the data needed to support it?
Generic Answers Are Practically Wrong
People expect an agent to remember what matters to them without being told their whole story every time.
A business owner who asks for a customer quote expects it at their prices, in their voice, and for a day they are available. If the agent gets those details wrong, the result is useless to the person who asked, even though the task was technically “done.”
Yet today’s evaluation benchmarks often miss this type of failure:
- The success criterion is whether the task was completed. But the same request from two different people can have two different right answers. A personalized success criterion should instead ask: “Would this person, with this history, consider the result correct?” Only the reality of a long-horizon exchange with a human user can supply that judgment.
- LLM judges can gravitate toward the average of the human interactions represented in their training data, rather than the months of context that determine whether a response fits a particular person.
- Outside evaluators who did not participate in the user-agent relationship may not know what the person wanted, which constraints were implicit, or why a plausible answer was still wrong. Years of a person’s messages contain emotion, negotiation, evolving preferences, and context that cannot always be reduced to names, dates, and explicit facts.
If the training signal cannot distinguish a personalized answer from a generic one, the model receives little signal pushing it toward personalization. It learns to satisfy the instruction for the median user—the user the available data most clearly describes.
A memory file at inference time can give the model relevant information in its context window, but it may not teach the model what to weigh, when an old constraint should override a new instruction, or how different pieces of personal context should influence an action. Those behaviors need to be represented in the training data itself.
Data Curation for Personalization
To capture the uniqueness of each user, personalization data should account for three things:
- Long “relational” horizon: A long-horizon task usually means thousands of interaction steps—for example, a computer-use agent working through a sandbox for an hour. In our setup, long horizon means relationship. To serve today’s request, the agent has to recall a conversation from months ago, weigh it against how the person’s thinking has changed since, prioritize among competing pieces of context, and translate the result into action. Personal agents fail on this far more often because training data rarely exercises it. Our tasks require both: multi-step work whose correct execution depends on months of accumulated context with one person.
- Implicit constraints: Personal requests are complex because they often contain requirements that are never stated directly. A person whose profile says they never schedule before 10 a.m., travel with a toddler, and keep Friday afternoons free has specific constraints when asking, “Book me an appointment before my flight.” Task completion is necessary, but personalization is what earns the reward and provides the relevant training signal.
- No synthetic neatness: A naive way to create personalized training data at scale is to generate it—invent personas, generate tasks, and fine-tune. But synthetic pipelines can struggle to reproduce the inconsistency, contradictions, changing priorities, and individual quirks that make real people difficult to model. Human annotators are important here because they introduce the natural complexity that synthetic personas often smooth away.
A Bit of Information Theory
The difference between generic and personalized data lies in the information distribution they encode.
Formally, let a persona be defined as a random variable P. It generates rollouts, modeled as a random variable X that typically depends on P. The total variety in the data decomposes into two entropic terms:
H(X) = H(X|P) + I(X ; P)
where:
H(X | P) is the within-persona entropy, measuring task variability within the same persona environment.
I(X ; P) is the cross-persona entropy, or the mutual information between persona and task. It measures how much a task rollout informs about the person who generated it.
Synthetic pipelines can struggle to maximize both terms at the same time, whereas real human data naturally exhibits variability across both dimensions. Caricatured personas may inflate the second term—high mutual information between persona and rollout—while reducing the first through low entropy within the persona, making them predictable.
This gives a cheap and honest test: try to guess which persona produced a held-out rollout. Its accuracy is a lower bound on I(X; P): Fano’s inequality says it cannot exceed mutual information. If the probe is near chance across personas, the pipeline is not producing enough diversity, whatever the rollouts look like.
Run the same probe on a real population, and the gap tells you how far the synthetic one has drifted toward the “average user.”
Mutual information has a known limit: it cannot tell a well-drawn character from a stereotype, since both are highly attributable. Nor can any intrinsic metric say whether a persona stays consistent while drifting across fifty sessions, or whether a synthetic population covers the tail rather than the bulk.
The most reliable anchor for each of those is a real person saying, “That is not me.” This is why our platform keeps the human at the center rather than treating interviews as merely a bootstrapping step for a generator.
The Agent Personalization Platform
The Agent Personalization Platform is a training and evaluation environment for personal agents, where all three ingredients above land in the same data packet.
- The agent starts its first task with a constructed history grounded in the participant’s life and work, as if it had already worked with this person for half a year. The model must learn to prioritize within a memory that is deliberately messy: conflicting, disorganized, with competing priorities and rules that have exceptions.
- We create tasks that depend on participants’ goals, histories, and context, and ask the same people to evaluate the results. A participant is interviewed in depth about their life, work, tools, habits, and how they communicate. That information becomes the agent’s memory and seeds the workspace. Task authors then write requests whose correct answer requires implicit constraints. Finally, the participant reviews the agent’s actions and explains why a result fits or fails. The person who built the context judges the output—the evaluator who can distinguish the personalized answer from the generic one.
- The platform runs agents in virtual workspaces with real tools, the user’s data, and a memory store. Under the hood is a fleet of virtual machines, each running a real agent harness plugged into one person’s story: inbox, calendar, messaging, files, CRM, invoicing, seeded to be consistent with that person. The workspace is stateful and mutable, so actions carry consequences in the long run.
- Checkpoint technology lets a team freeze memory, tool state, and even search context at a moment in time, change conditions, and rerun the task. Because the environment is controlled and the rubrics are explicit, the same task can be rolled out many times and scored consistently. Rubrics encode the participant’s rules as verifiable checkpoints; their feedback becomes preference data. Both hold across reruns.
Each task is one self-contained, replayable packet: the persona the agent runs as—identity, user it serves, values, voice, constructed memory, apps it operates in, the tools it can call, and the seeded data behind each—and the task it is asked to complete: an under-specified request, tool specs and backing data consistent with the persona, a rubric of verifiable checkpoints, and the participant’s verdict.
That structure turns subjective preference into something closer to a verifiable reward.
A single rollout with a thumbs-up is preference-based, while many rollouts of the same task against a rubric the participant wrote is RLVR applied to human preference. The data packet serves both use cases: preference pairs for RLHF and rubric-scored rollouts for RLVR-style training.
Concretely, every packet yields both:
- Evaluation data: trajectory-level scores on tasks that require months of context beyond completion; rubric checkpoints plus the participant’s verdict and reasons; frozen checkpoints so the same moment is comparable across model versions; and memory ablations that show which recalls break when storage changes.
- Training data: preference pairs from the people who hold the preferences; rules turned into verifiable rewards across repeated rollouts; a profile of how each person quotes, books, declines, and writes; and environments that grow richer with every session, sandboxed, portable, and re-runnable.
Who is This For?
Any team shipping an assistant that is supposed to know its user.
A shopping agent recommending a laptop to a software developer should account for different priorities than one helping a college student working within a budget.
A life-sciences research assistant should learn how a particular scientist reads a paper.
A small-business agent should quote jobs in its owner’s voice.
A writing assistant should respond the way its author would.
The recipe is the same in each case:
- Bring one user group, or let us recruit a sample of it.
- Bring the task your agent gets technically right but practically wrong.
- We build the workspace and the checkpoints around it, plug your agent in, and find out which tasks left those people unhappy, what the agent should have called, and which memory it should have consulted.
- Train on the annotated answers.
Personalization is the unlock for agent-as-assistant.
It is not just a prompt, and it is not just a memory file. It is a training problem—and to solve it, you need to make personalization personal.