The shifting locus of AI portability

Agents are all the rage in AI these days. But what, exactly, is an agent? The concept implies, well, agency given by a human to a piece of software. An agent is empowered to make decisions and take actions based on what it encounters when it uses tools that connect it to the internet and remote services. It is typically personalized to the user and directed in specific ways. And for data portability, it’s not the model but the agent that matters.

The fundamental AI data portability concern that has motivated me for nearly three years now remains: in the AI future, will people be empowered or held hostage by their personal data? In that 2023 article, I said that models would become commoditized, and thus the ability of personal data to provide a personalized experience would be a key value proposition for AI services. The context I had in mind when writing that piece was providing personal data directly to a model; thus, portability meant being able to move personal data from one model provider to another.

Where does personal data sit today? With agents, the locus of where and how we build infrastructure to promote empowerment through personal data may have shifted. From an architectural perspective, an agent is a model plus a harness (see Vivek Trivedy or Databricks). A harness is everything except the core trained inference engine. That includes tools that execute code and run terminal commands and look up information in databases and online; memory beyond that in the current context window, including user specific preferences and an understanding of interaction history; workspace details including files, system environments, and other context; and guardrails, any policies or permissions that are meant to limit action. Harnesses both condition and load context before the model runs – context windows being the only form of input to LLM inference other than model weights – and review and modify its outputs before the user sees them.

Under this definition, perhaps we’ve been talking about agentic AI portability for the entirety of our time talking about AI portability! As our focus is on an individual’s personal data, which is gained through processes apart from inference and stored in some form of long-term memory (whether in the sense of a hard disk or in the sense of an AI memory), that data has always been something a bit separate from pure LLM operations. It has always been loaded into the LLM’s context window for inference by a harness structure of some form; and that combination, then, is an agent. When those functions were blurred into the service offered by the model provider, the terminology blurred similarly.

These terms can be used loosely in practice. For example, Claude Code is described in Anthropic’s documents as an agentic harness around Claude. Is it a harness or an agent? Is Claude part of the agent or something separate from it? Also, doesn’t Claude as a “pure” chatbot, as distinct from its coding tool instantiation, also include at least memory and guardrails and some level of workspace?

Hugging Face usefully broadens agents into three concepts: model, harness, and scaffold. In this configuration, the scaffold is a behavior-defining layer, and the harness is an execution layer. As I see those terms, scaffolding is more tightly connected to the model from a user’s perspective, whereas harnesses are easier to view as something operating separately.

Another possible layer for personal data in this context is “context,” as used by DTI affiliate Koodos Labs. A user’s personal context includes not just their memories and conversation histories, but other data that is specific to their use and running instance of AI. This could be something kept separately from the model, harness, and scaffold; in theory it could be stored within one or more of them, or separately, and paged in and out as the user needs.

So to return to my earlier question: Where does personal data live? Where should DTI be working with service providers to build portability tools? Like the terminology, the locus for personal data portability is unclear. Chat histories are one thing; we can get a handle on that in many respects. We know where it’s stored, and we can even build a clean schema for it. But the personal data used in customizing AI agent experiences, as distinct from chat services, can be much richer and at the same time much less clear.

There has been some progress in portability, though it is thus far inconsistent and very incomplete. You can import and export “memory” from Claude. You can usually download your chat histories from service providers, and some other personal data, and Google has rolled out guidance to import such data into Gemini. Meanwhile, a mention of the value of portable personal memory in an OpenAI Developer Community forum didn’t get any traction. As Kevin Bankston at CDT wrote in June in “Don’t Let Perfect be the Enemy of Portable,” there’s more that can be done, especially when it comes to imports.

Establishing clarity, transparency, and documentation into where personal data lives and how it can be accessed by other services matters, and not just for the kinds of user personal data portability that DTI works on. The frontier services aren’t just shipping models, and haven’t been for quite a long time. They’re shipping agents themselves, integrated models and harnesses with their own context and infrastructure. This integration can make data portability more challenging technically, trying to disentangle data and services; but it also makes it more tractable for regulators, who can impose rules on a single service provider and require them to figure it out. That seems more and more likely every day.



Previous Post

Catch up on the latest from DTI

  • AI
The shifting locus of AI portability
  • standards
IETF Work Related to Personal Data Portability
  • news
A midstride check-in on DTI
  • trust-registry,
  • trust
What–or whom–do you trust?
  • trust-registry,
  • trust
Launching the DTI Badge of Accreditation
  • policy
Our regular regulatory roundup
  • social,
  • standards
ActivityPub and account portability
  • research,
  • public-benefit,
  • open
Data portability and researcher access
  • trust-registry,
  • trust
DTI's Data Trust Registry is now post-pilot
  • policy
Web browsers - a data portability patchwork