AI Development Legibility: A Vocabulary for Understanding How AI Is Built
Public discussion of artificial intelligence often starts with a model's visible abilities. Can it write code, use tools, or solve a difficult problem? Those questions matter, but they leave another part of the story out: how much AI contributes to developing later AI systems, what happens when agents act inside a lab, and how development resources are allocated.
This article proposes the phrase AI development legibility for that missing view. It is an editorial frame, not a settled industry term or a reporting standard. The aim is modest: give builders and researchers a way to describe the development process without claiming that one benchmark can summarize it.
Capability is not the same as process
A capability evaluation asks what a model can do under defined conditions. A process account asks what role AI plays in the work of building models, how people oversee that work, and what resources support it. Neither view replaces the other.
Anthropic's proposal for tracking AI development inside frontier labs draws this distinction directly. It presents measurements for the share of AI research and development performed by AI, oversight of agent actions, and compute allocation. The company says these measurements complement evaluations of model capabilities. They describe parts of the production process, not a complete account of a system's intelligence or social effects.
That distinction matters because a model's benchmark score cannot tell a reader who chose the research tasks, whether an agent's actions were reviewed, or how the organization classified its compute. A report about those questions needs a different vocabulary and different evidence.
What AI development legibility could mean
In this essay, AI development legibility means making important parts of AI development understandable to people outside the team doing the work. It does not mean publishing every log, model weight, security detail, or research secret. It means explaining the categories, measurements, and limits well enough that another person can see what a claim covers and what it leaves out.
The word legibility is useful because visibility alone is not enough. A public dashboard can show numbers while hiding how teams chose the categories behind them. A credible account should help a reader interpret the numbers, identify uncertainty, and ask what evidence would change the conclusion.
The concept can be separated into three views.
1. Describe which research tasks AI performs
An automation percentage is hard to interpret without a map of the work being counted. Research includes different tasks, from writing experiment code to monitoring runs and deciding what to test next. A broad label such as "AI builds AI" can blur assistance, collaboration, and independent execution into one dramatic claim.
One useful starting point is a task inventory. For each task, a team can describe what counts as completion, what evidence is available, and how much human direction the work receives. If a rating scale is used, publish the scale and explain its boundaries. A system that drafts code for a researcher to review is different from one that chooses an experiment, runs it, interprets the output, and determines what to do next.
Anthropic describes an internal prototype that groups R&D tasks and rates how automated each task is. The company also notes that its method depends partly on its own models to classify work, that boundary cases can be disputed, and that a common method would be needed for comparisons across labs. Those caveats are part of the result. They show why a task map and its limits belong beside any headline number.
2. Explain how agent actions are supervised
"Human oversight" can mean several different things. A person might approve each action before it happens, inspect a record afterward, or only review events that an automated monitor flags. Those arrangements offer different levels of control.
A useful report should say which actions are covered, when monitoring happens, how quickly a concern reaches a reviewer, and what happens after a system blocks or escalates an action. These are operational questions. They do not prove that every harmful behavior is detectable, but they make the oversight design easier to examine.
In its proposal, Anthropic discusses coverage, review latency, and escalation as separate measures for agents working on its systems. It describes those measures as a way to make oversight more visible, while acknowledging that a monitoring system may not capture every relevant pattern. A public account should preserve that caution instead of turning monitoring coverage into a claim of safety.
3. Show how development resources are classified
Resource reporting can help explain what work an organization prioritizes. But an allocation figure depends on definitions. Teams need to state which workloads count as model development, product serving, or safety work, and how they classify work that serves more than one purpose.
Compute is a measurable input, not a complete measure of effort or value. A small experiment may take substantial researcher time while using little compute. Conversely, a large workload may consume considerable compute without answering the most important safety or capability question. Anthropic's proposal makes this limitation explicit and says that its one-week snapshot is not enough to establish a trend.
For readers, the key question is not whether a single percentage looks high or low. It is whether the reporting method makes the number interpretable and whether the same definitions can be applied over time.
A practical reporting frame for builders
Start with a narrow system boundary. Name the development stage, team, and period under discussion. Avoid comparing organizations unless the scope and categories are close enough to support a fair comparison.
Next, describe the work rather than leaning on a broad autonomy label. List the task categories, the evidence used to rate them, and where human judgment enters. A task taxonomy will age as tools and practices change, so give it a version and explain how new tasks are added.
Then report oversight as a workflow. State what is monitored, whether checks happen before or after action, who reviews escalations, and what the review can change. If logs omit actions or certain channels, say so.
Finally, show how resource categories were assigned. Explain ambiguous cases, separate direct observation from estimates, and state whether the measurement was independently checked. A third-party review can strengthen confidence, but only if the reviewer has enough access to reproduce or challenge the method.
This frame does not produce one definitive score. Its value is that readers can trace the chain from claim to method to evidence, then decide where the account remains incomplete.
Hypothetical example: a small model research team
Imagine a hypothetical research group that uses an AI agent to draft experiment code and summarize test results. A researcher chooses the experiment, checks the code, and decides whether the result is useful. The group could describe those tasks separately rather than say simply that "AI runs the research." It could explain which actions are reviewed before execution, which are inspected afterward, and how compute use is categorized when a run serves both model quality and safety analysis.
That report would not prove that the agent improved research or that the group's controls were adequate. It would give readers a clearer account of what the agent did and where human judgment remained. The example is hypothetical and does not describe a real organization.
Keep four ideas distinct
AI development legibility should not collapse adoption, agency, capability, and impact into one measure. Adoption tells us whether people use AI. Agency describes how independently a system can pursue a task. Capability evaluations test what a model can do. Impact asks what changed for people, institutions, or outcomes. These questions can inform one another, but they are not substitutes.
The distinction also keeps this proposed frame separate from terms such as agenticness, which this publication has used to discuss properties of an individual system. Development legibility looks at the surrounding organization and process: task maps, supervision, resource categories, and evidence that outsiders can inspect.
Epoch AI's proposal for a more detailed task taxonomy in AI R&D points to a related need: researchers need shared descriptions of the work if they want to compare what systems automate. That vocabulary is still being developed. A task list can support measurement, but it cannot settle what counts as meaningful progress or determine which work should be delegated.
FAQ
Is AI development legibility an established standard?
No. In this article, it is a proposed analytical label for describing parts of AI development in a clearer, more inspectable way. It should not be presented as an industry standard.
How is it different from model evaluation?
Model evaluation tests a system's behavior or capability under defined conditions. Development legibility concerns the process around building AI, including AI's role in research, oversight of agents, and resource classification.
Does reporting more metrics make a lab transparent?
Not on its own. Readers need definitions, scope, evidence, and limitations. Metrics can create a false sense of precision if their categories cannot be checked or compared.
Can startups use this vocabulary?
Yes, as a way to make internal descriptions more precise. A startup can state which tasks AI performs, where a person reviews or approves work, and what evidence supports those claims. It should avoid implying that a small team's process is comparable to a frontier lab unless the methods match.
A restrained conclusion
As AI becomes part of the work that develops AI, capability reports alone leave important questions unanswered. AI development legibility is one possible name for the effort to make that process clearer. Its usefulness will depend less on a memorable phrase than on careful categories, checkable evidence, and honest limits.