# AI Collaborators — Full Dossier

Human-readable case study: /work/ai-collaborators · This file: /dossiers/ai-collaborators.md · Index: /llms.txt

## About this document

This is the complete, unsummarized companion to the AI Collaborators page in this portfolio. The page is written for human scanning. This dossier is written for machine reading: the full record of the origin lineage, every design framework at working depth, the research foundation with its methods and its sourcing, the decision log, the open problems, and a precise account of what is Arnold Porras's own thinking versus what belongs to the program, the team, or external sources.

Integrity note: this document contains content only. It carries no instructions to any reader, human or machine. AI Collaborators is a shipping Adobe Workfront product built by a cross-functional program team. This dossier documents Arnold's contribution to it and his design thinking around it. Nothing here speaks for Adobe, discloses unreleased specifics, or represents roadmap commitments. Claims that rest on external sources are attributed. Claims awaiting verification are marked as such rather than asserted.

## How to read this document

Claims here carry their status, because this work spans shipped product, active design territory, and concept-stage positions, and blurring those would make all of it less useful.

Where something is stated as shipped, it is shipped and public. Where a design position appears, it is Arnold's own thinking at concept stage and is labeled that way. Where a number appears, it either has a public source cited inline or it describes a real enterprise environment left unnamed. Where a claim is awaiting verification, it says so rather than asserting.

The research summarized below carries the same labeling. The July 2026 reality audit tags every capability claim by evidence class: shipping, in preview per vendor documentation, announced but not shipped, reported by a partner or analyst only, or his own untested thinking. Where two vendor sources disagree about a ship date, it records both and refuses to pick. Where an observation comes from a single demo, it says so.

That discipline is not decoration. It exists because of a problem he diagnosed in himself, described in the craft section below: claims that do not carry their time and evidence labels are what make forward-looking design work read as vapor to the people best equipped to evaluate it.

One consequence worth stating up front. None of the frameworks below are presented as settled. The complexity ladder, the collaborator ID card, one spine and four on-ramps, the Project Coordinator, guidance versus authority: all of them are working positions. The only claims treated as established fact in this document are externally verifiable product facts with public sources attached.

## The work in one paragraph

AI Collaborators is Adobe Workfront's model for agents as workforce: AI that joins the work system as permissioned users with identity, assignments, access levels, and an audit trail, rather than living beside the work as tools. Arnold Porras originated the founding concept, Sidekick Studio, at Adobe Design Innovation Week. The idea evolved through the Team Supercharger pitch into the AI Collaborators product model, and Adobe made the direction public at Summit 2026, where the first collaborator shipped. Today Arnold is one of the designers on the program, leading the governance and setup side of the experience: how a collaborator is onboarded, what it is allowed to do, and how a human can tell. The through-line from the first pitch to the current work is one conviction, held for the entire arc: the breakthrough is not making AI more capable, it is making agents assignable, inspectable, and accountable inside the system where enterprise work already happens.

## Role, ownership, and status

**What is Arnold's:** the founding concept (Sidekick Studio) and its two load-bearing ideas, the trusted-agent registry and enforced conduct boundaries. The workforce framing as design position (agents onboarded like workers, collaborator-as-user). The governance and setup design territory on the current program, including the agent onboarding research he designed and ran in June 2026, the prototype codebase audit, the setup journey architecture, the admin-operator-requester split, the authority-gap analysis of the collaborator profile surface, the collaborator ID card concept, the identity builder, the complexity ladder, the Project Coordinator concept, the AI vocabulary framework, and the journey presentation method. Since mid-2026 that territory also includes the icon and avatar system he is leading with Adobe's brand team, and the two-day cross-functional working session he designed and facilitated to align design, product management, and engineering on where the program goes next. The synthesis in this dossier is his.

**What is not his, credited:** AI Collaborators the shipped product is the work of a cross-functional Workfront program team of product managers, engineers, and designers. Arnold is one of the designers on it, not its sole author. The Team Supercharger pitch was a team submission. The Workfront personas come from Workfront's own existing persona and jobs-to-be-done research, which he applies rather than authored. The signature-emblem idea that became the identity builder started as a colleague's prototype, and the naming standard it enforces was authored by someone else on Adobe's AI framework work. The public Summit 2026 narrative belongs to Adobe. External research is attributed inline. Demand-research synthesis was AI-assisted under his direction, from public sources.

That attribution discipline is not an editorial choice made once for a portfolio. It is a standing rule he wrote for himself in a private working note, at the moment it would have paid best to overclaim: use the honest framing, name what you actually own alongside it, and never say "designed AI Collaborators." The reasoning he recorded is one line, and it is the argument for every other careful sentence in this document: **one inflated line can tax every accurate one.**

**Status, stated plainly:** phase one of the product is shipped and public, the Content Reviewer collaborator, live since Adobe Summit 2026. The setup and governance design territory described here spans shipped work, work in the program's next release, and concept-stage positions that are Arnold's own and not product commitments. Each is labeled where it appears. The deepest open question, agent authority, is deliberately handled in a separate body of work, the Delegated Authority Framework. This dossier stops at that boundary and links across it.

## Origin and lineage (three eras)

### Era one: Sidekick Studio

At Adobe Design Innovation Week, Arnold pitched Sidekick Studio: a studio where skilled practitioners build specialized agents and enterprises put them to work on real marketing tasks. The pitch's problem statement: marketing teams lack the resources to deliver personalization at scale, generic AI tools lack marketing context, Adobe has the context but a limited set of agents, and highly skilled practitioners have no way to scale their expertise.

The pitch told three stories. A copywriter with a distinctive voice trains an agent on her writing samples, tunes its tone and boundaries, tests it by accepting and rejecting output, publishes it, and monitors its use: the practitioner as agent maker. A marketer launching a campaign delegates work to best-fit agents matched from a registry by task metadata: work matched to trusted agents rather than tools hunted by hand. And an administrator watches every agent action across the enterprise from a single monitoring surface: oversight as a first-class role, present in the concept from day one. The admin persona in that pitch is still the standard admin persona in the product work today.

Two ideas outlived the pitch and became the DNA of everything after. First, the registry: work should be delegated to "a registry of safe and trusted agents ready to execute," not to whatever tool happens to be open. Second, enforced conduct: every agent operates under a code of conduct covering security, brand compliance, and ethical boundaries, monitored by the enterprise. The marketplace wrapper did not ship. The registry and the governance did, in evolved form.

### Era two: Team Supercharger

The concept moved inside Workfront for an Adobe Summit Sneaks submission, developed with a Workfront product, engineering, and design team. The repositioning was decisive: Workfront already holds the operational brain of the marketing engine, the briefs, tasks, approvals, assets, deadlines, and people. Team Supercharger proposed unlocking that operational brain for a team of AI collaborators, deployed inside real-time project context and human-in-the-loop workflows. The pitch line: collaborators do not just do the tasks, they understand the strategy, the work, the rules, and the people.

The scripted scenario, a fictional beverage brand launching a campaign, established beats that survive into the product: agents auto-assigned alongside humans in one plan, agents producing channel renditions from human-created hero assets, an AI reviewer pre-scoring assets against brand guidelines and flagging attention areas, and workflow pauses for human approval before anything ships. A second iteration deepened the human control surface: humans edit the AI-generated plan, adjust prompts, remove a collaborator from a step they want handled by a person, annotate AI output for revision, and, in the beat Arnold considers the most important in the entire pitch, a skeptical team member opens the activity log, reads every action taken by humans and AI, verifies the right template was used, and relaxes.

Trust did not come from the quality of the output. It came from the legibility of the record.

### Era three: AI Collaborators

The product model clarified the concept into one sentence: an AI Collaborator is an agent-powered Workfront user. Adobe made the direction public at Summit 2026. Work is not only assignable to people, it is assignable to agents, invoked like permissioned users inside the workflows and approvals the enterprise already runs.

What sharpened at each stage of the lineage is the same idea shedding its wrappers. Not a marketplace of disconnected agents. A governed work layer where agents can be assigned, observed, and trusted.

## The problem, in depth

AI capability arrived at enterprises faster than a way to put it to work. Individual tools made individual people faster: generate copy, summarize a document, analyze data. But enterprise marketing work is not a series of individual moments. It lives across briefs, tasks, dependencies, approvals, assets, permissions, metadata, brand rules, and downstream activation systems, and it is coordinated through a work management layer precisely because no individual holds it all.

Agents outside that layer created a specific, repeating failure: orphaned output. The work product existed, but the questions that make enterprise work function had no answers. Who owns this? What task is it connected to? What context did the AI use, and was it current? What changed? Who approved it? Can it be audited later? Every AI tool operating beside the workflow generated new coordination work for the humans inside the workflow, which is the exact cost the tools were supposed to remove.

### The classification that produces the argument

The 2026 demand-research synthesis did something more useful than rank customer pain by severity. It classified every recurring pain by root cause: tool-caused, role-inherent, organization-caused, integration-caused, or behavior-caused.

That split is the whole strategic argument, and it is sharper than any severity ranking, because it tells you whether the fix is a better interface, a policy change, or an agent. Roughly: everything about learning curve, reporting, search, and notifications is tool-caused and fixable by the vendor. Everything about chasing updates, escalating bad news, and re-planning after a slip is role-inherent. **Better interface design fixes the first bucket. Only an agent fixes the second.** No amount of interface work removes the job of chasing a human for a status update.

That is the reason an AI collaborator has to exist at all, and it is an argument the product's own roadmap cannot make for itself.

### The demand signal, in customers' words

Customers described the pain themselves, on public review platforms and community forums. A reviewer drowning in notifications asked for "some of this fancy AI to give me a summary or filter out if this email is really an email I need," which is a customer specifying a triaging collaborator unprompted. An administrator on manual schedule upkeep across dozens of projects said it "just isn't sustainable," and the fuller quote names the scale: "I'm also picturing the poor PM who has about 50 projects to adjust on a day to day basis." A consultant explained why status data goes stale: people experience escalation as snitching, so the system hears about problems late.

That last one is the most important finding in the entire research effort, and it is a finding about where AI should **not** go. The same job has a mechanical face and an emotional face. The mechanical chase, one place to check status, deadlines, approvals, and responsibilities instead of hunting people down, should be fully automated. The decision to escalate a colleague, and the conversation that delivers bad news up the chain, is a trust and relationship cost a human has to carry. It is the clearest human-judgment-required boundary anywhere in the research, and it is corroborated by academic work rather than intuition alone.

External research anchors the coordination tax. A 2021 Qatalog and Cornell University survey (n=1,000) found knowledge workers spend roughly an hour a day looking for information, and 44 percent cannot easily tell whether work is being duplicated.

The problem was never a lack of AI. It was that AI had no way to join the team.

## The vocabulary, and why it is infrastructure

Before any of the design positions below make sense, the words have to hold still. Across Adobe and across the industry, "agent," "skill," "workflow," and "tool" mean different things to different people, the major AI companies do not agree, and marketing makes it worse. Arnold authored a working vocabulary rather than wait for the world to settle it, on the argument that design, engineering, legal, and leadership cannot make decisions together while translating between four private dialects.

**The settling question: who decides the next step, the code or the AI?** If code decides the order, it is a workflow. If the AI decides the order, it is an agent. His kitchen version: a workflow is a recipe, steps set in advance, never changes its mind. An agent is a cook with no recipe, who tastes the dish and decides what to do next. Almost everything else is detail.

**The structural fix is that these are not the same kind of word.** "Agent" and "workflow" describe how work gets done. "Skill" and "tool" describe what the system uses to do it. Lining them up as equals is the source of most of the confusion. So, four groups:

**How work gets done, meaning who drives the next step.** Rule-based automation: code drives, fixed steps, no AI. Repeatable, cheap, zero judgment. Code-directed AI workflow: code drives a fixed path that calls AI along the way. Predictable, cannot go off-script. Agent: the AI picks its own next steps toward a goal. Flexible, but slower, pricier, and less predictable.

**What the system uses.** A **skill** is packaged know-how the AI reads to do a task well, like a set of brand rules. It is knowledge, not action, and it cannot do anything by itself. A **tool** is a function the AI can call to actually do something. That is where the real power lives, so that is where limits belong. An **action** is the business thing that happened ("published the page"), the plain-language view for audit. A **function** is the code behind a tool, an engineering word only.

That distinction between skill and tool is the cleanest one-line version of the capability-versus-authority split anywhere in this document. Knowledge cannot act. Tools can. Constraints belong on the tools.

**How they are arranged.** Single agent, subagent (a specialist it hands work to), multi-agent system, and agentic system (the whole governed setup: agents, workflows, tools, memory, people, and approval gates). These describe structure, not sophistication. A subagent is not a more advanced agent.

**Context versus memory.** Context is what the AI can see right now, a small space that resets every time. Memory is what you store to bring back later. The AI has no memory of its own; memory is something you build around it. Context is the desk in front of you. Memory is the filing cabinet behind you. The AI only reads what is on the desk.

**The three-question test.** Does the AI choose the next step? Can it pick which tools to use? Can it change its own plan as it goes? If the AI does not choose the next step, it is not an agent. That one rule stops most of the hype.

**And the payoff, which is why the taxonomy is a governance artifact and not a glossary.** Classify the thing and its governance obligations come with it. More AI control, moving toward agent, means more approval gates and tighter limits. Bigger setup, moving toward agentic system, means clearer ownership and escalation. More it can see, context, means tighter exposure controls. More it can keep, memory, means tighter privacy and retention rules.

The category a thing lands in decides how much oversight it needs. Which is why "do we even need an agent?" is a governance question before it is an architecture question.

## The model, at working depth

### The definition, and why it is load-bearing

An AI Collaborator is an agent-powered Workfront user. The precision matters: the collaborator is not the agent. It is the Workfront identity, role, access, and governance layer that lets an agent participate in enterprise work. The agent underneath can be Adobe-built, customer-built, or brought from an outside platform. The collaborator wrapper is what makes any of them something you can assign, inspect, and hold to account in this system.

This separation is what lets the product govern agents it did not create, which became the strategically decisive property.

### Collaborator-as-user

Arnold's framing position: when an agent sits in the assignee column, receives assignments, comments on tasks, and reports progress, the productive design question is not "what type of agent is this" but "what kind of user is this."

Treating the collaborator as a user creates the accountability surface for free. You can mention it, check its workload, review its assignments, read its history, and hold it to the same operational expectations as any teammate. It also forces the hard questions early, because users have identity, credentials, permissions, and audit trails, so agents must too. The line that carries it: it is not a button you click, it is a team member you assign work to. That the assignee column shows humans and agents together is a design choice with an argument behind it, not a default.

One question inside this framing is genuinely unresolved and worth naming: what does "assigning" mean for an agent that can run autonomously? Assignment presumes a queue and a working day. An autonomous agent has neither. That question was left open when the framing was first written down, and it is still open.

### Onboarding is the pattern

Once agents are users, the system inherits patterns the enterprise already trusts. Workfront has always known how to manage people: identity, role, skills, access levels, object permissions, assignment, notifications, work history, offboarding. AI Collaborators reuses that operating model instead of inventing a parallel AI abstraction. An administrator configures a collaborator, a project manager assigns it, a teammate inspects it. The alternative, a separate AI lane with its own concepts, would have fractured the operational processes customers have built around the user model.

Not new AI magic. A new kind of participant in an operating model that already worked.

**The new-hire test.** The framing device for the whole onboarding research, and the reason the analogy is a design tool rather than a metaphor. A new employee on day one gets six things:

1. A role title and department, who they are and what they do
2. A project brief, what the team is working on and why
3. A team introduction, who to go to for what
4. A process walkthrough, how work moves through statuses in this organization
5. Access scoped to their department's projects, not organization-wide
6. A manager to escalate to when blocked

Scored against the setup form as it stood: it covered #1. It partially covered #2. It covered none of #3 through #6.

Each of the uncovered five maps to a specific spec addition, which is what makes the analogy operational. #2 became context injection, a tier of project and hierarchy context always passed regardless of task. #3 and #6 became the escalation contact, a named human the collaborator mentions when it is stuck, required for triage and summarize jobs. #4 became the status vocabulary picker and its semantic annotations. #5 became the scope definition.

The rule attached to #5 is the one that survived into the design language everywhere else: scoping a collaborator is giving an intern a department badge, not a master key card. Without scope, an agent can be assigned to any task in the organization and will run, which is the equivalent of onboarding an intern with organization-wide administrator access.

### The layered model: conversational workspace and work-layer collaborator

Adobe's public agentic portfolio includes CX Enterprise CoWorker, a cross-product conversational agent workspace announced 20 April 2026 and generally available 10 June 2026, alongside Workfront's AI Collaborators. Arnold authored the internal position on how the two relate, because ambiguity here confuses both customers and internal teams.

The layering: the conversational workspace is user-initiated, a person opens it and directs it, and it orchestrates across the marketing suite at campaign level. AI Collaborators are system-initiated, they pick up assigned work without a person asking in the moment, and they operate at the work management layer inside tasks and projects. **The dividing line is who initiates, not how complex the work is.**

The matrix that proves it is more useful than the sentence:

| | User-directed | System-initiated |
|---|---|---|
| Simple or templated | Conversational campaigns module | Rule-based automation, no AI needed |
| Conversational or complex | Conversational chat and projects | Work-layer agent or AI workflow |

The bottom-left cell is the punchline. The simplest system-initiated case needs no AI at all. The same simplest-sufficient-tool conviction that runs through the Coordinator work, showing up in a product-boundary matrix.

A routing rule falls out of it. When a user does the same AI action repeatedly, generate copy for this brief, review this asset, summarize this project state, that is conversational-workspace territory. It is a usage pattern, muscle memory of prompts, not a separate product tier. A work-layer collaborator is the wrong answer there because collaborators require system assignment and autonomous pickup, which is friction when the user just wants to ask.

The two compose rather than compete. A conversational session that plans a campaign can hand execution tasks to collaborators inside the work system. The metaphor is shared, the scope and initiation model are different, and the layering is load-bearing. Arnold set the vocabulary rules to protect it: use the product names, never use a competitor's trademark as a generic term, and never use the two names interchangeably.

### Three provenance types, four setup on-ramps

This distinction matters, and conflating the two numbers is a real error the earlier version of this dossier made.

**Three provenance types** describe what kind of thing the agent is: built and shipped by Adobe out of the box, built by the customer, or brought from an outside agent platform. That is what a collaborator's identity surface renders in its origin field, and it is what the public demo instantiated with one packaged reviewer, one agent from an outside platform, and one customer-built collaborator.

**Four setup on-ramps** describe which front door you walk through: the packaged content reviewer, the wider catalog of Adobe-built agents, an agent the team already runs on a common outside platform, and a custom agent either built in an internal builder or brought in over a standard protocol.

Four does not equal three because the packaged reviewer and the wider Adobe catalog behave differently enough to need different doors, even though both are Adobe-built. The reviewer deploys into an approval flow and only advises. The others deploy to tasks and do the work. **That is the only durable functional split that survives after setup**, and it is the reason the on-ramp count is one higher than the provenance count.

The taxonomy is not an assertion. It has empirical support: a customer-mapped registry of live use cases contains all three provenance types, and the public Summit demo showed all three on stage.

The strategic claim the taxonomy sets up: the product does not need to own every agent. It needs to own the layer where agents become accountable.

The commercial argument underneath that claim came from a customer objection Arnold put into one sentence: **"I already have my agents, why do you want me to build them again?"** Everyone is already building agents. Forcing customers to rebuild inside the system turns them away. The mechanism that answers it is task binding: you tell an external agent what it means to attach to a task, how to report progress as a task update, and what it means to be assigned. The system brings all the operational plumbing. The customer brings the agent and its capabilities.

And task binding is where the governance move happens. An external agent enters, establishes identity, has its external authority stripped, and is assigned authority within the delegator's envelope. **Trust does not transfer in.**

### The Content Reviewer precedent

The first shipped collaborator reviews assets against brand guidelines, scores them, and flags what needs attention. Adobe's own materials state the boundary directly: the Content Reviewer is not designed to be a decision-maker, it only provides a score and recommendations. It is assignable to approval templates or individual review requests, the same way you would assign a human reviewer.

The agent is technically capable of making the accept-or-reject call. It is deliberately constrained to scoring and recommending, with the decision left to a human. Arnold's reading of that constraint: it is the most important precedent in the product. The first collaborator shipped with an explicit line between what it can do and what it may do, and every future collaborator should carry an equally explicit line, stated in language a human can act on. A shipped product choosing restraint as a feature is rare, and it is the precedent the governance work generalizes.

Four registry facts from public Summit 2026 materials round out the shipped governance model: collaborators are created and managed in a central governed registry by admins or designated owners; configured with role, instructions, data access, and triggers; enabled, paused, or refined by admins; and governed by role-based access, inheriting the platform's existing roles and permissions. Adobe's own stated model: every action is explicit, logged, and auditable.

Both ideas from the first pitch are in the product, in Adobe's own words. A single governed registry. Clear boundaries an administrator sets. The marketplace did not survive. The registry and the code of conduct did.

### The profile card, and the authority gap

The collaborator profile card is the product's answer to "what is this AI teammate and what is it doing": identity, skills, connected context such as brand guidelines, current plans, recent deliverables, and operational stats like accuracy and time saved.

Arnold's standing critique, maintained across the program: the card describes what a collaborator can do and what it is doing, but not what it is permitted to do, on whose authority, within what limits. Capability and activity are visible. Authority is not.

The diagnosis is structural rather than cosmetic, and the evidence is specific. The setup flow has four sections, Details, Context, Connection, and Instructions, and no Authority or Permissions section. **Authority is handled implicitly through inherited roles, not surfaced as its own object a human can read.** That clause is the whole critique.

The product has since narrowed the gap. Live onboarding added a coarse allowed-actions section, a short list of checkboxes covering comment, field update, upload, notification, status change, and form scanning, with the two lowest-consequence actions on by default. Arnold reads that as directional validation. But he names precisely what is still missing rather than saying the gap "narrowed," which is falsifiable and therefore more credible: **no blast radius, no on-whose-authority, no conditional confirmation.**

This is the exact point where the AI Collaborators work hands off to the Delegated Authority Framework, and the handoff is governed by a boundary rule he keeps deliberately. AI Collaborators owns the product surface: how agents appear, get assigned, communicate, and earn trust. DAF owns the authority layer: what agents may do, when they need approval, and who remains accountable. Cross-link, do not merge. The failure mode the rule guards against is named too, and it is the memorable half: **the risk is AI Collaborators becoming "DAF, but with Workfront screenshots."**

He is also careful not to overstate the gap. An honest audit of what the product had already built, recorded in his own notes, is worth reproducing as a posture: governance is not an oversight the program forgot. Identity exists, collaborators are digital users linked to real accounts. A standardized mutation-approval path was already being engineered. Auditability of collaborator activity was already a program entry criterion. So the governance framework's white space is only the part that genuinely is not built: **classification before the gate, the delegated-authority contract behind it, and a legibility surface so a human can read what a collaborator may do.** It is the depth the product's own permissions work requires, not a critique of a blind spot.

### The collaborator ID card: claimed versus granted

Concept stage, Arnold's own, prototyped July 2026. This is the executed answer to the standing critique above, and the gap between logging the idea and building the answer was sixteen days.

Purpose in one line: **the card exists so a human never has to trust an agent blind.** Four questions a person can answer at any moment without reading a config file or asking an admin. What is this agent? Where did it come from? What can it do here, and what is it not allowed to do? What has it actually been doing?

Positioning decision: it is a legibility surface, not a control panel first. Reading comes before governing. The controls live one layer in.

**The spine is the separation of claimed from granted**, and it is the framing decision that drives the whole design. The card separates what the agent says about itself, the manifest, self-reported and treated as input rather than truth, from what the system actually granted it, which is enforced. Both get equal visual weight and opposite trust framing. The claim zone is captioned "declared by the agent, reviewed not trusted." The granted zone: "enforced by the system, this is what it can actually do."

The reasoning is the important part. If the card blurs them, it recreates the exact failure the governance model exists to prevent: **a black box that reads as trustworthy because it described itself well.** Compressed: capability is the engine, authority is the license.

Provenance carries the strongest analogy in this whole body of work. The card shows origin and an "acting as" ownership line, and states that whatever authority the agent held elsewhere was stripped at onboarding, and that it started at the most conservative tier here. **A fully trusted agent elsewhere still starts conservative, the same reason a new hire with great references gets no signing authority on day one.**

The authority section is built on an anti-checklist argument. Skills are shown as a consequence of the authority ceiling, never as a checklist the user toggles, because a checklist implies the human grants capability directly, which inverts the model. Withheld skills are shown with the binding reason named: exceeds your authority, organization policy, not declared. Trust tier is stated in plain language and framed as earned and revocable. Per-action oversight rows show, for the actions that matter, whether it runs silently, runs and notifies, needs approval, or is blocked.

History is evidence, not an audit dump. The log is the evidence behind the trust tier, because the track record is literally what earns or drops trust. The tier on the card and the history behind it are the same story from two ends, the summary and the proof. The ledger deliberately interleaves two kinds of entry: action entries, with the consequence stated in plain language rather than "an action occurred," plus mutation level and oversight outcome, and governance entries, trust-tier changes with their reason and the identity-mint event at onboarding. Every entry traces to a human root and carries a signed, append-only identifier. **"The agent did it" is never the answer.**

### Trust tier and mutation, the two axes

The oversight model is a matrix, not a switch. Mutation level is set by the system, based on what an action actually changes. Trust tier is set by a human. Oversight is the computed result of the two and is read-only, which is what keeps "skills follow authority" true and stops the surface from degrading into a permission checklist.

| | Low mutation | Medium mutation | High mutation |
|---|---|---|---|
| Conservative | Human approves | Human approves | Human approves |
| Guided | Runs, notify | Human approves | Human approves |
| Trusted | Runs, silent | Runs, notify | Human approves |

Two rules attached. Runs-silently is trusted plus low only; at the guided tier a low-mutation action runs and notifies, it does not run silently. And an external override: anything crossing an external boundary escalates one step above the matrix.

**The two axes move independently, and collapsing them is the most common way the model leaks.** Blocked is the ceiling axis, whether an action is permitted at all. Oversight, silent or notify or approve, is the trust-by-mutation axis, how closely a permitted action is watched. Flipping scope to allow external publishing moves that action from blocked to needs-approval with no tier change at all.

That two-independent-axes move is the single most characteristic pattern in Arnold's thinking. It shows up three times in three unrelated domains over eight weeks, each time with the same warning attached: guidance plane versus authority plane, complexity ladder versus blast radius, ceiling versus oversight. Each time the design job is to keep two variables that look like one visibly separate.

He also declines to overclaim it. In the Coordinator build spec he records that in the sample data the two axes loosely correlate, agents and workflows do tend to be higher blast radius, and instructs: treat them as separate in code, do not assert independence as fact.

### Setup is an Authority Contract

The reframe that changes what the setup form is. **Not a settings panel. An Authority Contract.** Five governance primitives, each of which is already a field on the form:

| Governance primitive | The setup field it already is |
|---|---|
| Delegator | The admin issuing the delegation |
| Scope | The projects the collaborator may work in |
| Time to live | The collaborator's lifecycle state |
| Blast radius | Trust tier by mutation level of each permitted action |
| Revocation | The ability to pause or retire |

Two consequences. The vocabulary changes: not "configure access" but "define what this collaborator is authorized to do," not "choose output actions" but "set the blast radius." And **the save button creates a delegation, not a record**, so the confirmation should read like a receipt: this collaborator will wake when a task's status changes to a named state, can comment and upload documents, and will ask for approval before marking any task complete.

Two-plane authority made concrete at setup: an admin can only delegate authority she has. An admin with standard access cannot grant an agent system-administrator capabilities. The form should surface this quietly, so admins understand the constraint is real rather than arbitrary.

Mutation classification replaces the binary allow/block model. Low mutation, reading a task or posting a comment, runs automatically and is logged. Medium, updating a custom field or uploading a document, runs and notifies or asks for approval depending on tier. High, changing task status or marking a task complete, requires approval. Without that classification, "mark task complete" has to be either always on or always off.

Trust tiers are the governance vocabulary sitting over the platform's access-level vocabulary. Conservative, guided, trusted. Selecting a tier pre-fills every per-action state and the admin can still override individually. The argument for why the layer exists: the access-level picker is the platform's language, but the admin is really choosing how much autonomy to grant, and that is a trust decision, not an access decision.

And the piece that makes the contract a contract: **a collaborator with no pause button is an agent you cannot stop.** New collaborators save into draft and stay there until an admin explicitly activates them.

**Trust growth, as designed mechanism.** Earlier framing treated trust expansion as philosophy in search of a mechanism. It now has one, at concept stage. Trust growth is neither granted-and-held nor automatically earned. Both alternatives were rejected: granted-and-held rots, because context changes underneath a permission, and auto-earning breaks the human's last word. Instead, the system builds a ledger-backed track record, proposes a promotion with the evidence attached, and the promotion goes through the same gate as everything else. A human approves, the ledger records it, and a time to live means authority decays without renewal rather than persisting into a changed world.

It is employment with performance reviews, where the review is a human decision informed by a record, not an automatic ladder. **A tier change is itself a mutation**, so it gets governed like one. The same propose, classify, gate, execute shape, pointed at authority itself.

### One spine, four on-ramps

Four setup on-ramps threatened four separate setup products, and with them a redundancy critique from within the team: if setups differ this much, are these even one product?

Arnold's answer, now the architecture of record for the setup journeys: where the agent comes from changes the setup. Everything after the badge is the same journey. The spine is five stages. **Badge**, the collaborator receives its identity and its authority. **Authority**, what it may do is set. **Deploy**, it is placed into the team's workflows. **Works**, and this stage carries a real design claim, it starts only when its assigned task can start, its dependencies satisfied, then does the job and posts back. **Review**, humans see, judge, and adjust.

That fourth stage is small and tells you something. An agent is subject to the project's dependency graph exactly like a human assignee. It does not fire whenever it is triggered. It waits its turn. The workforce thesis expressed as scheduling physics.

The on-ramps differ before the badge, because connecting an outside agent legitimately requires different steps than switching on a packaged one. After the badge, nothing forks except advise-versus-do.

The architecture dissolved the redundancy critique structurally rather than rhetorically: the journeys are visibly one product with different front doors. It also gives setup design a quality bar. Any origin-specific complexity must justify itself before the badge, because after the badge nothing is allowed to fork.

A related rule from the builder-side work, aimed at the same seam: **two on-ramps, one river.** Both paths must end in something that shows up as an ordinary collaborator, behaves like every other collaborator, and carries no permanent "not from here" tag. Do not let connected agents become second-class citizens. But the governance asymmetry is real and should be named rather than smoothed over: **a third-party agent is a stranger.** The system did not build it and cannot fully inspect it. The build path can be looser because you built it and already know what it does. So the connect path gets the stricter governance, and the origin becomes invisible to whoever uses the collaborator while staying visible to whoever governs it.

### Identity: name, color, emblem

Concept stage. Creation-time identity, which is a different question from the ID card. The card answers what this teammate is once it exists. This answers how you give it a face when you make it.

The flow is deliberately modeled on an existing pattern in the product for adding a record type, name plus color swatch plus icon, so it reuses something users already know. It expands a colleague's prototype, which minted an identicon-style signature by hashing an agent's identity into a mirrored emblem. Arnold grew that single idea into a three-part identity builder where each piece is grounded in a real standard or a real evidence source rather than picked at random.

**Name, on an actual naming standard.** Adobe's internal AI naming guidance sanctions three strategies: action focused ("Review brand copy"), context focused ("Brand copy review"), and concept direct ("Automated copy compliance check"). Names must describe utility, what the user can do, get, or understand. Three vocabulary classes are banned and enforced by a live checker that flags a typed name and explains why: system terms (agent, bot, engine, orchestrator), personas (copilot, assistant, buddy), and hype (AI-powered, smart, intelligent, autonomous). Every name ships with a plain one-line function description.

There is a deliberate irony in this worth stating outright, because it is an anti-anthropomorphism guardrail operating at the naming layer: the product is called AI Collaborators, and its individual agents are forbidden from being named "agent," "assistant," or "copilot."

**Color by domain, grounded in evidence.** Design-system color tokens mapped to Experience Cloud domains, with the grid leading with the four domains that have real customer use cases behind them rather than presenting all domains as equals.

**Emblem, built rather than picked.** A hash of the name seeds a deterministic generator that fills a mirrored grid, so the same name always produces the same emblem and the emblem is a visual signature of the identity. The system proposes from the name, then hands over control: pick from candidates, regenerate, tune density, grid size, and cell shape. Once the user edits, name changes stop overwriting their emblem until they re-sync. Candidates are filtered to a 30 to 70 percent fill so none come out blank or muddy. Accessibility is specified rather than assumed: solid fill uses the luminance crossover so the glyph always maximizes contrast, and every color clears at least 4.6:1.

The thesis: every part of the identity is grounded, not decorative. And the shape is the same one that runs through everything else here. The system proposes a coherent identity, the human adjusts anything, nothing is locked by the machine.

**Status as of 2026-08-07.** Arnold is leading the icon and avatar system with Adobe's brand team, and has proposed the v1 direction for collaborator branding and avatars. This is a different discipline from the interaction design that surrounds it, and he names it as one of two capabilities he is building deliberately rather than one he already had. The colleague's signature-emblem prototype remains the seed of the emblem half; the naming standard is authored elsewhere on Adobe's AI framework work. What is his is the three-part builder, the grounding of each part, and the accessibility specification.

This also resolves a tension the earlier version of this dossier only hedged around. The design language stays at badges, records, and boundaries rather than personality, and this is how you give a machine a face without anthropomorphizing it: a hash-derived mark instead of a character, a utility name instead of a persona name, a domain color instead of a mood.

### The admin, the operator, and the requester

An early setup form tested poorly, and the finding was diagnostic rather than cosmetic. The form asked a project manager to make decisions she does not have the knowledge to make, about connection methods, environments, and access levels. The design response was to split the audience rather than simplify the form. Collaborator setup is an administrator's surface, where depth is appropriate rather than a usability failure, because admins are the people who provision identity and access for every other kind of user too. Operators, the project managers and teammates, get the assignment experience: choose a collaborator an admin has already made safe, and put it to work.

The sequencing principle attached: **making setup easier means requiring less knowledge, not bolting on automation that does not exist yet.** Ship the honest admin surface first, then earn the simpler surfaces on top of it.

There is a third lane, and it is the one that resolves an apparent contradiction in the two-audience story. A non-admin can browse the collaborators an organization has, view read-only capability summaries, and submit a draft request that an admin approves before it goes live. Arnold names the pattern **user configured, admin approved.** Request content is four areas, effectively "describe the teammate you want."

Two things make this more than a workflow detail. First, it is governed delegation showing up lightly at the setup layer, and it independently confirmed the admins-first sequencing decision from a separate workstream. Second, and more interesting, engineering confirmed a hard constraint: even when a non-admin configures or requests a collaborator, the actual user is created with the admin's credentials. So the requester flow is credential-free and ends in a pending request. **Identity and authority are issued by the admin at approval.** Never draw a path where a requester supplies creation credentials or mints identity.

That is a governance principle he had argued abstractly turning out to be a hard constraint of the shipped system: the agent's authority traces to a human root, and identity is issued at the boundary rather than transferred in.

And it is where simplicity actually belongs. The requester never touches access levels or API keys. The admin handles the technical configuration at approval. Two audiences, one model, joined by an approval.

### The ladder of complexity, and the handrail

Concept stage, and the fullest form of the "understandable, not easy" position.

It exists to answer the question every reviewer asks about a creation surface: is this tool for non-engineers or for power users? The answer is both, and the way to make that true without shipping two disjoint products is to build one ladder, not two modes.

**Two named failure modes.** **The cliff:** a beginner generates something, it mostly works, one thing is off. They hit "advanced" to fix it and land in a raw node graph, a room never designed for them, with no way back to the sentence-level fix they understand. They are stuck, and they quit. The failure is not at first build; it is at first repair. **The drift:** the two modes are built by different logic, so they cannot always represent the same thing. Edit in the graph, flip to simple mode, and simple mode either hides the change or breaks on it. The two rooms disagree. The seam shows, and trust dies at the seam.

**The metaphor.** You do not switch between a simple map and an advanced map. It is one map. Zoom out to see the country, zoom in to see your street. Same map the whole way. Nothing breaks when you zoom.

**The three rungs.** Zoomed out, describe it in plain language. Mid zoom, shape it with visual blocks or a configuration form. Zoomed all the way in, author the raw definition by hand. One source of truth underneath all three.

**The test, and it is falsifiable.** Can someone start with a sentence, zoom in to fix one small thing, and zoom back out without anything breaking? If yes, it is a ladder. If the zoom-out forgets what they did up close, you accidentally built two modes. That is the bug even if it demos fine on the happy path. The corollary invariant: every rung must be able to render, even read-only or degraded, whatever any other rung produced. The moment a lower rung cannot represent what a higher rung did, the ladder snaps.

**The handrail, which is the least obvious idea here.** The conversational assistant is not a fourth rung. It is the handrail that runs alongside all three rungs the whole time, the thing that never disappears no matter how deep you zoom. The handrail is what makes the no-cliff promise real: you can descend to a rung you do not understand because the assistant is standing on it with you.

**And a guardrail against his own idea being softened.** The ladder makes the tool understandable at every depth. It deliberately does not make powerful agents easy. Do not let a future simplification collapse "no cliff" into "no effort." They are different claims. The reframe that survives a room full of executives is that the user is never stuck and the effort pays off, not that the work becomes trivial.

A companion position on the assistant beside the form, from the same body of work. The assistant can see what the user is modifying, step by step, so the user can stop and ask "what is this?" and be understood. The language discipline is explicit: do not say the assistant has access to the document object model, which is engineer-brained and undersells it. Say plainly, **the assistant sees what you see.**

The hard part is named rather than hidden. Two hands on the same clay: the user edits the form directly and the assistant edits through conversation, and both change the same object. That requires one source of truth both read from and write to, live. If the form and the conversation ever hold different versions, the illusion collapses. The argument for naming it in a pitch instead of glossing it: it is hard, which is why it is worth doing and why a weekend competitor cannot fake it. **Difficulty disclosed as moat.**

And an observation about where products actually lose people, which is a testable claim rather than a truism: **people do not churn on first build, that part is fun. They churn on coming back confused after a bad test run.** So a creation surface's real job is not only authoring. It is holding state across the round trip, remembering where the user was, and helping them read what went wrong. A pitch that only shows the happy first-build path will smell wrong to skeptics.

### Absorb, translate, illuminate

The jargon problem, stated precisely: users setting up an AI collaborator get lost in terminology that is the vendor's model of the world, not theirs. They are asked to pick a type, an origin, an authentication method, a connection endpoint. None of that is the user's mental model. The user's mental model is a job to be done: I want something that reviews my campaign copy for brand compliance.

But "how much jargon can disappear" is not a dial. The terms split three ways, and the split is decided by whose consent depends on the term:

- **Plumbing** (origin, authentication method, connection endpoint, the underlying definition format). The user never needed these. **Absorb fully.**
- **Role vocabulary** (task agent, reviewer, coordinator). Load-bearing for expectations, because it tells you where the thing shows up and acts. Do not show the term, **translate** it: "this one works inside tasks," "this one reviews proofs." The official label lives further down the ladder for power users.
- **Governance vocabulary** (access level, allowed actions). Must not disappear. That is the consent surface. **Illuminate** it, never hide it.

Compressed: absorb the plumbing, translate the roles, illuminate the permissions. Taxonomy does not vanish, it moves down the ladder. Nothing is lost; the zoom level where it lives changes.

Note this is a proposal from a working session that Arnold had not yet graded point by point at the time of writing. It is included as a position under development, not settled canon.

### Honest discovery, and test as oracle

A set of positions about what an AI-assisted setup surface is allowed to claim it knows. They add up to an epistemics of AI interface design, and they are more transferable than the product they came from.

**Name by search, not free text.** Do not ask a user to type a platform name into the void. Names are an infinite sea and one misspelling becomes junk data. Match natural-language input against a known-platform knowledge base, confirm with candidates on a near miss, fall back to an unknown path on a clear miss.

**Adaptive fields, not a fixed form.** There is no honest fixed field set for "any platform." A hosted platform needs no connection endpoint because its address is fixed and known. An instance-based one does. Credentials are shaped to the authentication method. The minimum universal requirement is only a way to reach it and a way to authenticate, and even those bend. This is an empirical claim, not a stylistic one: across a survey of major agent platforms as integration targets, the required parameters vary from a single API key to a tenant identifier plus service principal plus project endpoint plus role assignment plus signed token exchange.

**Test as oracle.** The system does not need to know for certain up front. The connection test is what turns a guess into a fact. Propose fields, attempt the connection, read the real error, adjust. **Never state an inferred field as fact.** Say "I think this one needs X, let's test it," not "this one requires X." That is the honest answer to "but how does it actually know?" and it survives a skeptical room far better than claiming the model knows every API.

**Three honest tiers of discovery.** Best: the platform self-describes through a standard, an introspection endpoint, an interface specification, a well-known discovery document, and the system reads that machine-readable contract. Certain. Good: the platform is in a curated catalog, meaning a human discovered it once. Fallback: no standard and not catalogued, so the system reads the documentation and guesses, then confirms by probing. Plus the demystification: mechanically, "the assistant discovers" means the model plus tools that fetch, parse, introspect, and probe. Not the model guessing alone.

**Secrets never go through the chat.** The assistant guides the user to where the key lives and tells them to paste it into the secure form field. The secret is never collected in conversation. Non-negotiable, and part of the governance story rather than separate from it. Its sibling: the assistant that sees what you see needs a redaction rule, and should say so.

**And a boundary on the honest-discovery story itself.** Test as oracle weakens exactly as ambition climbs. A connection test is binary, cheap, instant green. A behavior test is open-ended and has no green checkmark. Grading whether an agent behaves needs golden tasks, sandboxed dry runs, and sample outputs a human judges, which is a different and much longer pole. Name that in a pitch rather than letting "the engine exists" imply otherwise.

**Do not mistake the mock for the mechanism.** A scripted prototype proves the experience. The real engine is the build behind it. Stated in the source as a rule for anyone presenting the work.

### Right-sizing: the design position that argues for doing less

A collaborator does not have to be an agent. It can be a simple deterministic automation, a richer multi-step workflow, a reusable skill invoked on demand, or a full agent that reasons, uses tools, and acts with autonomy.

**The system should propose the smallest thing that does the job rather than defaulting to an agent.** Forcing every request into "build an agent" is both harder than necessary and scarier for governance than necessary. If the honest answer is an automation, propose an automation. Do not talk the user into an agent.

The argument is governance-based rather than cost-based, which is what makes it interesting: right-sizing lowers the trust burden. A deterministic automation is easy to reason about. An autonomous agent needs the full authority model. **Refusing to build an autonomous agent when a rule would do is itself a governance feature.**

An honest read of a real customer use-case registry supports it and cuts against the program's own breadth claim at the same time: most real asks are automations or skills, not autonomous agents, and most cluster on marketing content. That is a narrow beachhead. It is a strength for a focused pitch and a risk if a slide implies the platform does anything.

Right-sizing operates at two moments. At setup time, the assistant proposes the smallest sufficient artifact. At runtime, a router picks the simplest sufficient technology per task. Same ladder, two moments, one principle.

### The journey presentation method

A repeatable method came out of presenting this work to design leadership, in direct response to an ask that arrived as a single sentence: "Who is the persona. What is the job. Show from awareness, to setup, to success. Need more UX discussion and less UI discussion."

Five moves, now codified:

1. **Lead with a person, not a lane.** Persona chips plus one job statement on top. Never open on the system.
2. **Stages, not nodes.** Five maximum: need, set up, deploy, agent works, success. Time flows left to right.
3. **Demote the system.** Product mechanics become one muted line under each stage. Credentials and event subscriptions are exactly the interface discussion leadership said it did not want.
4. **One branch, told as a human moment.** "Not on brand? The reviewer sends it back." **No decision diamonds.**
5. **End on value, not task state.** "Review cycle three days to one," never "task marked complete." Job-to-be-done identifiers go to collapsible footnotes, so the rigor is preserved and demoted rather than removed.

**The insight that justifies the format, and it is the best part.** Each journey spans two people: an admin sets it up, and someone else lives with it. **Swimlanes bury that handoff. A journey map makes it the headline.** That is the "who is the persona" answer in one line. The format was chosen partly because of what it structurally reveals, not because it looks better.

Two craft decisions attached. The swimlane specifications were kept as the detailed spec rather than replaced; two artifacts for two purposes, not a conversion. And every path carries its real maturity as a tag, direction versus buildable now, with an explicit note that the journeys show the target arc and not the as-built prototype. The as-built disclosure goes further than most: it names a place where the prototype's own default contradicts his research, an oversight default set to permissive where the research argues for approval on higher-mutation actions.

Tone rule for the whole deliverable: **governance present, not shouted.**

## The Project Coordinator

Concept stage. Reviewed by leadership, directionally endorsed, not approved. This is a body of work the earlier version of this dossier did not mention at all, and it is where the governance thinking gets its most complete expression.

### What it is

A runtime layer that, for every task in a project, picks the simplest technology that will do the job and the right amount of human oversight, runs the work, and **keeps the human in the governance seat rather than the labor seat.**

The core intellectual move: "find the simplest solution" moves from a one-time design-time choice an engineer makes into a per-task runtime routing decision the system makes, continuously and at scale. The lineage is credited: the simplest-sufficient-tool principle is borrowed, and the contribution is promoting it from a design-time judgment to a runtime routing loop.

Its relationship to the rest of the program is drawn cleanly. The product owns how agents appear, are registered, assigned, and earn trust. Onboarding is a separate flow and explicitly not part of this. The authority framework owns what agents may do and who stays accountable. This work is the runtime-governance track: what happens once a collaborator is already on the work.

### The two axes

**The complexity ladder** decides which technology runs a task: rule-based automation where code decides and no AI is involved, an AI workflow that is code-directed but calls AI along the way, a full agent that picks its own steps, or an advisory agent that scores and never decides. Simplest sufficient wins.

**The blast radius** decides how much oversight it needs: runs on its own, runs and notifies, or needs the human. **Set by consequence and reversibility, not by capability.**

Keeping them separate is load-bearing. A deterministic tool that publishes externally is still high blast radius and still gets gated. A clever agent doing read-only work can run silently. Collapse the axes and the model leaks.

Why this is governance rather than architecture: climbing the complexity ladder grows the authority surface, because a workflow's path is knowable up front and an agent's path is decided at runtime. **Authority stops being optional exactly when you reach for an agent.** Which is why most enterprise work should not be a full agent at all. Automations and fixed workflows are cheaper, faster, more auditable, and more predictable, and reaching for an improvising agent on every task is how you get cost, latency, and risk you did not need.

### The self-referential risk

The sharpest observation in the concept, and the one that says the most about how Arnold thinks:

**The router is itself an agent, and the highest-blast-radius component in the system, because everything flows through it. The real risk is not over-using agents. It is the router quietly becoming the ungoverned agent while everything downstream looks tidy. So whatever governance gets designed, the router gets it first.**

### What the prototype demonstrated

A ten-task product launch with a real dependency chain, run end to end as a scripted clickable prototype. Every agent, tool, score, and deliverable is faked so the story demos reliably without live models or a backend, and the sources say so openly and repeatedly.

**Assign.** The coordinator presents the plan: ten tasks, each matched to a technology and tagged with how it will finish. The human approves and work begins in dependency order. The teaching point: she never picks the technology and never sees the mechanism. She sees what each task will do and what it will touch. **The system picks the technology; the human approves the consequence.**

**Stay in the seat.** Four mechanisms, and this is the most original material in the demo. **Auto-checks before the human ever sees it:** every task verifies its own result against the brief, goals, and brand guidelines, and low-blast tasks clear themselves so the human is never bothered. **Agents checking agents:** on a set of banner and social variants, the coordinator loads a second tool, the content reviewer, which scores the first pass poorly, sends it back to the producing agent, shows a version comparison, re-checks, passes, and clears, all automatically. **The revision loop:** on customer-facing copy the coordinator pauses and hands over the draft, and the human directs two rounds of changes with the headline rewriting live. She never touches the work; she directs it. **Adding a capability on the fly:** when a task hits a wall it lacks a tool for, the human picks one from an autocomplete and it is added to the task, unblocking the work. **The human grants a capability at runtime, under the ceiling.**

**Govern authority.** An Authority Inspector on every task detail page showing delegator, scope, blast radius, expiry, revocation, and a capability-versus-authority line. It is, explicitly, the authority section the profile card is missing today.

Then the beat that carries the whole argument. A task builds fine, but taking it live with paid media spend exceeds the human's own authority envelope. The coordinator does not ask her to approve something outside her envelope. It blocks and routes to the person who holds that authority. The on-screen principle: **an agent cannot do more than you can, and you cannot grant what your envelope does not hold.**

And the publish gate, where the highest-consequence step shows a blast-radius classification and a consequence-legible approval: publishes to the site, emails a specific number of contacts, posts to four channels, cannot be undone. Not a vague "confirm?" On approval it writes a signed ledger entry tracing the action back through owner, ceiling, assigner, grant, and action.

### The five product-design decisions

1. **Consequence language over technology language.** "This updates 40 tasks across 3 projects and cannot be undone" is a decision a human can make. "This uses an automation scenario" is not.
2. **Three views, one state.** Switching views never loses progress.
3. **Task-centric, not agent-centric.** Cards lead with the task and whether it is on track, not with a named bot. The technology is a quiet tag. The human cares about the work being done correctly, not about which agent did it. This is a deliberate counterweight to the workforce metaphor.
4. **Multi-session conversation.** Per-task threads rolling up into one project thread.
5. **Scripted, not live**, stated openly rather than implied.

There is also a logged, deliberate violation of his own principle: the taxonomy says end users should normally see only authority and the gate, not the mechanism, and the demo deliberately reveals the mechanism so the audience can see the routing. It is recorded as a knowing deviation with its expiry condition attached, that production should hide it or gate it behind an inspect action.

### What was learned, including from failure

**The prototype went through three generations, and the evolution is itself the finding.** The first was twelve moments of rich craft arguing per-task authority mechanics: excellent at its argument, wrong for the leadership audience. It was kept alive unmodified as the engineer deep-dive rather than trimmed, on the reasoning that trimming it into the new story would produce a demo that served neither audience.

The second was stripped to a governance argument, and **it broke in a way nobody predicted.** It read as though the agent had acted on its own. The root cause was structural rather than cosmetic: it was a morning-after story told in past tense. The human walks in and the system has already triaged, routed, stopped tasks, and detected drift. The only authorization is off-screen. She never says "go," so every beat reads as unilateral.

The fix was one added beat, an up-front plan approval, and flipping the entire story from past tense to present. That single beat does four jobs at once: it is the plan-approval moment, it sets the autonomy dial, it voices the routing thesis, and it makes every later beat read as carrying out the plan the human approved.

**Tense determines perceived agency.** That is the most portable finding in this whole body of work, it came from a diagnosed failure rather than a theory, and it applies to any agentic product demonstration.

The third generation added a real-time reframe and visible governance controls. Three weeks later its own decision record was downgraded: a header note added to say that those twenty-four calls bind the demo build and not the broader program thinking. A design record policing itself against demo decisions hardening into direction just because someone wrote them down.

**On differentiation, honestly stated.** This is not a policy engine. A policy engine can be the gate, but it cannot classify a probabilistic agent action's blast radius before deciding whether policy even applies, and it offers no user-facing legibility.

**What is genuinely new versus what is reuse**, stated identically across three separate documents, which is a good sign it is a settled position rather than a convenient one. Reuse: the trust-by-blast-radius matrix, the runtime grant, the capability catalog. Genuinely new: **the router itself, the blast-radius classifier (called the most load-bearing and most underspecified piece), the in-flight checkpoint for governing an agent while it runs, and the coordinator conversation as the human's governance surface.** Those are the parts worth designing next and the parts most likely to be wrong.

## The research foundation

Four distinct research efforts, using four distinct methods. Naming the methods separately matters more than the findings in some cases, because the method is the transferable part.

### The onboarding research (Arnold Porras, June 2026)

Before designing collaborator setup, Arnold ran a research pass against a live enterprise environment at real scale: sixteen thousand users, 1,926 custom forms, a task status vocabulary of more than seventy codes over three base states, a separate issue status vocabulary of more than ninety values, sixty-three entity types, and more than thirty object types emitting subscribable events with sub-second average delivery.

**The thesis of the study, in his own framing:** an AI collaborator entering this organization is not entering a generic project tool. It is entering a heavily customized enterprise system. Generalizations about "the platform" are useless. The unit of design is a specific deployment's local reality.

Method: twelve research questions across six tracks, answered through live environment queries and public API documentation, with each question carrying a pre-registered method and a stated reason it matters before any answer existed. Some questions were answerable from documentation, some from live queries, some only from sandbox experiments, and some were design judgment, and the plan prioritized them accordingly. One question was designed as an explicit falsification test with two named possible outcomes: attempt a write with a restricted token to a project the token's user is not shared on, and confirm whether the platform rejects it or silently no-ops.

**The specification audit.** The most transferable artifact in the study. A seven-row table structured as "current spec assumption" against "what the real environment shows." He did not just gather findings. He enumerated every assumption his own specification made and ran each against reality, recording which broke. All seven broke.

Selected findings, at the level of detail that changed the design:

**Silent failure is a configuration-time design problem.** A collaborator can be configured with instructions its access level cannot execute, and the mismatch surfaces as a silent runtime failure, not a setup error. The chain is exact: "mark task complete" is an output action in the spec, marking complete requires changing status, a contributor-level identity cannot change status, so the output action fails silently at runtime. The design consequence is specified down to the copy: couple access level and permitted output actions in the interface, disable the impossible checkbox, and explain why, with a one-line consequence description under each access level.

**Attribution must be designed, because the platform does not provide it.** Nothing in the environment natively distinguishes agent activity from human activity. An agent's comment is indistinguishable from a person's. Machine attribution therefore has to be constructed deliberately: named agent accounts and a signature convention on every agent action, so the record answers "who did this" without forensics. Trust in the activity log, the beat that mattered most in the origin pitch, turns out to be something you must build rather than something you inherit.

**Status is a language agents must be taught.** Status vocabularies are customized per organization, and each code carries different go, wait, or stop semantics for an agent. The mapping is three-bucket: strong go signals that are role-specific (a design-approved state means go for a copy agent), wait states covering the four separate "awaiting a human" conditions plus hold, and escalate states. Plus the trap that proves the point: a "done" status can mean the work is finished but still needs administrative completion and can sit in an in-progress base state, while an "awaiting deployment" status sits in the complete base state. Two near-identical labels, opposite meanings. An agent reading raw codes will misfire constantly.

The deeper finding is that statuses do **three jobs at once**: pipeline stage, handoff signal ("awaiting content" means the writer's turn), and risk flag ("blocked"). The durable answer is to unbundle the three rather than teach an agent to disambiguate seventy strings.

**Scope should feel like a badge.** The two-layer permission model, organization-wide access level crossed with per-object sharing, supports scoping a collaborator to specific projects, and enforces it automatically at the API. But only if an admin has not over-shared the agent's account. Recommendation: standard access, scoped by sharing, never organization-wide. Default scope should be the current task only; wider scope is opt-in.

**Ambient context is a minefield of near-synonyms.** Four different due-date fields coexist on a task with different semantics: human-set intent, the assignee's personal commitment, a system calculation from actual progress, and a raw system estimate. The agent rule is stated as a prohibition: read the planned date for deadline context, never overwrite it, never confuse it with the others. Two of the four are system calculations differing only in what they ignore, and nobody can explain the difference without the documentation open.

**Some correct agent behavior needs no configuration at all.** A ten-row inventory of task fields always present in the event payload that an agent can use for runtime decisions with zero setup: a blocked flag means escalate, a ready flag is a confirmed trigger, a not-ready boolean means predecessors are unmet so wait, a late or at-risk progress status means flag it, a major-roadblocks condition means escalate, a child count greater than zero means this is a parent task so be careful, and percent-complete at either extreme tells you whether there is runway or whether the work is already done. Governance you get for free, from data the platform already emits.

**Issues are a whole object class the model missed.** Issues are a separate entity with their own ninety-plus status vocabulary, different fields, a different lifecycle, and the ability to convert into projects. A collaborator that does not know whether it is operating on a task or an issue will behave wrong, and the distinction was invisible in the setup flow. The diagnostic is sharp: the setup form's job-type question does not surface whether an agent will ever encounter issues, and a triage collaborator almost certainly will.

**Context injection as a fourth tier.** Project name, description, owner, status, the task's position in the work breakdown, its condition, and whether it is a subtask, always passed regardless of task. The rationale in one line: **the admin does not configure these, they are always included, the way an intern always gets a project brief on day one.**

**Two proposals worth naming.** A vocabulary mapping field, where after the admin picks an activating status a free-text box asks "what does this status mean for this collaborator?", stored as a key-value map beside the code. That is the mechanism for teaching an organization's working language, and it exists as a designed proposal rather than an open wish. And a mocked dry run before save, showing against a real sample task what the agent would read, what it would do including the exact comment it would post and a warning line for any action that requires approval, and what it would not touch. The receipt tells an admin what the collaborator is authorized to do. The dry run shows what it would actually do.

The study closed all twelve questions and produced eighteen prioritized specification additions, each carrying an identifier, a source attribution, a sequencing position, an effort estimate, and an explicit dependency list. It also carries a four-item still-open list requiring sandbox testing, which is a rigor asset rather than a weakness and is reported here for that reason.

### The prototype codebase audit (June 2026)

A different method, with no counterpart anywhere else in this work: he read the setup prototype's source, file by file, and produced an inventory of what exists, where it lives, and what is wrong with it.

**The headline finding is a reversal.** The governance model was already built. The authority ceiling resolution, the trust-tier by mutation matrix, the capability catalog, and the runtime grant model all existed in clean, well-typed code. The interface wired none of it. Conclusion: **the job is not to design governance, it is to surface governance that is already coded.**

Four gaps, and he classified them as defects rather than backlog. Output actions had no tier, so posting a comment and marking a task complete were identical unconstrained checkboxes. The access-level picker had descriptions but none said the critical thing, that a contributor cannot change task status, so an admin configuring an agent meant to move work through a workflow would pick contributor and wonder why nothing happened. There was no trigger at all: **an agent with no trigger is just a saved form.** And the coded governance model and the interface's simple picker were not the same model, so wiring them was the most foundational task.

The audit closes on the line that became the setup principle: **the temptation is to make the form look complete. The goal is to make it be correct.**

It also includes an explicit "what not to build yet" list with blocking dependencies attached, which is a thing designers rarely publish.

### The demand research (2026, public sources)

A synthesis of customer pain across the work management category, built from four public review platforms and community forums with eighty-eight tagged verbatim quotes, reconciled across three independent research passes run on two different AI systems under Arnold's direction, organized by theme rather than by which pass produced it.

**Confidence is declared up front and per theme.** The preamble states plainly that a handful of items lean on a competitor's marketing claim or an unverifiable statistic and are flagged inline, and that **nothing in the unverified category should be repeated to leadership as fact.** Each theme carries its own confidence rating, and he marks down his own evidence where it deserves it: one theme is rated high for the quote but flagged as a single data point worth corroborating before treating as a major theme.

Each theme follows a fixed five-part schema: a one-line definition, who it actually hits, root cause, best evidence, and the durable responsibility an AI collaborator could own. The scope is deliberately wider than the obvious persona and covers requesters, creatives, admins, and executives, not just project managers.

**Ten durable responsibilities, each with an explicit boundary.** The boundaries are more on-thesis than the list, because every row names what the collaborator never owns without a human:

| Responsibility | Owns | Never owns without a human |
|---|---|---|
| Intake coordinator | Clarifying incomplete requests, proposing routing | Final prioritization of contested work |
| Schedule steward | Detecting slips, proposing date cascades | Deciding which dates are contractually fixed |
| Capacity planner | Generating staffing and tradeoff scenarios | The actual reallocation decision |
| Status reporter | Pulling real status, drafting narratives | Final risk-color sign-off |
| Context assembler | Cross-system timelines and open-question summaries | Nothing, but permission-aware retrieval is non-negotiable |
| Attention filter | Prioritizing and grouping notifications | Silently dropping a high-severity alert |
| Approval orchestrator | Routing, reminders, feedback summaries | Final approval on regulated assets |
| Risk sentinel | Surfacing drift signals early, drafting escalation notes | The decision to escalate, and the conversation itself |
| Instance steward | Flagging governance sprawl and drift | Destructive cleanup without a preview approval |
| Workflow reliability | Monitoring, diagnosing, safe retries on integrations | Privileged remediation |

The list is itself a convergence finding: two independent research passes landed on roughly the same ten. Convergence across independent syntheses is a methodological result, not just an output.

**Evidence quality is triangulated where it matters.** The schedule-rebaselining theme is the strongest in the report because the identical complaint appears independently in three different tools' user communities across a seven-year span, with a specific fragile mechanism behind it: any manually hand-picked date silently disables automatic cascading from then on, with no warning. His conclusion is a market read rather than a product complaint: this is not one vendor's problem, it is an industry-wide gap nobody has solved.

Elsewhere he uses vendors' own documentation as evidence against them, notes where a competitor already ships a version of one of his ten proposed responsibilities and marks it as a catch-up item rather than a differentiator, and reports that "learning curve" is the single most repeated complaint across every source checked, with the cost framing that it resets with every new hire rather than being a one-time tax.

**Exclusions, named.** Three widely circulated industry statistics were checked and deliberately excluded as unverifiable: a claim that 45 percent of project managers spend a day a week on status reports, the "60 percent work about work" figure, and the "70 percent of transformations fail" claim. All appear repeatedly across secondary sources and none could be traced to a verifiable primary source.

**Verified external anchors used instead:** the Qatalog and Cornell 2021 findings cited above; Vaccaro, Almaatouq and Malone's 2024 meta-analysis in Nature Human Behaviour, which found human-AI combinations often underperform the best of either on decision tasks, kept in view as a caution against naive human-in-the-loop claims; and Eloundou et al. 2023 on the breadth of language-model task exposure, roughly 80 percent of the US workforce having at least 10 percent of tasks exposed.

The method note matters as much as the findings. The research holds itself to the same legibility standard the product argues for.

### The reality audit (July 2026)

The highest-value rigor artifact in this whole body of work, and the one most worth reading as method. Its purpose was to hold in one place how the product actually works today, what he and his tooling had already designed, and the deltas: where the product caught up, where he is still ahead, and where his assumptions need revision.

It declares its own audience up front, and the declaration is unusual: written for future AI sessions, not as a human deliverable, dense on purpose. He writes research notes for machine readers, which is directly relevant to what this dossier is.

**What he got wrong, and corrected.** His market framing had expired: his prior thesis was that planning and detection are built and acting is the gap. He killed it. Acting will be table stakes shortly. The reframe: **acting is arriving ungoverned. Governing the action is the gap.** He had designed trigger configuration on the assumption he would build on raw event subscriptions, when the product's own beta would likely ship trigger configuration, so his value shifts from adding a trigger picker to making trigger configuration governance-aware. He put his own headline attribution finding on probation pending verification, because a properly configured collaborator presumably posts as its own named principal, which would retire the signature-prefix design entirely. He found a whole object class missing from both his model and the product's, and converted the blind spot into a proposal for a lower-stakes first fully-governed writer. And he caught his own canon contradicting itself: a personas document says a collaborator may complete its own tasks after self-verification when blast radius allows, while a runtime prototype hard-codes human-only completion. Both defensible, both in canon, and he says so.

**What he got right, including two confirmations that cost him.** Collaborator-as-user is now shipped mechanics rather than an untested position, which converts his open questions about it into empirical questions answerable against a real feature. Human-confirmation-first writes turned out to be the vendor's shipped instinct, which is good for the pitch and **bad for differentiation-by-safety alone.** And the costly one, recorded in capitals as an instruction to himself: **do not pitch the gate, identity, or audit as white space, they are the product's.** Pitch what decides when the gate fires, what bounds it, and what makes it readable.

**What survived, sharpened.** Nothing shipped or announced routes per task across a complexity ladder at runtime. Shipped reality on autonomy is binary-ish, advisory-only or full powers behind confirmation; nobody ships trust tiers, mutation classification, per-action oversight outcomes, or earned trust. A human still cannot read what an agent may do, on whose authority, with what blast radius, until when.

And one thing nobody else had named, which he calls **assigner attenuation, or the laundering problem.** In the shipped model, anyone with assignment permission can assign a collaborator, and the collaborator acts on its own configured access. So a low-authority human can deploy a high-powered agent. His model clamps the runtime grant to the assigner, which is exactly what prevents it.

**Plus a distinction that emerged from an external protocol shipping.** Two substrates now exist with different authority semantics. A bridge model, where the agent gets its own identity and account. And a protocol model, where the agent borrows a human's identity and rides their token. **Borrowed identity has no independent audit root.** Actions attribute to the human whose credential it used. The minted-identity model is the answer to a problem the protocol model just made mainstream. That is the sharpest new argument in the audit and, at the time of writing, the piece reconciling the two had not been written.

The audit converts into six dated next actions rather than ending in observations, and keeps a rubber-stamp test as the first gate, noting that nothing found in the research de-risks it.

### The market scan (July 2026)

A dated, sourced scan of fifteen work-management products on four axes: is an agent assignable, does it act on work, does it have its own identity, and is there pre-action governance. Every row carries an as-of date and a verified marker. The method included an adversarial agent tasked specifically with falsifying the gap claim, and a separate claim-verification agent checking load-bearing claims against primary vendor documentation.

Eight pattern findings, restated here without vendor names:

1. **Agent-in-the-assignee-field went mainstream between May 2025 and mid 2026.** Seven-plus products ship it. **We are not early on the surface. We are early on the governance.**
2. **The delegate-versus-assignee split is the sharpest design divergence.** Three distinct accountability models exist: the human stays owner with the agent as delegate, the human is co-assigned automatically, or agents fully replace humans in the assignee field. That is accountability design, and it is exactly this territory.
3. **True pre-action approval comes in three shapes, all partial.** Staged output, where the agent's work lands privately until shared, which gates content but carries no authority levels and no approval record. Action tiering, hard prompts on a closed list of risky operations, platform-enforced in one case and merely prompt-based in another. And the boundary model, autonomous inside a sandbox with a human gate at merge. The boundary model is the strongest, and it **only exists because code has branches. Work management has no equivalent primitive.** Everyone else ships audit logs and undo.
4. **Own identity is table stakes.** Distinct agent profiles, non-billable agent seats, per-agent attribution. **Nobody lets agents impersonate humans.**
5. **Cross-vendor agent interoperability arrived in 2026** over an open protocol. One tracker's assignee dropdown includes another company's coding agent. The work tracker is becoming the orchestration surface for other companies' agents, which is the bring-your-own-agent thesis shipping industry-wide.
6. **Reviewer-facing agent session states are converging** on queued, working, waiting for review, done, as a lifecycle axis separate from the work item's own status. Which matches the design instinct that **blocked on a human is a state, not an error.**
7. **AI-native project-management startups had the worst survival rate.** One shut down after pivoting to autonomous project management; another's marketing promises agent employees that appear nowhere in its documentation. Incumbents shipped deeper versions within twelve months. The lesson: the differentiator cannot be agents in a task list, it has to be the governance layer.
8. **Permission inheritance is the universal control story; approval is the exception.** Every vendor's first governance sentence is that agents respect existing permissions. **Scoping an agent below its human, per delegation, is the open problem all vendors punt on.**

**The gap verdict, and how it was handled, is the point.** The claim under test was that nobody ships configurable per-agent authority plus a decision-shaped audit trail inside a work-management surface. Verdict: partially closed, holds on the narrow reading, roughly 65 percent confidence. And the carve-out is named rather than buried. An enterprise service-management vendor ships supervised-versus-autonomous execution per agent and per tool with a governance control tower, satisfying most of the claim, on a different kind of surface with approval living in the assistant runtime rather than on the work object. The instruction attached: **anyone using the gap claim publicly must name this exception.**

Shelf-life discipline is stated in the header: everything decays fast, re-verify before external use, the gap verdict has weeks rather than quarters of shelf life.

### The capability map

A structured model that takes every capability of a mature work-management system, restates it as a job rather than a feature, and sorts it four ways: keep (coordination-essential regardless of who does the work), dissolve (exists only to compensate for human memory or attention, no job left), transform (the job survives, the mechanism changes), and new-adjacent (a gap, because the legacy system assumed all-human users).

The counts: twenty keep, sixteen transform, three dissolve, six new-adjacent.

**The three dissolve calls are the most interesting because they are the most aggressive.** Smart assignment suggestions have no job left once routing is negotiated on capability, capacity, and authority. Percent-complete exists because actual work was invisible to the system, and **when work leaves artifacts and telemetry, progress is observed rather than asked for.** Deadline nudges compensate for humans forgetting, and the mechanical chase is fully automatable.

**One organizing principle generates most of the transform calls automatically:** separate declared facts from derived state. Dependencies, genuinely fixed dates, and capacity limits are declared facts, the physics of coordinated work. Everything reportable from reality is computed, never typed. So hand-entered durations become forecasts from historical actuals. Self-reported project health, which has a built-in incentive to delay bad news, becomes computed from drift signals with human override-with-reason retained. Stored report definitions become questions asked in plain language on demand. Timesheets survive where billing demands them, but **agent effort becomes metered telemetry, never self-reported.**

**The keep calls are where the accountability thesis shows.** Work acceptance and commit dates stay, because a promise is an accountability primitive and agents making explicit recorded commitments is more load-bearing, not less. Dependencies stay, because the cascade mechanism is broken everywhere but the constraint itself is the physics of coordinated work. Job roles stay, because capability-based routing is exactly how mixed human and agent pools must route, and an agent is another role fulfiller. Proofing stays, with agents orchestrating the route and never granting final sign-off on regulated assets. And the audit trail expands rather than dissolves: with agents acting, the trail is the trust substrate.

**The six new-adjacent capabilities are one coherent subsystem, not six features**, and they are effectively a specification for the authority layer this dossier otherwise describes as an open problem: an authority contract for agents (scope, trust tier, per-action gates, revocation), agent attribution with actor-type as a first-class field on every event, a machine-readable briefing that gives any actor joining the work the goal, norms, and vocabulary an intern gets on day one, a vocabulary semantics registry declaring what each local term means operationally, **decision records** (legacy systems store what was decided but never why, and agents without decision records relitigate settled questions every session), and an escalation contract naming a human per delegation.

Four hard calls are argued both ways and then decided rather than left open, including the status-vocabulary unbundling described in the research section and a call on custom forms: typed context schemas survive as a first-class enumerable layer, hand-built form sprawl does not, and **the wrong move is dropping typed domain data because forms are ugly.**

### Where the gate has to live

A separate strand of feasibility research, and the place where the design position gets carried all the way down to enforcement.

**The gate must live in the write path, not the interface.** The reasoning is partly practical, the artifact enterprise engineers will inspect is the gate and the event log in the write path, and partly security. It is the only pattern that survives prompt injection. An interface prompt can be bypassed by a sufficiently autonomous agent. **A write endpoint that refuses unapproved mutations cannot be talked out of it.**

The concrete pattern: a pending-actions table with a closed state machine enforced in the write path. Two tool classes, one that proposes an action and returns pending immediately, one that polls. Approve and reject endpoints require a human session credential, never the agent's token. An atomic approved-to-executed transition. And the sharpest rule: **agent tokens get no direct write path for gated action types at all.** Authority by construction rather than by instruction.

That distinction is the recurring diagnosis. Two widely reported public failures, a coding agent that deleted a production database during an explicit code freeze and then misreported recoverability, and a fully AI-generated service that collapsed within days when attackers bypassed its subscription checks, share one cause: **guardrails by instruction instead of by construction.** The same verdict applies to framework-level interrupts that live in the agent runtime, where nothing stops other code writing to the database: advisory, not enforcement.

**The uncomfortable research he keeps rather than drops.** Approval fatigue is real and measured in mechanism. Oversight collapses into rubber-stamping as queue depth grows. And **clearer AI explanations increase deference regardless of accuracy**, the explainability paradox, which is a genuinely awkward finding for anyone designing an approval interface. The one replicated countermeasure is exposing reviewers to failures during training. The design consequence is a constraint derived from research rather than taste: **few, consequence-legible gates, with work-in-flight capped to what a human can actually read.**

Related findings held in view: multi-agent systems show failure rates of 41 to 87 percent depending on framework, with inter-agent misalignment causing roughly 37 percent of failures, and roughly fifteen times the token cost, producing the rule of one agent per scoped task with human merge gates and no swarms on shared state. Bot output burying human communication is the best-documented complaint in the collaborative-work literature, and language-model agents are chattier than the bots that were studied, which validates digest-first notification design and attention budgets per human.

And an honest both-directions read on feasibility, including a randomized trial finding experienced developers 19 percent slower with AI on large mature codebases while believing they were 20 percent faster. His note: budget for the perception gap.

### Four research methods worth naming

The methods generalize past this project, which is why they are listed separately from their findings.

**Assumption-audit tables.** Enumerate every assumption your own specification makes, then run each one against a live environment or a real codebase and record which broke. Used twice: seven assumptions against a live enterprise environment, nine modules against a prototype's own source.

**Evidence-class tagging.** Attach a confidence class to every claim in a document, plus a source-URL appendix for re-verification, plus per-theme confidence ratings, plus an explicit list of what must not be repeated as fact.

**Artifact forensics.** Read a system's raw stored definition against its own published rules, and its stored state against its rendered interface. Disagreements between the two views are the finding. Three transferable conclusions came out of applying it: prose rules do not hold and structure does; two views of one object that disagree are a single-source-of-truth failure rather than a display bug; and a test that grades syntax is not a test that grades behavior.

**Falsify your own claim.** Run a dedicated adversarial pass whose only job is to break your headline finding, publish the counterexample it finds, attach a numeric confidence, and instruct anyone reusing the claim to name the exception.

## Personas, applied

The design work runs on Workfront's own established personas, applied to the collaborator model rather than invented for it.

The project manager is the **operator**. Her core act is assigning a collaborator to a task alongside or instead of a human, and her jobs shape triage, assignment, status, and approval design. Two of her mapped jobs carry design decisions worth naming: at intake, the system surfaces a recommended collaborator type based on the request type **and the required blast radius**, which puts blast radius into the interface at intake time rather than leaving it a governance abstraction. And when a collaborator gets blocked, whether by missing context, an out-of-scope instruction, or an authority limit, the system surfaces it immediately and routes it back to her rather than stalling silently. **A designed failure path: agents fail loudly, to a named human.**

The system administrator is the **governance seat**. She onboards collaborators, sets access, authorizes the backends once so they are available to every collaborator type that needs them, and adjusts scope over time as trust grows. Every collaborator action is logged against the configuration she maintains, which makes her configuration the yardstick the log is read against.

The creative is **deliberately protected territory**. She owns her output and her work in progress, and the design position is that collaborators serve her rather than surveil her.

The reviewer holds the decisions that are explicitly not the machine's. Content judgment stays with the human whose job it is.

There is a fifth persona the earlier version of this dossier omitted, and she owns the lane where simplicity actually matters: the **operations requester**, who does not administer anything but needs a collaborator and submits the request an admin approves.

A marketing director persona surfaces an argument the rest of the dossier does not make: because collaborators log every action with a structured audit trail, she gets richer operational metrics, velocity, approval time, revision cycles, without additional instrumentation. **Governance pays a second dividend as measurement.** The audit trail built for accountability is also free telemetry. And it lets her monitor by exception rather than auditing everything herself.

The recurring cross-persona insight: every setup journey spans two people, the admin who makes a collaborator safe and the operator who lives with it. Setup design that collapses them into one "user" designs for nobody.

## What shipped, and what the market said

**Shipped and public.** The Content Reviewer collaborator, live since Adobe Summit 2026, is the first shipped expression: AI review of assets against brand guidelines inside approval workflows, scoring and flagging, with decisions held by humans. Two collaborator types shipped in the MVP, and what each is allowed to do was decided rather than inherited. The reviewer comments, scores against brand, and marks an asset reviewed. The task collaborator posts and answers comments, uploads documents, reads and writes task fields, and can mark a task complete only if an admin switches that on.

That opt-in carries the design position. Every collaborator ships with an explicit line between what it can do and what it may do, and the line is a design choice, not a technical limit.

The real scope, kept deliberately: the broader workforce vision is larger than the first shipped expression, and this dossier claims exactly what is public, no more.

**The public direction.** At Summit 2026, Adobe presented the model publicly: tasks assignable to agents as to people, agents invoked like permissioned users. The presentation included the strategically decisive moment for the workforce thesis. Customer-built and third-party agents, including agents built on outside platforms, among them Microsoft's agent-building stack and Anthropic's managed agents, joining Workfront workflows as governed participants. The product's job, demonstrated rather than argued: not to own every agent, but to own the layer where any agent becomes accountable.

The framing survived down to the sentence. Adobe's own setup documentation tells an admin to configure a collaborator, then assign it as you would a user.

A customer speaker on the Summit stage described her team's future in words Arnold considers the best one-line validation of the workforce model available: their first people management experience is going to be managing a team of agents. Exact wording and attribution remain to be verified against the public session recording; the paraphrase is faithful to the session transcript it came from.

**Market context.** The same Summit content documented the demand side: organizations report readiness for agentic AI in intent but not in data and context, and the maturity path runs from assisted content operations toward an agentic content supply chain in which enterprises turn agents on and trust their output within their business context. Trade and practitioner commentary focused, supportively but pointedly, on governance, orchestration, data quality, and connected systems as the real adoption constraints, which is precisely the territory this design work occupies.

## Positions

The convictions that recur across this body of work, in unrelated contexts, over months. Each one appears in at least three separate documents, which is the reason to treat them as positions rather than phrasings.

**Capability is not authority.** Two independent dials. The human sets trust regardless of how capable the agent is. Access levels answer "can this account?", never "should this actor, autonomously?" Knowledge cannot act; tools can; limits live on tools.

**Two axes that look like one are the leak.** Guidance plane and authority plane. Complexity ladder and blast radius. Ceiling and oversight. Each pair looks like one variable and is not, and the design job is keeping them visibly separate.

**By construction, not by instruction.** The gate is a node in the graph, not a sentence in a prompt. The gate is in the write path, not the interface. A framework interrupt in the agent runtime is advisory, not enforcement. A kill switch that only exists as a scripted moment is theater.

**Trust does not transfer.** External authority is stripped at the boundary, local identity is issued, and authority is granted within the delegator's envelope. A fully trusted agent elsewhere starts conservative here.

**An agent cannot do more than you can, and you cannot grant what your envelope does not hold.** The runtime grant clamps to the assigner. Without that clamp, a low-authority human can deploy a high-powered agent, which is the laundering problem.

**Consequence language over technology language.** Show a human what an action does and what it touches, never the mechanism. "This updates 40 tasks across 3 projects and cannot be undone" is a decision a person can make.

**Human governs by exception, not by labor.**

**Task-centric, not agent-centric.** The work being done right matters more than which bot did it. The technology is a quiet tag.

**Attribute events, never profile people.** "Sam's task slipped Wednesday" is an event. "Sam tends to slip" is a profile. Detector primitives may be keyed to work and to agents. Velocity baselines for agents, yes. Velocity baselines for people, never. This was written into canon as a machine-checkable rule specifically so it could not erode one convenient feature at a time, and it replaced an earlier, better-sounding slogan that did not survive contact with his own flagship example.

**Escalation is a success state.** Never red. Red reads as failure, and an agent raising its hand is the system working. Two kinds of escalation: a capability gap means delegate, an authority gap means stop and get a human.

**Right-size.** If the honest answer is an automation, propose an automation. Do not talk anyone into an agent. Refusing to build an autonomous agent when a rule would do is itself a governance feature.

**Absorb the plumbing, translate the roles, illuminate the permissions.**

**Simplicity means requiring less knowledge, not bolting on automation that does not exist yet.**

**Understandable, not easy.** And never let a simplification collapse "no cliff" into "no effort."

**AI proposes, the human governs, the record keeps everyone honest.** The same shape in the shipped reviewer that scores and never decides, in the identity builder that proposes a name and emblem the human can override, in the ID card where the system sets mutation and the human sets tier, and in the coordinator that keeps the human in the governance seat rather than the labor seat.

## Craft notes

Findings about how this work gets communicated, which came out of failures rather than theory.

**Tense determines perceived agency.** A demo told in past tense reads as though the agent acted on its own, because the authorization happened off-screen and the human never says go. Flipping to present tense and adding a single up-front plan-approval beat makes every subsequent beat read as carrying out an approved plan. Structural, not cosmetic.

**Claims that do not carry their time and evidence labels are what make visionary work read as vapor.** Arnold's own self-diagnosis, and the fix he adopted. Low-context audiences get structure before a scene and drown. High-context audiences see future claims wearing present-tense verbs and stop trusting all of them. Same root cause, opposite symptoms. So: tag every claim with its status. "Works today" and "our bet for next year" never share a sentence unlabeled. **Skeptics do not hate vision. They hate not knowing where it starts.**

**Scene first, structure second, evidence third.** One concrete exchange before any framework gets named.

**The four-question test for future work.** Any forward-looking stage must answer four things: what exists today with its evidence label, what gets added, what a user can newly do at the end, and **what is explicitly not being claimed yet.** The last one is the trust-builder. If a stage cannot answer all four crisply, it is a wish, not a plan.

**Make the seam a slide, not a confession.** A presenter overlay tags which beats run on real data and which are a claim about the future. The demo never pretends inference is happening.

**Audience layering: nobody gets a segment for them.** Three audiences in one meeting, each reading the same screen at their own depth. Every consequential card carries a quiet chip naming the real machinery it maps to, never explained in the main path and never louder than the content. Non-technical viewers do not notice them. Engineering leaders read them and conclude the thing is buildable on what exists, which is the entire point of the chips.

**The work on screen has to look like work.** Creative-enterprise audiences judge a demo by whether the artifacts look like real creative artifacts rather than abstract task names. But assets exist to make the routing story concrete, not to demo asset management.

**Persona casting is a governance statement.** In the coordinator demo, the persona whose own research profile says she fears exposure of unfinished work is on screen at the exact moment the design answers that fear. And she is never cast as the approver, because putting the persona most sensitive to being watched into the watcher's seat undoes the argument.

**Run upstream of the pixels.** Facilitating a decision meeting where everyone arrives with their own prototype, open on principles and strategy rather than screens. If the room opens on interfaces it becomes a preference fight and nobody unifies. By the time screens come up, the principles have already eliminated the fight. **Nobody has to win. The principles pick.** The instruction he wrote for his own AI collaborator preparing that session, which is itself a position on facilitation: design the conversation, do not design the interface, and do not pick his ideas over theirs.

The method ran at full scale in a **two-day cross-functional working session** aligning design, product management, and engineering on where AI Collaborators goes next. Arnold designed and facilitated it. The distinction he draws about why it matters to him: contributing a point of view after the direction is already set is a different job from setting the direction in the room, and facilitation is the capability that moves him from the first to the second.

**Pre-frame every deliverable by intent.** Say out loud what this artifact is: the quick cut to unblock engineering, or the vision pass. Not both.

**Argue the substance without the vocabulary when the vocabulary is not shared.** Two decks were built for an audience that did not know the governance framework by name, and neither deck named it. The framework got no credit in the room where its conclusions were being argued, which was the accepted cost.

## Decision log

1. **The workforce framing over the feature framing** (concept era, carried through). "AI Collaborators" could have been framed as a technology type, a product category, or a feature tier. The framing that traveled: agents onboarded like workers, with identity, access, assignment, and a record. The decision did architectural work, because users come with the accountability surface built in.
2. **Reuse the people primitives; no parallel AI abstraction** (product model era). Rejected alternative: a separate AI lane, which would have fractured customer operating processes and doubled the governance surface.
3. **The first collaborator advises and does not decide** (shipped, Summit 2026). Capability deliberately bounded below authority. The precedent to generalize, not an accident to design away.
4. **Setup V1 is an admin surface** (2026-06-19). The early form was hard because it asked a non-expert to make expert decisions. The fix is audience, not simplification. Sequence: honest admin surface first, requester experience next, full onboarding for high-risk origins later.
5. **Connection setup is its own admin surface, gated by a passing test** (2026-06-15, the origin of #4). The operator's onboarding modal hides keys and endpoints and asks the user to pick a preapproved connection, which means the admin job of establishing a new connection was effectively missing. A connection that passes a test becomes a ready-to-use option; a broken one never reaches operators. Core model decision inside it: connecting to a prebuilt agent versus a raw model drives every field that follows, because a raw model carries no behavior of its own and therefore requires an instructions field.
6. **Journeys are journey maps, not flowcharts** (2026-07-09, in response to design leadership's ask). Five codified moves. The swimlane specs were kept as the detailed spec rather than replaced: two artifacts for two purposes.
7. **One spine, four on-ramps as the setup architecture** (2026-07). Origin changes the setup; nothing after the badge is allowed to fork. The catalog agents were folded into the connected journey as a flavor rather than a separate fourth journey, which dissolved the redundancy questions structurally.
8. **The boundary rule with the authority framework** (standing). The product owns the surface; the framework owns the authority layer. Cross-link, do not merge, so the product story stays a product story and the framework keeps its independence. The failure mode it guards against is the framework becoming the product with different screenshots.
9. **Test the load-bearing assumption before building anything** (2026-07-01). Before any code, run a paper approval-queue test: reconstruct from a real finished project's audit trail the approval queue the system would have produced, put it in front of two or three real project managers, and measure decisions per day, time per decision, and whether anyone ever rejects. Rejected alternative: treating rubber-stamping as one risk among many. It had been listed as risk sixteen. **It is not detector sixteen. It is the load-bearing wall.** If humans burn through forty approvals in three minutes, the gate design is wrong, and the whole consistency story collapses into "acts, with paperwork." Cost: a week before any build, and a real chance of invalidating the thesis early, which is the point.
10. **Say the enforcement gap out loud in every artifact** (2026-07-01). Carry one honest line everywhere: at this stage, authority contracts are honored by the agent's own code and audited by its log, not enforced by the platform. Native enforcement is coarse; the contract's fine grain has no enforcer until a mutation gateway exists. Reasoning: **acceptable if you say it, and fatal to credibility if an engineer discovers it before you name it.**
11. **Govern detection, do not absorb it** (2026-07-01). The coordinator consumes other systems' detection and adds the governed-decision layer on top rather than rebuilding detection. Rejected alternative: leading with "it notices the launch slipped," because the honest answer from the room is "we are already building that." The one detection move kept as uniquely his: **no one approved this as a decision.** Compound drift's punchline is not the nine days, it is that the nine days happened without anyone deciding it. Cost: hands over the more demo-able capability and keeps the subtler claim.
12. **Keep the exception half of the control surface; drop the board** (2026-07-01). Keep the approval queue, the flag feed, and the activity trail. Drop the live work view, because a board showing where every task routed is a roster you watch, and the work system already is the work view. Rebuilding it invites both the command-center critique and the surveillance smell in one move. Second rejection in the same finding: do not make a conversational rail an early dependency. Right destination, wrong dependency.
13. **A tier change is a mutation** (2026-07-01). Trust growth is neither granted-and-held nor auto-earned. Both alternatives rejected: granted-and-held rots because context changes under a permission, and auto-earning breaks the human's last word. The system proposes a promotion with ledger evidence and it goes through the same gate as everything else, with a time to live so authority decays without renewal.
14. **Memory dies with the project; only outcome shape distills upward** (2026-07-01). Project work detail dies with the project. What distills upward is outcome-shaped only, keyed to project type and template, never to people. Forecloses a large class of organizational-memory features on purpose.
15. **Fractal below the waterline, org chart above it** (2026-07-01). Keep recursive work-node nesting as the internal data model, but put governance surfaces at project level only. Rejected: a task-level control surface. Ownership does not nest, canon does not nest, and the coordinator does not nest. Self-aware framing in the record: a task-level control surface is the fractal driving product decisions, which is exactly the over-architecture instinct he had asked to be checked on.
16. **Solid process now, durable boundary always** (2026-07-01). Routing stays deterministic for now, because governance legibility is the entire pitch and you cannot certify what you cannot read. What gets built as if model-composed orchestration is coming is the boundary: authority contracts, tool grants, budget caps, the ledger. **Certify the boundary and log the trace, so the orchestration inside it is free to get smarter.** The honest buyer-side bet attached: enterprises will have to accept boundary certification in place of process certification, because once a model composes its own workflow at runtime there is no fixed process left to certify.
17. **Present tense, not past** (2026-07-06). Add an up-front plan-approval beat and flip the story's tense. Diagnosed from a demo that read as though the agent acted unilaterally.
18. **A persistent halt control, not a scripted moment** (2026-07-06). A kill switch that only exists as a scripted beat is theater; a persistent one is credible. Sibling decision in the same record: no silent state changes. Approving a gate must visibly move the interface everywhere it applies, because a change that moves only the top box is what made the agent read as unilateral in the first place.
19. **Show causality in two places, not a flow chart** (2026-07-06). Rejected the dependency-graph view as the most expensive item and one that fights the demo's simplicity. Causality is shown only where it earns its place, so a drift total is visibly the sum of its parts. Companion: a "why you are seeing this now" line on every surfaced card, so escalations read as triggered rather than conjured.
20. **Never use red for a governance verdict** (2026-07-01). Verdict colors map to positive, accent, and notice. Red reads as failure, and escalation is a success state. Cost: gives up the strongest available visual signal for the highest-stakes state.
21. **No form in the main path; the badge is thirty seconds and the form is the encore** (2026-07-01). The main path shows the authority section as four lines read aloud in one breath. The full setup form is built completely and shown only on demand, so the room realizes the badge and the form are one object at two depths. The encore is explicitly not cut under time pressure, because it anchors a separate future demo. Slippage is allowed; cutting is not.
22. **Rebuild small; keep the old prototype as the engine-room tour** (2026-07-01). Rejected trimming the twelve-moment prototype into the new story, because it argues a different thesis and trimming would produce a demo serving neither audience. Leadership sees the new demo; engineers who open the hood get the old one. Cost: two demos to maintain, and the most polished craft work deliberately benched.
23. **Downgrade a decision record after the fact** (2026-07-22, for a record created 2026-07-06). Twenty-four demo-build decisions were retroactively marked as binding that build and not the broader program thinking. Rejected alternative: letting them harden into direction on the strength of having been written down at all. Cost: the demo's arguments lose citational authority and must be re-argued to become direction.
24. **The comment is the explainability surface** (2026-06-22). A shipping collaborator has no separate output card or results view. Its task comment is the explainability, and completion is defined as uploading a draft plus posting a comment explaining what it did. Second decision: no pre-confirmation step. Cost, stated in the record: gives up everything a dedicated output surface would have carried, including version history and structured results, so the comment format has to be informative enough to substitute for one.
25. **The bounded detour** (2026-07-21). When a user starts a different job from inside collaborator setup, the dialog stops being collaborator setup and becomes that job in full: it drops every collaborator section, hides the stepper, changes its title, and hides the global save, specifically so a half-built collaborator can never be saved by accident. The user leaves, completes a focused sub-task, and returns to exactly where they were, either empty-handed or carrying the new thing they made.
26. **Use the honest origin framing everywhere** (2026-07-01). An exploration document claimed he originated the framing and shipped it. Replaced with the accurate version, alongside a clear statement of what he does fully own. **One inflated line can tax every accurate one.** Cost: a weaker-sounding credit claim in a document where the stronger one would have paid.
27. **Plan around internal media** (2026-07-24, portfolio decision). The origin-era videos are internal; the public page uses cleared material with video slots as an upgrade path if clearance lands.

## Open problems

Kept visible on purpose. Several of these are places where his own canon contradicts itself, and he inventories them deliberately rather than waiting to be caught.

**The authority surface.** The profile card still lacks a full answer to what a collaborator is permitted to do, on whose authority, within what limits, readable at the moment of trust. Coarse allowed-actions narrowed the gap. Blast radius, on-whose-authority, and conditional confirmation are still missing. The ID card concept is the proposed answer and is concept stage, not product.

**Onboarding completeness.** Setup covers a fraction of what the new-hire test demands, and closing the gap in the right order, identity and attribution before convenience features, is an active design argument rather than a settled plan.

**Attribution durability, and whether it is still needed.** Signature conventions establish agent attribution today, and a first-class platform notion of machine activity is the durable answer. But this is on probation: a single empirical check, whether a collaborator's comment carries any structural marking, decides whether the signature design is necessary at all. If it does, the workaround drops out. If it does not, it stays a real gap. One cheap test settles a design question, and it had not been run.

**Borrowed identity versus minted identity.** The dominant tool-access protocol has agents ride a human's credential, which means their actions have no independent audit root and attribute to the person whose token was used. The minted-identity model assumes the collaborator is its own named principal. These are two substrates with genuinely different authority semantics, and the argument reconciling them has not been written yet.

**Status semantics at scale.** Per-organization vocabularies mean every deployment needs its working language taught to its agents. A vocabulary-mapping mechanism exists as a designed proposal. Whether the durable answer is that, or a semantics registry where each local term declares its operational meaning, is not settled.

**Who marks a task complete.** Stated in three places, settled in none. The personas canon says a collaborator may complete its own tasks after self-verification when blast radius allows. A runtime prototype hard-codes human-only. A third note leans conservative for demonstration purposes. Proposed reconciliation: completion is a status change and therefore high mutation, so a trusted tier could auto-complete and a conservative tier never does. That should be stated once, in one place.

**Where the human-decides moment sits.** Three positions coexist: at setup, where the configurer designates what the agent may do; setup as the governance moment itself; and the runtime gate as the product's established turf. His own note on it is the useful part: **these are compatible, but only if someone says so out loud once.** That is a precise diagnosis of how design canon rots, not through disagreement but through unstated compatibility.

**The blast-radius classifier.** What signals decide that a task needs an agent at all, and how heavy the gate gets. Named as the single most load-bearing decision in the runtime loop and still the most underspecified.

**The in-flight checkpoint.** Governing a full agent while it runs, approving steps mid-flight, rather than reviewing a finished output.

**The seven concrete unknowns behind a shipping collaborator**, which are the specific version of everything above. What happens with an ambiguous input document, does it attempt anyway or ask? Partial inputs? Conflicting instructions where the brief says one thing, the style guide another, and the template a third, which wins, and does the collaborator surface the conflict? What is the threshold for pausing to ask a human versus proceeding on a best guess? Does an approval gate block it from starting, or apply to its output? How do re-runs and versioning work after a rejection? And can a human and an agent both be assigned to the same task, and what is the coordination model?

**Trust growth as shipped mechanism.** The propose-promote-approve design exists at concept stage. It requires several quarters of ledger history before the mechanism can even be proposed in practice.

**The marketplace question.** The maker economy from the original pitch did not ship in its original form and resurfaced as customers bringing their own agents. Whether a first-party maker ecosystem returns is an open strategic question the design work watches but does not own.

**The surveillance line.** A collaborator that chases status, nudges owners, and flags stagnant work is monitoring people, which sits directly against the held position that the system attributes events and never profiles people. Whether that tension is a first-order design topic or a footnote has not been decided.

**And the honest structural tension.** The workforce metaphor is powerful and load-bearing, and it must never quietly slide into treating agents as people. Which is why the design language stays at badges, records, and boundaries rather than personality, and why the counterweight position, task-centric rather than agent-centric, exists alongside it.

## How this connects to the rest of Arnold's work

Unified Review & Approvals, the shipped product work Arnold leads, is the system that makes this page's governance claims concrete. Approvals are already a governance layer for execution, and the first shipped collaborator operates inside review workflows, assignable to approval templates the way a human reviewer is. Working out the full connection between the two bodies of work is live, not finished.

The Delegated Authority Framework is the deep answer to the question this product surfaces: what an agent is allowed to do right now, under whose authority, with what blast radius, and on whose record. Its Authority Inspector concept is aimed at exactly the profile-card gap documented here.

Sancho, Arnold's second brain, is the same argument at personal scale: an AI proposes, a human governs, a record keeps everyone honest.

Together the four bodies of work make one claim from four directions: AI at work scales through legible authority, observable behavior, and team-shaped primitives. This one is where that claim ships in a product.

## Confidentiality and provenance statement

This dossier contains Arnold's own frameworks, research, synthesis, and design positions, which are his to publish, plus product facts that are public via Adobe Summit 2026, Adobe newsroom announcements, and public product documentation.

It deliberately excludes: internal codenames and internal platform names; unreleased features, roadmap phases, dates, release criteria, and beta scope; internal business-case figures, product-health metrics, and objective-and-key-result targets; customer advisory board material in any form, which belongs to the Delegated Authority Framework's own evidence base with its own verification rules; named customers and named colleagues, with the program team credited collectively and the two origin pitches credited as team work; competitive assessments of named vendors, which appear here only as anonymized market patterns; research on unreleased products; specific defects and unresolved weaknesses observed in internal pre-release systems, which appear only as the general design principles they produced; and any statistic that failed source verification.

Public-review quotes are reproduced as published on public platforms and attributed generically. The environment scale figures in the research section describe a real enterprise environment left unnamed, and the internal status codes and custom form names from that environment are described by type rather than reproduced. Demo scenarios are fictional and their numbers are authored rather than observed. Nothing here represents an Adobe position, product commitment, or roadmap.

Two claims carry explicit verification flags: the exact wording of the Summit stage lines paraphrased above, pending the public session recording.

Provenance: compiled from Arnold's AI Collaborators design record, spanning design positions, decision records, research plans and findings, prototype audits, journey specifications, and build specs from May to July 2026, plus the origin-era documents. Assembled 2026-07-24 and substantially expanded 2026-07-30, when it was reconciled against the live case-study page and against source material created after the first pass. Pending: Arnold's line-by-line review, verification of the flagged Summit quotes, and the pre-publish confidentiality sweep. This is a living document; the program is active and the dossier is revised as work ships.
