Agentic Delivery Field Notes
The Bottleneck Moves
How agents, skills, owner decisions, and scarce capacity determine product, project, and program outcomes.
A practical framework for selecting the most valuable feasible combination of work, agents, skills, sequencing, and delegated authority under limited capacity and uncertain evidence.
Central thesisThe goal is the best feasible combination of work, capabilities, sequencing, and delegated authority that produces the greatest verified outcome under current constraints—and can be revised when evidence changes.
Capability testEvery agent, skill, review, and approval should expand the feasible set, reduce total cost or risk, or improve decision evidence. Otherwise, it is overhead.
Terms in use
Keep the evidence chain visible
- Outcome
- A meaningful change experienced by a customer, beneficiary, operation, or organization.
- Output
- An artifact, feature, report, decision, or completed task produced by delivery activity.
- Accepted deliverable
- An output that meets explicit acceptance conditions and is usable by its intended recipient or downstream process.
- Realized benefit
- An observed improvement attributable, with appropriate qualification, to delivered change and adoption.
- Constraint
- A scarce resource, dependency, rule, or capacity limit governing feasible choices or current throughput.
- Bottleneck
- The current constraint most directly limiting accepted outcomes or benefit realization.
- Constraint migration
- Movement of the governing bottleneck after another part of the system is improved.
- Owner
- The accountable human responsible for consequential choices, reserved matters, and acceptance at the appropriate level.
- Best feasible
- The strongest permissible alternative identified under stated objectives, evidence, alternatives, constraints, and decision window.
The operating lesson
The bottleneck moved
Agents can increase execution, skills can prevent procedural reinvention, and owner approvals can preserve control. Each choice may be sound alone. Together, they can move the bottleneck: more work reaches specialist review, integration, evidence validation, owner decisions, or adoption faster than those stages can absorb it. Activity rises without a matching rise in accepted outcomes.
This is a generalized pattern, not a measured program result. Improving one component does not guarantee improvement in the constrained system. If execution falls while specialist or owner waiting rises, the intervention moved delay rather than improving end-to-end delivery.
The first Field Note argued that more agents do not automatically create more accepted delivery. The consequence is broader: even a simpler agent configuration helps only when the delivery system identifies and manages its current constraint.
The objective
Optimize accepted and realized outcomes—not activity
Completed tasks, features, projects, documents, agent calls, and skill reuse describe activity. They do not prove that an output met acceptance conditions, became usable, changed behavior, or created benefit. Verification can turn an output into an accepted deliverable; adoption and operating conditions influence whether its benefit is realized.
The management question is combinatorial: with limited people, expertise, money, time, review, owner attention, and compute, which combination of work and capabilities creates the greatest verified value? The highest-value initiative alone may block a stronger combination. Maximum utilization is not maximum value.
Deliberate slack leaves room for discovery, interruption, rework, risk response, owner decisions, and integration failure. A plan that consumes every modeled hour may be tidy and operationally fragile.
EXTERNAL EVIDENCE The Scrum.org Evidence-Based Management Guide distinguishes inputs, activities, outputs, outcomes, and impacts, and advocates adapting through evidence under uncertainty. That taxonomy supports the distinction here; it does not validate a particular allocation.
Decision language
Say what “optimal” actually means
Model optimality asks whether a plan is best within the encoded objective, alternatives, constraints, assumptions, and decision window. Decision quality asks whether those inputs and the authority to choose were defensible at the time. Realized outcome asks what followed implementation and adoption. These are different judgments.
A model-optimal plan can fail when an interaction is omitted or an estimate is poor. A sound decision can have a poor outcome; luck does not prove good reasoning. Feasibility merely satisfies stated constraints, and model optimality does not establish real-world fit.
Use evidence-bounded adaptive optimality: choose the best feasible action under current evidence and constraints, state assumptions and solution status, and define what change requires reconsideration. This is a disciplined claim about a decision boundary, not a promise of the best real-world result.
EXTERNAL EVIDENCE Google OR-Tools documents separate FEASIBLE and OPTIMAL statuses: a feasible solution may exist without established optimality. This technical distinction is used here as a truthfulness rule for management language.
What optimality does—and does not—mean
Solver status language is useful discipline for management claims: “feasible” and “optimal” are not interchangeable.
- Local vs. global
- A locally improved component can leave the whole system unchanged or worse.
- Subsystem vs. system
- Utilization inside one function is not end-to-end outcome optimization.
- Constrained vs. unconstrained
- A recommendation is meaningful only with its binding limits.
- Pareto-efficient
- No objective can improve without worsening at least one other objective; preference may still be required.
- Second-best reasoning
- When an ideal option is infeasible, optimize inside the constraints that actually remain.
- Feasible vs. optimal
- Feasible satisfies the model; optimal is best only when that status has been established for the encoded model.
The complete system
Connect investment choice, value, delivery, capability, and accountability
Portfolio management chooses investments; product management defines and tests value; program management coordinates benefits; project management organizes delivery; agents perform bounded adaptive work; skills provide reusable capabilities; accountable owners resolve consequential trade-offs.
These are complementary contributions, not mandatory departments or exclusive boundaries. One person may perform several. The distinctions matter because excellent execution of the wrong project combination does not create a good portfolio, and project completion does not prove program benefit.
Preserve both a delivery path and an accountability path. An agent may prepare evidence or execute an authorized decision, but tool access does not transfer human accountability for material spending, publication, security, legal commitments, or other reserved matters.
EXTERNAL EVIDENCE The PMI Lexicon provides standardized project, program, and portfolio terms. The contribution map below is this article’s synthesis, not a claim that every organization needs these as separate departments.
| Element | Primary contribution |
|---|---|
| Portfolio management | Selects, balances, defers, and stops investments under scarce resources. |
| Product management | Defines, tests, and steers toward customer and product value. |
| Program management | Coordinates related changes and dependencies needed to realize benefits. |
| Project management | Organizes a bounded change across scope, schedule, cost, quality, and risk. |
| Agents | Perform bounded adaptive execution or delegated decisions. |
| Skills | Supply reusable, validated procedural capability. |
| Accountable owner | Resolves reserved trade-offs and accepts consequential outcomes. |
Constraint inventory
Find what governs throughput now
Look beyond engineering. Discovery, expertise, security or legal review, research access, quality assurance, integration, readiness, adoption, data, compute, money, decision rights, owner attention, approval latency, sequencing, evidence quality, and maintenance can each govern the feasible plan.
Measure capacity by period, time, and sequence. Twelve engineer-weeks across a quarter may look sufficient while one specialist is needed by three initiatives in the same week. Aggregate capacity can hide an infeasible queue.
Separate hard constraints from preferences. Safety, security, privacy, legal, and production boundaries are eligibility conditions, not small penalties that a high benefit estimate can overwhelm. Speed, cost, breadth, resilience, and convenience may be traded only by an accountable decision maker. Choosing what “best” means is a governance decision before it is a computational problem.
Constraint migration: measure where the queue moved
A representative sequence, not a prediction: an improvement can expose a new queue and governing constraint, which requires remeasurement.
- Execution capacity
Agents increase the amount of work that can start.
- Specialist review
More outputs arrive for scarce expert evaluation.
- Integration
Approved parts compete for coordination and shared environments.
- Owner decision
Consequential trade-offs wait in a serial approval queue.
- Operational adoption
Released change exceeds the organization’s ability to absorb it.
Improvement → new queue → new governing constraint → remeasurement. Constraints can skip, recur, or coexist; the sequence above is illustrative.
Causal intervention
Choose the intervention before adding capacity
Capacity is only one response. A team may gain more by stopping low-value work, reducing work in progress, resequencing, removing a dependency, narrowing scope, clarifying acceptance, improving evidence, delegating a bounded decision, or running a smaller experiment. It may build or retire a skill or agent, use deterministic software, acquire expertise, redesign the workflow, or preserve slack.
Every intervention needs a causal claim: name the governing constraint and evidence; explain the relief mechanism; choose the expected metric; identify the next bottleneck or control burden; include lifecycle cost; define exit; and set remeasurement. Without that chain, capacity is being added to a story rather than a demonstrated constraint.
Remeasure end to end after every material change. Faster review may reveal integration delay; faster integration may expose owner-decision latency; faster approval may exceed adoption capacity. Never assume that removing one bottleneck removes the system constraint. Measure where the queue moved.
Capability allocation
Optimize agent deployment—not agent count
Ask which execution arrangement delivers an acceptable result at the lowest justified total cost and risk. Use deterministic software for fixed rules, workflows for predefined paths, skills for reusable procedures, agents for adaptive execution, and accountable people for consequential trade-offs and approvals.
Before adding an agent, record the current bottleneck, the causal relief mechanism, the observable result expected to improve, and the condition for removing, merging, replacing, or redesigning it. A role label or delegable task is not enough.
Total cost includes tools, setup, context, handoffs, review, retries, coordination, verification, correction, security, governance, and downstream defects. An agent adds capacity only when its value exceeds those costs. Parallel agents can suit genuinely parallel, high-value work, yet lose to one agent, a workflow, or software when coordination dominates.
EXTERNAL EVIDENCE OpenAI’s agent-building guide recommends incremental orchestration, evaluation, guardrails, exit conditions, and human intervention; Anthropic’s guidance likewise says complexity should be added only when it demonstrably improves outcomes. Neither source establishes one universally best topology.
Reusable capability
Treat skills as lifecycle investments
A production skill is a versioned package with creation, validation, invocation, dependency, failure, maintenance, monitoring, ownership, and retirement costs. Reuse amplifies value and error: a popular defect can spread farther than a one-off mistake.
Require repeated demand, a named capability gap, a stable-enough procedure, explicit inputs and outputs, acceptance tests, known exceptions, lifecycle ownership, monitoring, and a retirement condition. Then test whether reuse creates net improvement; reuse count cannot answer that.
Evidence can include verified uses, first-pass acceptance, time saved per accepted result, review reduced, escaped defects, rework, maintenance, dependent workflows, cross-context performance, validation age, and net lifecycle benefit. No new skill should enter without repeated demand, an owner, a validation plan, and a retirement condition.
EXTERNAL EVIDENCE Anthropic describes Agent Skills as organized instructions, scripts, and resources. The lifecycle investment test in this article is a proposed governance extension, not a reported vendor benchmark.
Choose the strongest combination—not the highest individual score
Hypothetical benefit units illustrate selection logic; they are not money, forecasts, or observed business results. Capacity is 12 engineer-weeks and 5 specialist-review days.
| Initiative | Benefit units | Engineer-weeks | Review days |
|---|---|---|---|
| A: Premium feature | 100 | 8 | 4 |
| B: Reliability improvement | 70 | 6 | 2 |
| C: Onboarding improvement | 60 | 6 | 3 |
Initial feasible set
A has the highest individual benefit. It cannot pair with B or C. B + C exactly uses both constraints and produces 130 units, the best feasible listed combination under the initial assumptions.
| Combination | Benefit | Engineering | Review | Status |
|---|---|---|---|---|
| None | 0 | 0 | 0 | Feasible |
| A | 100 | 8 | 4 | Feasible |
| B | 70 | 6 | 2 | Feasible |
| C | 60 | 6 | 3 | Feasible |
| A + B | 170 | 14 | 6 | Infeasible |
| A + C | 160 | 14 | 7 | Infeasible |
| B + C | 130 | 12 | 5 | Feasible |
| A + B + C | 230 | 20 | 9 | Infeasible |
The capability changes the feasible set
- Agent interventionB engineering demand falls from 6 to 4 engineer-weeks.
A + B uses 12 engineer-weeks but 6 review days: still infeasible.
- Benefit
- 170
- Engineering
- 12
- Review
- 6
- Status
- Infeasible
- Validated skill interventionA review demand falls from 4 to 3 days.
A + B uses 12 engineer-weeks and 5 review days: feasible.
- Benefit
- 170
- Engineering
- 12
- Review
- 5
- Status
- Feasible
Optimize the combination and the governing constraints—not an isolated productivity figure.
Serial capacity
Make owner back-and-forth a first-class constraint
Owner attention is a scarce, high-cost, and usually serial resource. Agents may execute in parallel, but an accountable owner cannot responsibly resolve unlimited consequential decisions at once. Unnecessary involvement adds waiting, context switching, repeated explanation, stale approvals, delayed acceptance, fatigue, and rework.
Classify every escalation. Reserved matters include material spending, production release, credentials, legal commitments, destructive action, external publication, strategy, irreversibility, and defined capital or customer risk. A reversible bounded decision may reveal missing delegation; repeated clarification may reveal an incomplete packet.
A decision packet should state timing, alternatives, recommendation, evidence, benefit, cost, risks, reversibility, inaction, exact authorization, prohibitions, and reconsideration conditions. No owner escalation without a reserved matter, complete evidence, and an exact decision request. This protects accountable judgment.
| Metric | Management question | Working definition |
|---|---|---|
| Owner touches per accepted deliverable | How much owner attention does delivery consume? | Count distinct owner interactions from request to acceptance; divide by accepted deliverables in the same cohort. |
| Median owner decision latency | How long does work wait for an accountable decision? | Measure elapsed time from a decision-ready packet to the recorded decision; report the cohort and excluded pauses. |
| Avoidable escalation rate | How many escalations could have been delegated or preauthorized? | Classify resolved escalations against the current reserved-matter policy; divide avoidable cases by all escalations reviewed. |
| First-pass decision completeness | Could the owner decide without requesting missing information? | Share of decision packets decided without a request for required evidence, alternatives, scope, or authority detail. |
| Reserved-matter precision | Did escalations actually meet owner-only criteria? | Share of escalations that satisfy at least one documented reserved-matter rule after review. |
| Approval staleness rate | Did source or evidence change before the approved action occurred? | Share of approvals invalidated or reconsidered because a material candidate, evidence, risk, or policy condition changed first. |
| Decision reversal rate | Were choices made with sufficient evidence? | Share of recorded decisions later reversed; classify the reason rather than treating every reversal as failure. |
| Owner-queue percentage of total lead time | Has governance become the dominant constraint? | Owner-queue elapsed time divided by request-to-acceptance lead time for the same accepted cohort. |
Measurement architecture
Keep five kinds of evidence separate
One optimality score hides the causal structure. Outcome measures ask whether improvement occurred; flow measures show how work became accepted; constraint economics identifies scarcity and displacement; capability measures test future feasible choices; control and evidence measures establish authority and trustworthiness.
The levels connect but are not substitutes. A shorter queue matters only if accepted delivery and outcomes remain sound. A reused skill can cut cycle time while increasing escaped defects; a compliant release can fail to create adoption. Activity is not proof of customer value.
Define each measure with a population, source, cadence, owner, limitation, and decision trigger. No metric without a decision it informs. Preserve the levels long enough to see which intervention changed the system and where the constraint moved.
Outcome
Did the intended improvement occur?Task success, adoption, retention, revenue, cost or risk reduction, reliability, user effort, satisfaction, willingness to pay, benefit realization.
Delivery flow
How efficiently did work become accepted?Request-to-acceptance lead time, accepted deliverables, first-pass acceptance, rework, queue time, work in progress, dependency waiting, terminal states, handoffs, context repetition.
Constraint economics
Which scarce resource governed throughput, and what was displaced?Specialist and owner queues, engineering or integration capacity, compute budget, operational readiness, total cost per verified result, opportunity cost, time in queue, migration after intervention.
Capability
Did the intervention improve future feasible choices?Cycle-time reduction, combinations added, review demand or failure reduced, verified reuse benefit, maintenance burden, cost avoided, degradation, net lifecycle benefit.
Control and evidence
Was the work authorized and is the result trustworthy?Authority exceptions, verification failures, evidence completeness, unsupported claims, defect escape, rollback, reversals, approval staleness, independent review, unresolved risk.
Recurring allocation
Manage optimization as a decision cycle
Define the beneficiary, desired change, baseline, target, horizon, guardrails, owner, and whether the investment seeks benefit, risk reduction, or learning. Map interventions, capabilities, dependencies, scarce resources, authority, evidence, and do-nothing, defer, and stop options.
Compare feasible combinations rather than ranking initiatives alone. Test sensitivity to benefit, effort, review, owner availability, adoption, dependencies, and capability cost. Commit manageable work while preserving capacity for uncertainty and correction.
Verify delivery and benefit separately: acceptance evidence tests the deliverable; outcome evidence tests the intended change. Then continue, expand, revise, pause, defer, stop, change capability or authority, or reopen the model. Record what changed. Work should not retain resources merely because it started.
Define the outcome
Name beneficiary, desired change, baseline, target, horizon, guardrails, owner, and whether the aim is benefit, risk reduction, or learning.
Map the system
List alternatives, capabilities, dependencies, scarce resources, authority boundaries, evidence, and do-nothing, defer, and stop options.
Compare feasible combinations
Test combinations and sensitivity to benefit, effort, review, owner availability, adoption, dependencies, and capability cost.
Commit selectively
Authorize manageable work and preserve capacity for uncertainty, correction, and justified interruption.
Verify delivery and benefit separately
Use acceptance evidence for the deliverable and outcome evidence for the intended change.
Reallocate deliberately
Continue, expand, revise, pause, defer, stop, change capability or authority, or reopen the model—and record why.
Reusable operating tool
Delivery Constraint and Capability Allocation Canvas
Use this canvas to prepare an allocation decision before adding work or capability. Entries remain only in this page while it is open; they are not submitted, saved, or analyzed.
Evidence boundary
The framework does not guarantee an optimal outcome
Value resists one scale, estimates can be wrong, and constraints can change. Models can omit alternatives, interactions, sequence effects, or constraints; decision makers can mistake preference for a hard limit or treat a safeguard as negotiable.
Several alternatives may be Pareto-efficient, leaving leadership preference unresolved. A defensible choice can have a poor outcome; a weak choice can benefit from luck. Agent and skill investments can reveal maintenance, dependency, and governance costs only after reuse.
This is a contextual management framework, not a universal model or guarantee. Test its proposals against the consequence of error, evidence quality, actual authority, and observed delivery behavior.
Primary and authoritative sources used
Links reviewed September 5, 2026. Each supports only the narrow statement shown.
- Scrum.org · Evidence-Based Management Guide
The distinction among inputs, activities, outputs, outcomes, and impacts, plus adaptation through evidence.
- PMI · Lexicon of Project Management Terms
Current standardized project, program, and portfolio terminology.
- Google OR-Tools · CP-SAT solver statuses
The technical distinction between a feasible solution and an optimal solution.
- OpenAI · A practical guide to building agents
Incremental orchestration, deterministic alternatives, evaluation, guardrails, exit conditions, and human intervention.
- Anthropic · Building effective agents
Matching agentic complexity to task needs, latency, cost, and demonstrable outcome improvement.
- Anthropic · Equipping agents with Agent Skills
Skills as organized, composable instructions, scripts, and resources rather than separate workers.
Operating principle
Allocate, verify, then reallocate
Select the most valuable feasible combination of work. Equip people and agents with validated capabilities. Coordinate delivery around the system constraint. Protect owner attention for consequential decisions. Verify the benefit separately from the output. Then revise the allocation when evidence changes.
Before adding another initiative, agent, skill, review, or approval, identify the current bottleneck and name the metric that should improve. If the intervention cannot be tied to a constraint and an observable delivery outcome, it is not yet justified.