Operational blueprint · AI systems: agents, teams and results
Blueprint: rebuilding your content operation with AI agents (team, costs, risks and adoption)
Sebastián Ocampo · 2026-07-24
This is not a case: it is the recipe. Everything prescribed below is running live in this house and documented with its numbers and failures in the sibling case study. Read it the way you would read the plans of an installation that already works.
- 3 Real cost line items: API/subscriptions, platform and supervision
- x10-100 Reduction in production cost per piece (supervision does not drop the same)
- 6 Models with published prices in the calculator (Claude and OpenAI, July 2026)
The diagnosis: where your budget is leaking
Before buying anything, measure where it hurts. In a manual content operation the budget leaks through four places almost nobody accounts for separately. One: repetitive production, hours of qualified people writing variants, adaptations, translations and summaries of things already written. Two: duplicated research, every piece starts from zero because knowledge lives in heads and scattered documents, not in a reusable structure. Three: artisanal quality control, someone re-reads by hand what a rule could check alone (links, formats, mandatory data, style). Four: manual or forgotten distribution, publishing without fresh sitemaps, without notifying search engines and without surfaces for AI engines, which is producing without delivering.
The quick test: ask your team for the hour breakdown of the last important piece. If more than half went into mechanical production and checking (not deciding what to say or verifying facts), you have a leak an agent system closes. If almost everything went into judgment (what to tell, to whom, with what proof), your bottleneck is not this blueprint: it is strategy.
The team rebuilt: from producers to governors of the system
The rebuild does not eliminate the team: it changes what each role does and where quality lives. The general rule: everything that was typing moves to the system, everything that was judgment concentrates and gets paid better. In a small operation the four roles in the table can be two people, or a single one with the system well built (the reference case runs three languages with one).
| Role before | Role after | What it gains, what it loses |
|---|---|---|
| Production writer | Editor-supervisor: reviews, corrects and signs what agents propose | Gains reach (more pieces, more markets); loses the typing. Editorial judgment becomes their entire product. |
| Researcher | Source verifier: validates what agents bring, hunts the false figure | Gains depth (verifies instead of compiling); takes on the system's most critical role: systematized errors get caught here. |
| Translator / localizer | Market reviewer: adjusts nuance, register and local reality of already-generated versions | Gains volume per hour; language stops being a bottleneck and becomes cultural quality control. |
| Content lead | System owner: playbook, query map, governance rules and the P&L | Gains leverage (the playbook scales, meetings do not); answers to management for the full account, supervision included. |
The exact process: the three implementation phases
The implementation order matters more than the tools, because each phase creates the condition the next one needs. Phase 1 (weeks 1-2), the bounded pilot: choose a single repetitive task with clear output and cheap errors (drafts of one specific format, localization of existing pieces, quality control against your style guide), define the baseline (what it costs today in hours) and run the pilot with total human review. Phase 2 (weeks 3-5), the system: write the playbook, structure content as data with automatic validation and build the one-query-per-page map. Phase 3 (week 6 onward), the operation: agents produce inside the frame, the person reviews and signs, distribution is automated, and a weekly data review decides what gets deepened. The detailed step-by-step of phase 2, with every rule, is in the reference case.
- Bounded pilot with baseline
- Written playbook
- Content as data + validation
- Query map
- Agent production
- Human review and sign-off
- Automated distribution
- Weekly improvement loop
| Example | First workflow to build | Why that one and not another |
|---|---|---|
| E-commerce with 2,000 product pages | Agents rewriting descriptions from product data (attributes, dimensions, uses), with mandatory-data validation and sampled review | Content springs from structured data you already have, errors are cheap and detectable, and the volume makes manual unviable |
| B2B services firm with 10 client cases | Agents researching every buying query in the sector and drafting sourced answers; the person verifies and adds first-hand experience | Here the value is expert judgment: the agent prepares the ground and the expert adds what nobody can copy |
| Niche publisher | This house's complete system: entity graph, playbook, trilingual production agents, tests and automated distribution | It is the case documented piece by piece, costs and failures included, in the reference case |
Tools to evaluate, with their cons in plain sight
No tool in this table is the system: the system is the playbook, the query map and the governance, which are yours and travel with you if you switch tools. Evaluate every category by its exit cost, not just its entry cost. The cons are written in the same ink as the pros, because a recommendation without cons is advertising.
| Category | Options to look at | Pros | Cons |
|---|---|---|---|
| Publishing platform | Static generator (Astro, Hugo) · WordPress · Webflow | Static: zero cost, speed, versioned content. WordPress/Webflow: editing without engineers. | Static requires someone comfortable with git. WordPress adds maintenance and attack surface; Webflow locks content into its platform. |
| Production agents | General-purpose agents (Claude, and equivalents) operating your playbook · AI content SaaS | General-purpose: total format control, pay-per-use, no lock-in. SaaS: hours to start, non-technical interface. | General-purpose demands writing and maintaining the playbook. SaaS locks your rules inside its product, and its quality changes when they switch models, not when you decide. |
| Validation and quality | Automatic tests on structured content (schemas, style rules) · pure manual review | Tests: repeatable errors become impossible and human review concentrates on judgment. | They require structuring content first (the hard part). Pure manual review does not scale and skips things on Fridays. |
| Distribution | Sitemaps + automatic notification (IndexNow, GSC) + AI surfaces (llms.txt) · publish and wait | Automatic: indexing in minutes, content citable by answer engines from day one. | Requires initial technical setup and verifying it actually works (our documented failure was exactly here). |
The system orchestration, in one diagram
Four layers, each with its parts. The reading rule: agents (layer 2) only touch what layer 1 defines, layer 3 decides what gets published, and layer 4 delivers hands-free. If, when drawing your version, a part fits no layer, be suspicious of it.
- 1 · Data and rules
- Entity graph
- Query map
- Editorial playbook
- Source register
- 2 · Agents
- Research
- Drafts
- ES/FR/EN localization
- Internal-link proposals
- 3 · Governance
- Automatic tests
- Source verification
- Human review and sign-off
- Failure log
- 4 · Distribution
- Publishing
- Sitemaps and search notification
- AI-engine surfaces
- Weekly measurement
The real economics: tokens, subscriptions, platform and people
First, the complete list of what gets paid, because no line item is zero unless your specific case makes it zero. Models: either API consumption by token (published prices in July 2026: Claude between 1 and 25 dollars per million tokens depending on model and direction, OpenAI GPT-5.6 between 1 and 30) or flat-rate subscriptions to agent tools, often the simplest route for small operations. Platform: a static site can run on free tiers, but a CMS adds its monthly fee, plugins and, if you publish via API, the cost of each integration; neither route is the smartest in the abstract, it depends on who edits and what you already run. And supervision: the line item nobody invoices but everyone pays.
The calculator below does the full math with all three terms and real Claude and OpenAI prices. The piece sizes are this operator's estimates (producing a piece consumes far more tokens than its final text: research, drafts, corrections and languages); adjust them to your own measurement as soon as you have one. The reading rule: if moving the numbers makes your cost per piece look too good, you have almost certainly set supervision too low. How to turn this cost into defensible ROI, with a baseline, is in our ROI answer.
And the star metric is not cost per piece: it is cost per resolved case. Divide the period's total cost (all three terms) by the units that passed the human gate and served the business, not by everything the system produced. Discarded attempts go in the numerator, never the denominator: they make each resolved case more expensive instead of vanishing from the account. It is the only metric a CFO can compare against the human alternative without tricks, and it is how we use it in our own research pipeline, which reports cost per qualified firm, not cost per API call.
Cost-per-piece calculator
API prices as published in July 2026 (sources at the bottom). The full formula: tokens + prorated subscriptions + human supervision. Adjust every assumption to your case.
| Tokens per piece | 0.68 € |
|---|---|
| Prorated subscriptions per piece | 5 € |
| Human supervision per piece | 11.67 € |
| Total cost per piece | 17.34 € |
| Total cost per month | 346.86 € |
| Versus the human cost per piece | x8.65 |
The line that moves the total most is supervision. If your result sets supervision to zero, you have not found a saving: you have found a risk.
Risks: governance, dependency and business continuity
A badly governed system does not fail like a person: it fails at scale, silently and with your name on it. And a well-governed but badly designed system creates another debt: dependency. The table lists the five risks we have seen up close and each one's concrete mitigation; in regulated markets (Switzerland, the European Union), this table is literally what a buying committee will ask you to show, with the governance framework written down.
| Risk | How it materializes | Mitigation |
|---|---|---|
| Systematized error | A false figure or a bias replicates across a hundred pages before anyone sees it | Human gate before publishing, mandatory sources per claim, tests that reject the invalid |
| AI vendor dependency | Price changes, model behavior changes or service outage | The playbook and data are yours and portable; the vendor is replaceable by design. Budget a margin for rate changes |
| Key person | The whole system lives in the head of whoever built it | The written playbook and tests ARE the executable documentation: if the person leaves, the rules stay |
| Compliance and traceability | Being unable to show what the AI decided, what a person reviewed and with which data | Versioned content (every change with author and date), source register, named owner per system |
| Judgment atrophy | The team signs without reading because the system is almost always right | Rotate deep review, measure the correction rate (if it drops to zero, be suspicious), and publish the failures found |
Warnings: content at scale without getting burned
The capacity to produce a hundred pages a week is also the capacity to destroy a domain in a quarter. Google has an explicit policy against scaled content abuse (mass content created to manipulate rankings without adding value), and its helpful-content systems reward exactly the opposite: pages someone would read even without coming from a search engine. The four rules this blueprint treats as non-negotiable at scale: publish below your real supervision capacity, never above it; one query per page, with a single map that keeps your own pages from competing with each other; structural depth in every piece that touches money (data, tables, steps, sources), not thin variants of the same thing; and delete or merge what does not perform, mercilessly, because a hundred mediocre pages drag down the ten good ones.
The least-told warning: brand risk beats search risk. A factual error on one page is a typo; the same error templated across fifty is your reputation. That is why layers 1 and 3 of the diagram exist before layer 2 accelerates, and why publishing pace is a governance decision, not a capacity one.
The first wins, and how to bring your colleagues along
The first two weeks' wins are not articles: they are proofs. An automatic quality check that catches errors in the content you already have (broken links, outdated figures, style inconsistencies) produces a list of visible fixes nobody argues with. Localizing your three best pieces into a second market demonstrates reach without editorial risk. And the weekly report that used to cost a morning and now generates itself is the win your boss understands without explanation.
To bring colleagues along, the classic mistake is announcing the tool; what works is distributing the system. Give the quality skeptic the verifier role with veto power: their skepticism turns from brake into paid quality control. For whoever fears for their job, show the roles table above: the typing goes away, judgment gets paid better, and whoever governs the system is worth more than whoever produced inside it. With management, agree the metrics before the pilot (baseline hours versus system hours, corrections caught, cost per piece with supervision included) so the result is a data point, not an opinion. And one rule that disarms half the resistance: the system's failures get published internally, as the reference case does publicly. A system that shows its errors is a system that can be trusted; one that only shows demos is not.
What to expect, and where this is moving
Honest expectations by quarter. First month: the system feels slower than promised, because you are paying for the design (playbook, structure, rules) that later scales; the win is the measured pilot. Months 2-3: production per supervision hour multiplies and the first search results appear; it still does not pay back the setup time. From month three: the system produces more than the team can review, and the bottleneck inverts, which is the signal it works; from there the limit is your supervision capacity, not production capacity. If at month three the bottleneck is still producing, something in the design failed: go back to phase 2.
And where this is moving: search is splitting into the click economy and the citation economy, readers will increasingly be AI agents acting on someone's behalf, and production cost will keep falling until producing is nobody's business. All the value shifts toward what this blueprint makes you build: proprietary data, verification, governance and a published record of hits and failures. The full argument, dated and sourced, is in the reference case's evolution section. The operational translation: if you build this today, you are not automating your content operation; you are building the asset that will still hold value when cheap content holds none.
Frequently asked
- What does producing content with AI agents really cost, all included?
- Three line items: AI tokens or subscriptions (tens to a few hundred euros a month for a small operation), infrastructure (zero with static hosting) and human supervision, which dominates: review minutes per piece times the hourly cost of whoever signs. Per piece produced, the machine share lands between cents and a few euros; the honest total depends on how much review your risk level demands. Any figure that excludes supervision is marketing.
- Does this replace my content team?
- It replaces the typing, not the judgment. Roles convert: the writer becomes editor-supervisor, the researcher becomes verifier, the lead becomes system owner. A smaller team operates a bigger system, and the people who stay do the better-paid work: deciding, verifying and signing. What does disappear is the position whose only content was producing variants of what was already written.
- I am an SMB with no technical team: where do I start?
- With phase 1 of the process: a two-week pilot on a single repetitive task (quality-checking your existing content is the safest start), with a baseline measured in hours and total human review. Do not buy any platform yet: the pilot runs on a general-purpose agent tool and a rules document. If the pilot proves the hours saved, then you choose system and tools with this blueprint's table; if it does not, you have spent two weeks and zero euros finding out.
- What if the AI vendor raises prices or changes the model?
- If you followed this blueprint, you absorb the hit: your playbook, your query map, your structured content and your tests are yours and work with any competent model, so switching vendors is reconfiguring, not rebuilding. The real risk falls on whoever built their operation inside a closed SaaS, where the rules and data live in someone else's product. Budget a rate margin and review the account quarterly, like any other vendor cost.
Sources
- Claude API pricing (precios públicos por millón de tokens) · Anthropic · 2026
- OpenAI API pricing (precios públicos por millón de tokens) · OpenAI · 2026
- El sistema de referencia, operando en vivo con costes y fallos publicados · y8y.ai · 2026-07
- AI Risk Management Framework · NIST · 2023
- Spam policies for Google web search (scaled content abuse) · Google Search Central · 2026
More cases and the method, in AI systems: agents, teams and results.