Case study · AI systems: agents, teams and results

One person, three languages, 300 URLs: y8y's editorial system with AI agents

Sebastián Ocampo · 2026-07-20

This case doesn't come from a client: it is the house you are standing in. Everything claimed below can be checked by browsing the site itself, reading its public content index or inspecting its markup. That is the lab's rule: auditable cases or no case.

The problem: a serious publication demands a team that didn't exist

A competitive niche publication needs, at minimum: documented writing, claim verification, translation for each market, technical SEO (markup, hreflang, sitemaps), distribution to search engines and a content map that prevents cannibalization. With a traditional team, that is four or five people before the first reader. The y8y experiment was to ask: how much of that can become a system, leaving the one available person the work that genuinely requires judgment?

The answer was not "use a chatbot to write". It was to design an editorial chain where every link (research, draft, localize, validate, interlink, distribute) has a checkable definition of done, and the AI executes inside that definition.

The system: a content graph that agents can operate

The central piece is not an AI model: it is the data structure. Every entity (robot, story, comparison, term, answer) is a JSON file with trilingual fields, a search query it owns and an extractable answer. On that graph, AI agents do the competitive research, draft to a written editorial playbook, generate the three language versions and propose the internal links. A battery of automatic validations rejects any invalid entry: data schema, broken references, duplicate queries (the anti-cannibalization guard) and even style rules.

Distribution is also a system, not a chore: the content index is published as a machine-readable file, search engines are notified automatically after every deploy, and the full answers are served to AI engines in a single file. The person stays exactly where they should be: choosing topics, verifying facts against primary sources, signing verdicts and deciding what does not get published. The full editorial method is public in the methodology.

  1. Target query
  2. Agent research
  3. Draft per playbook
  4. Automatic validation
  5. Human review
  6. Publish
  7. Search distribution

The account and the governance: what it costs and what is never delegated

Direct costs: monthly infrastructure is zero (a static site on its host's free tier) and the real spend is AI tool subscriptions plus the person's hours. That is the pattern that matters for a marketing team: the dominant cost of an agent system is not the software, it is the supervision, which is why the entire design aims at making human review fast (automatic validations before, clear diffs after).

Governance in three rules. One: the AI proposes, the person publishes; no content reaches the site without human review. Two: every factual claim carries a cited source, and figures only a manufacturer asserts are labeled as such. Three: there are hard rules the tests make impossible to skip, from the data schema to style. What this case does not claim yet, out of honesty: revenue and consolidated traffic; the publication is young and those numbers will be added here, dated, once they exist. The system that will produce them is what stands documented today.

How to build a system like this in your team, step by step

The order matters more than the tools: each step creates the condition the next one needs. With one person's partial dedication, the full journey takes weeks, not quarters. The full prescriptive version (budget-leak diagnosis, rebuilt team, tools with pros and cons, token economics and risks) is in the content operation blueprint.

  1. Inventory your entities List what your business genuinely knows: products, cases, customer questions, proprietary data. Every entity with verifiable facts is a potential page; what you cannot prove does not get in.
  2. Assign one query per page Before writing anything, decide which search each page answers and keep a single map. Two pages for the same query compete with each other; that map is your anti-cannibalization contract.
  3. Write the playbook before the content Answer format, style rules, source requirements, what is forbidden. Agents execute what is written down; what lives only in your head does not scale and gets broken.
  4. Structure content as validatable data One file per entity with defined fields, and automatic tests that reject the invalid: schema, broken references, duplicate queries, style rules. If your guide cannot reject a piece, it is a suggestion.
  5. Let agents produce inside the frame Research, drafts, localization and internal-link proposals: all inside the playbook, with sources cited for every claim. Speed comes from here; safety, from the previous steps.
  6. Close with a human gate and measurement Nothing publishes without human review; distribution (sitemaps, search-engine notification) is automated after every deploy; and a weekly review of search data decides what gets deepened or fixed. The system improves through that loop, not through volume.

Three content team models, compared

A manager's real decision is not which tool to buy but which team model to operate. All three work; they differ in where they place the cost and the risk.

Traditional teamTeam with AI toolsAgent-operated system (this case)
Dominant costProduction salariesSalaries plus licensesHuman supervision and system design
Where quality livesIn each person's judgmentIn judgment, with more speedIn the playbook and tests; judgment concentrates on review
Main riskHigh fixed cost, little elasticityProductivity without a system: uneven qualitySystematized errors if the human gate is missing
Multilingual scaleEach language multiplies the teamAssisted translation, manual reviewLanguages in parallel from the same data graph
When to choose itA brand whose voice is irreproducible and high-riskTransition: learning what to delegate before systematizingEntity- and data-based content, several markets

The decisions: why this and not that

A case that doesn't show its alternatives teaches nothing: any architecture looks inevitable told backwards. The principle that ordered every choice was this: boring, replaceable technology wherever there is no competitive advantage, and custom design only where there is one. In this system the advantage lives in exactly two places, the data structure (the entity graph with its query map) and the editorial playbook with its tests. Everything else was chosen for cost, robustness or easy exit, and every piece has an honest substitute: the static generator could be Hugo or Eleventy instead of Astro; the JSON file store could migrate to a CMS with a UI (Keystatic, Contentful) the day a team writes; the general-purpose agents could be any competent LLM operating the same playbook.

The table summarizes the six structural decisions with the reasonable alternative that was discarded and the real reason. None is dogma: in another context (a large team, high-risk regulated content, a brand without proprietary data) several flip, and the last column says when.

DecisionChosenReasonable alternativeWhy here (and when the opposite)
PlatformGenerated static siteClassic CMS (WordPress)Maximum speed, zero cost, no attack surface; content is reviewed like code. WordPress wins if a non-technical team edits daily.
Content storeOne JSON file per entity, in gitHeadless CMS or databaseEvery change is a reviewable diff, tests validate before publishing and no vendor holds the data. The CMS wins with several simultaneous editors.
ProductionGeneral-purpose agents under a written playbookAI content SaaSFormat, rules and quality live in our playbook, not in someone else's product; no vendor lock-in. The SaaS wins if nobody exists to write and maintain that playbook.
Quality controlAutomatic tests + human gateManual review onlyRepeatable errors (schema, links, duplicates, style) become impossible and the person reviews judgment only. Pure review wins for unstructured opinion content.
HostingStatic on a free tier, global CDNOwn server or VPSZero fixed cost and nothing to administer; our outage came from domain configuration, not a server. The VPS wins when there is real server logic (accounts, payments).
DistributionProgrammatic: sitemaps, search-engine notification, AI-engine surfacesWait for natural crawlingNew content reaches indexes in minutes and answer engines get the citable content in a single file. Waiting only wins at doing nothing.

What failed (and what it changed in the system)

Three real failures, documented because each one rewrote a rule. One: a story shipped with an unverified embedded video (the ID came from a search result and the player had a styling defect); a reader found it as a black box on mobile. Since then no external content gets embedded without prior technical verification, and the player is tested on mobile before publishing. Two: the first draft of the first data study came out thin and one of its key figures, one hundred thousand vacancies attributed to a public agency, did not survive verification: the real source was an employers' association quoted in a press release. It was caught before publication, the study was rewritten around the discrepancy (which turned out to be the most interesting finding) and the primary-source verification pass stopped being optional.

Three: automated distribution to search engines failed silently for days. The cause was not in the code but in a hosting domain configuration, and it only surfaced by examining real behavior from a neutral environment. The resulting rule: after every deploy, the system checks the site's live behavior, not the assumption that the deploy worked. None of the three failures was caught by the AI alone; all three were turned into rules by a person. That is the honest account of operating with agents.

Where this evolves (what I expect and why)

This section is the author's analysis, dated July 2026, and will be revised in public like everything else. The underlying bet: search is splitting into two economies, the click economy and the citation economy. Engines that answer (AI summaries, assistants, agents) don't send visits, they cite sources, and they cite best what they can verify and extract cleanly. That is why every piece of this system has a double output: the page for people and the surface for machines (extractable answers, structured data, the content index, the full answers file following the llms.txt convention). I expect those surfaces to standardize and become as mandatory as the sitemap; whoever has them first accumulates citations while everyone else debates.

Second expectation: readers will increasingly be agents acting on someone's behalf (comparing vendors, preparing a report, verifying a claim). An entity graph with dated, sourced data is exactly what an agent can consume with attribution; a blog of running prose is not. The opportunity I see: small sites structured like this one function de facto as trust APIs for those agents, and that depends not on domain size but on the quality and verifiability of the data. It is the first era of SEO where the small, rigorous operator holds a structural advantage over the big content farm.

Third: the cost of producing content will keep falling toward zero, which means production stops being the business. All the value shifts to what this case calls the account and the governance: proprietary data, verification, named accountability and the published record of hits and failures. What I will add to this system next, in this order: scheduled re-verification of dated claims (the system flags when a cited figure ages), per-surface answers (the same entity answering differently to a search engine, a voice assistant and an agent), and this same graph operating client content. When it happens, it will be documented here with its numbers.

Frequently asked

What tools make up the y8y system?
A static site generator with data in versioned JSON files, general-purpose AI agents operating under a written editorial playbook, automatic tests validating data, links and style, plus automated deployment and search-engine notification. No piece is exotic: the advantage is in the structure and the rules, not secret tools.
Does this system work for a brand that isn't a publisher?
Yes, because the pattern is general: entities with verifiable data (products, cases, customer questions), a one-query-per-page map, agents producing inside a playbook and one person governing publication. A B2B brand would apply it to its product pages, comparisons and pre-sales questions exactly as y8y applies it to robots.
What is the biggest risk of running content with AI agents?
Publishing unverified claims at scale. A manual error is one error; a systematized error is a hundred wrong pages with your name on them. That is why the two non-negotiable investments are primary-source verification before publishing and tests that structurally prevent repeatable failures. Speed without that brake is not productivity: it is reputational debt.

Sources

More cases and the method, in AI systems: agents, teams and results.