Operational blueprint · AI systems: agents, teams and results

Blueprint: identifying and prioritizing AI use cases (inventory, decision matrix and pilots that get measured)

Sebastián Ocampo · 2026-07-26

Identifying use cases is the duty that appears most in real AI leadership postings after strategy: nearly one in three asks for it in writing. It is also where the most money gets burned, because the alternative to a method is a wish list. This is the method: the same one this house used to decide its three production systems.

The diagnosis: why your list of AI ideas goes nowhere

An AI use case is not a technology ("use a chatbot") but a business task with an owner, a volume and a measurable current cost ("answering the 400 monthly customer queries about order status, which today consume 60 hours of two people"). Most AI idea lists fail by confusing those two things: they are born from demos somebody saw, not tasks somebody suffers, which is why none reaches production. The classic symptom has a name in the job market: in our corpus of 79 real AI leadership postings across Switzerland and Europe, identifying and prioritizing use cases appears as an explicit duty in nearly a third, precisely because companies already tried the wish list and it did not work.

The quick test for whether this is you: ask for your company's current list of AI ideas and check how many have three numbers next to them (hours or euros the task costs today, monthly volume, and who owns it). If the answer is none, you do not have a technology or talent problem: you have a method problem, and it gets fixed in two weeks with what follows.

The full process, from first interview to roadmap

The whole process fits in seven steps and two calendar weeks for a mid-sized company. The reading rule: nothing enters the matrix without its three numbers, nothing enters a pilot without a baseline, and nothing scales without its decision date honored. It is the same circuit the Chief AI Officer blueprint prescribes for the first 90 days; this is the fine mechanics of the use-case part.

  1. Interviews per function
  2. Task inventory
  3. Three numbers per task
  4. Risk filter
  5. Prioritization matrix
  6. Two pilots with a baseline
  7. Decision: scale or kill

Step by step: the inventory that finds the real tasks

The inventory is a week of structured conversations, not an ideation workshop. The goal is not for people to propose AI uses (they will propose demos they saw): it is to extract the tasks that fit the profile AI solves well today, which are almost always the ones nobody brags about doing.

  1. Interview each function owner with three fixed questions Which repetitive task consumes the most hours of your team each month? What information do you look up again and again in different places? What work piles up when someone is out? All three point at the same thing: volume, scattered knowledge and dependence on one head, the three profiles where AI performs.
  2. Add the anonymous shadow-AI form Ask which AI tools people already use on their own and for what. Where there is clandestine use there is a demand-validated use case: someone already decided it pays. You regularize it, you do not punish it.
  3. Write every candidate task in one line with an owner The fixed format: verb + object + monthly volume + who does it today. "Classify 900 support emails a month, service team, 45 hours". If an idea cannot be written like that, it is not a use case yet: it is an intention.
  4. Put the three numbers on every line before forming opinions Current cost (hours per month times hourly cost), volume, and variability (is the task identical every time or is every case a world of its own?). This step silently kills half the list, which is exactly its job: what is cheap to do by hand does not get automated.
  5. Run the risk filter before the matrix Two knockout questions: does the task touch personal or regulated data without an approved circuit for it? Would an error reach a customer or a regulator without human review in between? A yes without mitigation sets the case aside until governance exists. The order matters: safety first, enthusiasm second.

The prioritization matrix, with its scoring rubric

Three criteria, each scored 1 to 5, and the total score is value plus feasibility minus risk. Three criteria and not one more: every extra column added to a prioritization matrix is a new place to hide a personal preference. The full rubric, so that two different people score alike:

Criterion1 point means5 points means
Annual valueUnder 50 hours a year, or value impossible to express in hours or euros.Over 1,000 hours a year, or direct revenue or regulatory-risk impact with a defensible figure.
Risk (subtracted)An error is internal, visible and cheap to fix; no personal data.An error reaches customers or regulators, or touches personal or regulated data; requires the full governance circuit before piloting.
FeasibilityScattered or paper data, task different every time, integrations that do not exist.Accessible digital data, clearly patterned task, pilotable with general-purpose tools in two weeks without deep integration.

Three worked examples through the matrix

Three typical tasks of a mid-sized services company, scored with the rubric above. They are worked examples to calibrate your eye (your company's numbers will differ), and they show the pattern that surprises almost every committee: the glamorous case loses to the boring case with volume.

Candidate taskValue / Risk / FeasibilityTotal and decision
Sales chatbot on the website (the idea the committee brings)Value 2 (few conversations a month), Risk 4 (talks to customers unreviewed), Feasibility 3.Total 1. Set aside: lots of shop window, little volume and an error is expensive.
Summarize and classify the month's 900 support emailsValue 4 (45 hours a month), Risk 2 (internal draft, a person replies), Feasibility 5 (patterned digital text).Total 7. Pilot one: boring, measurable and with a happy owner.
First draft of sales proposals from the archiveValue 4 (revenue impact), Risk 3 (reaches the client, but always reviewed), Feasibility 4 (digitized archive).Total 5. Pilot two, with the human signature as a design condition.

The pilot charter: one page or no pilot

Every approved pilot gets written on one page with the table's fields, and that page is the contract. The field almost everyone omits and that saves the most pilots is the last one: the date on which the pilot stops existing as a pilot, whether scaled or killed. Without a date, every pilot tends toward eternity, because killing it looks like failure; with a date, killing it is following the plan.

Charter fieldWhat gets written
Task and ownerThe inventory line, with the person from the function accountable for the outcome (not the AI lead: the business owner).
Measured baselineThe task's current hours and quality, measured before touching anything. Without a baseline there is no honest accounting later: it is rule number one.
Success metric and thresholdOne number and its scaling threshold, agreed before starting. How to measure without inflating is in how to measure AI ROI.
Estimated cost with the three line itemsTokens and subscriptions, platform, and human supervision. The full structure with a calculator is in the content operation blueprint and applies to any function.
The pilot's governance rulesWhat the system may do alone, what requires human review, and where failures get logged. The full system is in governance that signs.
Decision dateThe day, 4-8 weeks out, when the committee decides to scale or kill with the numbers on the table. Both outcomes are success: one produces a system, the other produces cheap learning.

The five decision rules (against pilot theater)

One: no task enters the matrix without its three numbers; what cannot be measured cannot be prioritized, only wished for. Two: two pilots maximum at once; the third steals the supervision the first two need to be conclusive. Three: no pilot without a prior baseline; a baseline estimated from memory after launch always comes out inflated. Four: every decision date is honored, and scale or kill are the only two valid outcomes; "keep piloting" is the forbidden answer. Five: every quarter, the whole matrix gets rescored, because feasibility changes every few months with each model generation, and January's unfeasible case can be June's obvious one.

The fifth rule is what turns this from a one-off exercise into an operating system: the living matrix is the company's AI roadmap, and whoever maintains it (internal or fractional) holds the map that job postings call strategy.

How we use it in this house (the method, eaten at home)

This blueprint is not consultant theory: it is the filter the three production systems of this lab passed through. The editorial operation won its pilot because it was the highest-volume task with owned data and review-controllable risk; the market research system scored high on feasibility (public sources, digital text) and its risk was mitigated with the data contract; and more than one flashy candidate was killed on time, its decision date honored, because the numbers did not work. This house's matrix gets rescored every quarter, and from it will come the next cases published here, with their numbers and their failures, as always.

If you want the matrix applied to your company under our honesty rules (real baseline, dated decision, governance from day one), the usual door is below, on the card of the person who signs.

Where this is shifting (as read in July 2026)

Two shifts will change your matrix within a year. The first is the feasibility column: agents able to execute complete tasks are raising the score of cases that were unfeasible a year ago (the ones requiring navigating systems, crossing sources or sustaining long processes), so quarterly rescoring stops being hygiene and becomes advantage: whoever rescores first, captures first. The second is the risk column: the EU AI Act's obligations apply in phases and will turn filters that are best practice today into requirements with sanctions, making customer-facing cases without governance more expensive and, in relative terms, the boring internal ones cheaper. Our read: the competitive edge of the next two years will not be having AI ideas, which are already free, but the speed and honesty of the idea-pilot-decision circuit. That circuit is exactly what this blueprint leaves built.

Frequently asked

Which AI use case should my company do first?
The one that wins in the matrix, which is almost never the one the committee brings: it is usually an internal, boring, high-volume task (classifying emails, summarizing documents, preparing internal drafts) where an error is cheap and the saving is measurable in hours. As a general pattern: first an internal efficiency case with a clear baseline, and only then customer-facing cases, once governance exists and has been proven.
How many use cases should I prioritize at once?
Inventory every one that appears (twenty or thirty lines is normal in a mid-sized company), score them all in the matrix, and pilot only two. The limit is not ambition but supervision: every serious pilot consumes human review hours and the function owner's attention, and a third simultaneous pilot steals exactly that from the first two. The matrix keeps the rest queued for next quarter; nothing is lost, it just waits its turn with its score attached.
How do I know if an AI pilot is working?
By comparing against the baseline you measured before starting, with the metric and threshold agreed in the charter: task hours after versus before, on the same volume, subtracting the supervision and correction hours the system adds. If you did not measure a baseline, you cannot know, and any figure presented will be an opinion; that is why the baseline is rule number one of the method.
Do I need a consultant to identify AI use cases?
For the inventory and the matrix, no: this blueprint is the complete method and an internal person with a mandate can run it in two weeks. Where external help does pay is on three concrete points: calibrating the feasibility column (knowing what current AI can do requires operating it), setting up the first pilot's governance, and holding the decision-date discipline when the pilot belongs to someone powerful. If you buy that help, internal or fractional, demand the usual: cases with numbers, failures told and a method that stays in your house when the advisor leaves.
How often should the exercise be repeated?
The matrix gets rescored every quarter (a meeting, not a project) and the inventory gets fully redone once a year or after a major tool or business change. The reason for the quarterly rhythm is the feasibility column: with each model generation, tasks scoring 2 move to 4, and the company that rescores on time captures those cases one or two quarters before its competitors.

Sources

More cases and the method, in AI systems: agents, teams and results.