You can hire an AI CMO today. Ours has shown up 24 working days in a row

The line "you can hire an AI CMO today" is doing the rounds on LinkedIn. It is true, and it is also the least interesting part. A CMO who onboards herself and writes a 90-day plan is a demo. A team that reads your inbox, your CRM and your campaign tool every night, and has a briefing on your desk at 06:30 for 24 working days straight, is an operation.

We run a link building agency across Spain, Latin America, Germany and Italy with a very small human team. The reason that works is a group of six agents built in Claude Code. This piece shows how the team is structured, the one rule that decides which model gets which job, and where it broke in three months of running. At the end is the starter structure you can set up tomorrow.

What an AI marketing team actually is

Most tutorials give every agent a first name and a job description. Fine as a starting point, but it hides what makes the thing a team. Three parts:

  • A folder, not a prompt. Context lives in files: rules, client profiles, voice, proof sentences with their source. Every agent reads the same folder before it works. If your context lives in the prompt, you are re-explaining your preferences every week.
  • A schedule, not a chat. The agents run without being asked. Briefing at 06:30, content run Monday 10:00, publisher checks overnight. An agent you have to start is a tool. An agent that is waiting for you in the morning is a colleague.
  • One rule about who decides. No agent sends, posts or books anything. It writes a card with options, a human picks one, and the decision is saved as a file. Our internal name for it is standing rule 4: nothing leaves the folder without approval.

Everything else is casting.

The six roles

This is our team as of September 2026. The names do not matter. The cuts do.

  1. The morning briefing. Reads every inbound channel, writes the daily brief with a ranked list of proposals and exactly one thing that happens first today. Monday to Saturday 06:30, Sunday as a weekly review at 07:00.
  2. The inbox agent. Works through publisher replies and client requests, records terms in the CRM, drafts answers. Never sends.
  3. The publisher check. Verifies overnight whether booked placements are live and whether the link is still set the way it was ordered. Reports deviations instead of fixing them.
  4. The content run. Mondays: collect ideas from the knowledge base, write two of them in three languages, open one approval card per piece. The text you are reading came out of that run.
  5. The treasurer. A script computes payment status, the agent reads the related email threads and decides line by line whether to follow up. The agent does not calculate. It judges.
  6. The CEO agent. Twice a week, a bottleneck diagnosis: where the value chain is stuck and whether that bottleneck is worth solving. The only agent whose product is a judgment.

Six roles across four functions: buying, sales, content, finance. No generalist. Every role has a trigger, a reporting duty and a place where its output lands.

The six roles of the AI marketing team: morning briefing, inbox agent, publisher check, content run, treasurer and CEO agent

The rule that picks the model

The expensive question in an agent team is not which model is best. It is which model gets which role. Our rule has been in a note since 22 August and has not changed:

Do not ask how hard the task is. Ask what the product of the run is.

  • If the product is a judgment, use the strongest model. For us that is exactly one role, the CEO agent.
  • If the product is material (a draft, a summary, a sort order), stay in the middle. Briefing, inbox, content run and treasurer all run there.
  • If the product is a check against fixed rules, go down. Or drop the model entirely: a script on a cron beats any agent once the rule can be formalized. The publisher check is half script.

Two additions that saved us money. First: deterministic beats any model. Before every model step we check whether a query, a filter or a script gives the same result. Second: upgrade only with evidence. Not "this feels complex", but "this run missed X, here is the line". Otherwise the whole team drifts upward and the invoice with it.

Picking the model by the product of the run: judgment, material or check against fixed rules

Where it broke

A team that has been running for three months has a failure history. Ours in three points, because they explain the build rules.

Nine runs failed silently. The briefing ran as a pilot over the summer. At the end of August it failed nine days in a row and nobody noticed, because an agent that does not write also does not write an error message. Since 1 September every run has a reporting duty: it logs itself with a timestamp even when it found nothing. The 24 consecutive runs are only countable since that rule.

A number that was wrong for three days. The briefing reports the lead stock across campaigns. On 28 September it showed 465 approved leads. 124 of them sat in paused campaigns and could not send. The honest stock was 341. The error surfaced when we added a check question for every number in the brief: what would someone still have to do for this number to be true? An agent counts correctly. It just does not know what counts.

A watchdog that reported "all clear" every day. A check agent was supposed to flag stale publisher contacts and excluded an entire category, because the code carried an exception whose justification was no longer valid. Result: out of 1,139 records in that category, 473 had not been contacted in over twelve months, and the watchdog reported green daily. Since 23 September every exclusion rule in a watchdog has to carry its own evidence of when it was last verified.

The lesson from all three: an agent team does not need a better prompt. It needs one agent that watches the others work and reports where the same correction was needed for the third time.

Three failures in three months of running and the rule that came out of each

The starter structure for tomorrow

You do not need a framework or a platform. You need one folder, six entries in it, and one cron job.

  • rules.md: who decides, and what never goes out without approval.
  • context/: client profile, voice, proof sentences with sources. What every agent reads before it works.
  • agents/briefing.md: trigger, sources, output format, reporting duty. One file per role.
  • inbox/: cards with options that a human decides.
  • decisions/: answered cards, with date and reasoning.
  • log.md: one line per run, including "nothing found".

And the order in which to build the team:

  1. Start with the briefing, not with sales. An agent that only reads and reports cannot break anything. It teaches you what your context folder needs to look like.
  2. Give it a schedule before you give it more tasks. It only becomes a team member once it runs without you.
  3. Write the approval rule before the first agent could send anything. Card, options, decision as a file. It is the one rule we have never loosened.
  4. Pick the model by the product of the run, not by how difficult the task feels.
  5. Build the watcher second. An agent whose only job is to read the others' logs and report repeated corrections.
The order in which to build the agent team, in five steps

What we do not have yet is the watcher. It has been an open question in our notes since 31 August. When it runs, I will write up what it found in its first week.