Agentic Marketing Platforms in 2026: Scored, Not Ranked

By

Eight agentic marketing platforms scored against a published rubric: which archetype each one really is, the exact rung where it breaks, and the dated vendor evidence behind every call. Not one of them reaches the compounding tier.

13 min read

No agentic marketing platform we scored compounds. Eight of them, measured against a published rubric on 29 August 2026, and every single one breaks below the rung where outcomes start changing output. That is a stronger claim than a ranking, so here is how it was reached: the rubric is published before the results, every call links the vendor's own dated page, and the rubric's author is deliberately not in the ranking. For contrast, the highest-authority board on this query, Profound's seven-platform list from 31 July 2026, carries a comparison column titled "Where the loop closes" and never answers it, discloses no scoring method, and was written by Profound's own Founding Marketing Engineer, who placed Profound first with roughly 800 words against 300 to 400 for each of the others.

The five archetypes

The buying mistake is never picking the 7 out of 10 instead of the 8. It is buying a workflow tool while believing it will get smarter. So the verdict here is a category, not a score.

Archetype What it really is Buy it when
A0 Generator Makes things when asked. Nothing persists. You are the loop. You need volume of artifacts and you supply all the judgment and all the memory.
A1 SaaS Automation Runs workflows a human designed, on a schedule. Rules with AI inside them. The process is known and stable and you want it executed without staffing it.
A2 Assistant Talks, drafts, suggests, and remembers you. Still waits to be asked. You want leverage on one operator's throughput, not an unattended system.
A3 Executor Given a goal it plans and acts. Each run starts cold. It does, it does not learn. The work is multi-step and variable and you accept re-teaching it forever.
A4 Compounding Engine Remembers, measures what its output produced, and changes its own future behaviour. You want an asset that appreciates. The only archetype where month 12 beats month 1 for free.

The archetype is derived from the scores, never assigned by opinion: a platform reaches A4 only when Evolution, Memory and Outcome all clear rung 3. Below that it falls to A3 if it chooses its own steps at runtime, A1 if it only runs a sequence someone authored, A2 if it mainly remembers you, and A0 if nothing persists. The line between A1 and A3 is where most "agents" die: running a sequence a human wrote is automation, and choosing the sequence is agency.

Nothing in this cohort reaches A4. That is the finding.

The board

Archetype is the verdict; the grade is the arithmetic. A platform's grade is its weakest dimension out of eight, not its average, because a system compounds at the rate of its weakest link.

Platform Archetype Grade (0-5) Breaks at Last verified
Salesforce Agentforce A3 Executor 2 Evolution: no learned rule read at generation 29 Aug 2026
Microsoft Copilot Studio A3 Executor 2 Evolution: the same wall 29 Aug 2026
Jasper A1 SaaS Automation 1 Evolution: outcomes not stored against cause 29 Aug 2026
HubSpot Breeze A1 SaaS Automation 0 Eval: no output check found 29 Aug 2026
Profound A1 SaaS Automation 0 Eval: no output check found 29 Aug 2026

One clarification the table cannot hold. A grade of 0 from "no output check found" is weaker evidence than a grade of 0 from a tested absence, and we mark the difference. For HubSpot and Profound it means no such mechanism appears in their own documentation or changelogs, which is not the same as proving none exists. Where a claim was tested and failed rather than simply missing, the platform's section says so.

How it was scored

Eight dimensions, each a five-rung ladder, published in full on the Agentic Compounding Index:

  1. Memory. Does state persist, and is it read before acting?

  2. Evolution. Does outcome data change future output, automatically?

  3. Distinctiveness. Is output measured against the machine default, or only against your own style guide?

  4. Production. Can it make the asset end to end, at channel standard?

  5. Direction. Can a human command it without approving every step?

  6. Outcome. Is it accountable to a business result, or only to activity?

  7. Receipts. Can you audit what it did, why, and on what evidence?

  8. Eval. Does it test itself before it changes?

Three rules make the ladder hard to game. It is monotone, so rung four requires rungs one through three and nothing skips. Evidence class caps what a source can prove, so marketing copy can never establish a rung that requires observed behaviour. And evidence expires, so every row carries the date it was last checked against a dated vendor surface.

Platform by platform

Salesforce Agentforce: the strongest criteria in the cohort, and nothing that blocks

Agentforce holds test cases as version-controlled metadata, which nobody else does. The AiEvaluationDefinition type stores each case with an utterance and "a set of expectations (such as an expected action sequence)", generated as a YAML spec and extensible with custom criteria. It detects its own degradation through Agent Health Monitoring rather than waiting for a human to notice. Then it stops. The testing guide recommends adding tests to a CI system and documents nothing that blocks a deployment when they fail. And when a test does fail, its own instruction is to "use all the information you gathered to fine-tune your agent instructions, actions, or subagents." The human is named as the learning mechanism, in the vendor's own words.

Microsoft Copilot Studio: the same wall, reached from the other side

Agent evaluations reached general availability in March 2026: customizable test sets, graders measured against defined reference answers, multi-turn conversation tests, and side-by-side version comparison to spot regressions. In June 2026 it shipped memory that "captures user preferences and patterns, stores them per user, and applies them" to shape later responses. Seven months of shipping, and nothing lets performance data change what the agent generates without a person editing instructions. The one outcome-to-behaviour loop in the release notes runs inside the vendor: thumbs feedback on evaluation results "drive ongoing improvements to evaluation reliability." Microsoft's graders learn from your feedback. Your agent does not.

Jasper: the only marketing platform here that checks its own output

Brand IQ "flags brand violations and suggests on-brand replacements", which reads generated output and returns a verdict. That is more than HubSpot or Profound documents, and it moved Jasper's Eval score up against our expectation while researching this page. It flags and suggests; it does not block. Jasper also markets a brand score, and its own page discloses no methodology, no metrics and no thresholds, and the score measures conformance to guidelines you wrote rather than distance from the machine default. Those are different questions, and only the second one predicts whether a machine can tell you apart.

HubSpot Breeze: ahead of the field on the criterion that costs money

HubSpot is the only vendor in this cohort whose revenue is tied to your outcome: $0.50 per resolved conversation, down from $1.00 per conversation, and $1 per qualified lead, effective 14 April 2026. That is real commercial accountability and no competitor here matches it. Its AEO product tracks "brand mentions, prompt coverage, and citation patterns over time" across ChatGPT, Gemini and Perplexity. And the loop still ends at a person: you "take action based on recommendations, including automatically generating content with AI that you can review before publishing", with the documentation adding that manual review is strongly recommended. Charging for outcomes is not the same as learning from them.

Profound: knows which page is decaying, cannot act on it

Profound genuinely upgraded during this scoring window. Citation Decay shipped on 14 August 2026, letting you "sort your owned pages by half-life" and filter by domain or model, which attaches a citation outcome to the specific owned page rather than to the brand. Then five changelog entries from 24 July to 21 August contain no path from that outcome back into what the next brief says. Profound now knows exactly which page is dying and still cannot make the next brief smarter because of it. Independently, a tester trying to build "brief-to-publish gated on AEO score" found the gate absent rather than undocumented.

Why Ivanooo is not on this board

Ivanooo built the rubric, so Ivanooo does not appear in the ranking. You cannot referee a race you run in, and that is the exact failure this page opens by naming: the highest-authority rival board was written by a company that placed itself first on it.

Being off the board is not the same as being exempt from the standard. The instrument itself scores us, on the same eight ladders, and we do not pass. We fail our own Distinctiveness requirement in public, because it demands measuring against two poles and our category pole does not reliably execute. An audit on 27 August 2026 returned "competitor corpora too thin" and produced no category reading. We are on the wrong side of the same Evolution wall as every platform above: our outcome memory holds rules with sample sizes and confidence bands, and its liveness check failed on its first real run. Wired is not learning.

Two rungs we genuinely hold were also scored down, because their evidence sits in a private repository and a reader cannot check it. Our full row, including every failure, is published on the Agentic Compounding Index. Read it before you trust anything on this page.

What this page is not claiming

Thirty registered platforms are still unscored, and unscored means unscored, never absent. Two reference frameworks, the Claude Agent SDK and LangGraph, were measured for calibration and are deliberately not ranked against products, because a framework hands you primitives and a product hands you a system. The panel behind this page was captured through web search rather than through a live answer engine, so it establishes who competes and what they claim, and it does not establish what ChatGPT or Google AI Mode cites. Ivanooo is scored on the same instrument and is deliberately absent from the ranking above, for the reason given in the previous section.

Questions buyers ask

What makes a platform "agentic" rather than automated? Running a sequence a human authored is automation; choosing the sequence at runtime is agency. Gartner's term for the gap between the two is agent washing, which it defines as rebranding assistants, RPA and chatbots as agents. Gartner estimates only about 130 of the thousands of agentic AI vendors are real and predicts over 40% of agentic AI projects will be cancelled by the end of 2027.

Is a grade of 2 out of 5 bad? It is where the category is. The ladder tops out at behaviour no vendor here has built, which is the point of publishing it. A platform grading 2 can be excellent to use. It just is not an asset that improves on its own.

Why is the grade the weakest dimension instead of an average? An average hides the break. A platform with perfect memory and no evaluation will confidently repeat a mistake forever, and the average will look fine while it does.

Why does HubSpot grade 0 when it is ahead on outcomes? Because the ladder is monotone and it breaks at the first rung it misses. Being the only vendor whose revenue depends on your result is genuinely the strongest commercial position on this board, and it does not substitute for a check that reads the output.

How do I check a claim on this page myself? Every call links the vendor's own dated surface. Open the changelog or the documentation, search it for the mechanism, and see whether it is described or merely asserted. That is the whole method, and it takes about ten minutes per vendor.

What should I ask a vendor in a demo? Ask where the loop closes: does the system store an outcome against the thing that caused it, does it read a rule derived from that outcome before it generates the next thing, and can anything stop a bad output without a person noticing? Ask for the release note, not the answer.

The question under all of this

Every platform on this board can tell you what happened. The category has built the scoreboard and not the loop. Given the chance to build a rule that refuses, four independent vendors each built a person who approves instead. A person can be tired, rushed, or agreeable on a Friday afternoon. A written standard cannot. That gap is why AI gives such generic answers to buyers researching your category, and why the average voice keeps winning the output no one checked.

At Ivanooo, Firoz Azees built this rubric to answer a buying question we could not answer for ourselves, and then published the instrument with our own failing row on it. The full instrument, including every ladder and the anti-gaming rules, is at the Agentic Compounding Index, and the practice it belongs to is Growth Engineering.

Want to know where your own brand sits before you buy any of these? Run the free AI visibility audit and see what the machines say about you today.