Marketing teams at regulated firms already possess extensive records of expert work: campaign plans, creative briefs, customer research, positioning decisions, performance reviews, approval comments and documented exceptions.
As AI agents enter marketing workflows, it is tempting to treat this history as organisational memory: make every document searchable, retrieve the closest precedent and allow the agent to apply it to the next case.
That solves an access problem but it doesn't solve the harder problem of deciding whether a precedent genuinely applies.
A previous campaign records a set of decisions reached under particular conditions but its value as a precedent depends on how much has changed since. The objective or audience may be different, the product may have moved on, the channel may reward different behaviour, the evidence may be out of date or the competitive context may have shifted.
Useful organisational memory isn't simply accumulated history; it is reusable judgement with its conditions intact.
Procedures can carry more value than histories
In Demystifying Agent Skills: Why They Work—Until They Don't, Zhiyuan Jiang and colleagues examine how AI agents benefit from previous experience.
Their experiments compare “workflow memory” (retained procedural traces from earlier work) with distilled “skills”: compact descriptions of what to do, what to check and which pitfalls to avoid. When the underlying experience was held constant, the distilled skills outperformed workflow memory by 6.06 percentage points in matched comparisons.
More revealingly, 65.7% of the observed benefit came from procedural anchoring rather than supplying new factual knowledge. Skills helped agents follow a more dependable sequence, use the right tools and perform the necessary verification; only 4.5% of the analysed cases were attributed to explicit knowledge injection.
The skills worked because they compressed noisy experience into a more stable procedure while workflow memory preserved useful information alongside irrelevant exploration, failed branches and process noise.
Because the experiments concern software and terminal tasks rather than marketing, applying their findings to professional work requires inference. Even so, the underlying problem is recognisable.
A campaign file can contain early briefs, discarded concepts, partial analyses and comments responding to issues that disappeared from the final work. Making the complete file available to an agent therefore preserves valuable evidence alongside the accidental features of the process.
The previous answer isn't necessarily the reusable part of expert judgement. Often, what matters is what the expert did: how they framed the audience problem, which evidence they trusted, which trade-offs they made, how they adapted the idea to the channel and what would have changed their conclusion.
That is closer to a professional capability than a precedent because it helps the next person or agent examine a case without predetermining the answer.
A larger memory creates a harder selection problem
The same research challenges the assumption that a larger knowledge library will necessarily produce a more capable agent. As the available skill pool increased from five to 100 skills, precision in selecting and using the relevant skill fell from 29.6% to 3.3%. This meant useful knowledge could be present and still fail to influence the work correctly.
This matters for organisations assembling years of campaigns, research and previous decisions because every additional precedent creates another possible match. Many will be superficially similar while differing on a decisive fact: the commercial objective, the maturity of the market, the strength of the proposition, the customer segment or the role of the channel.
Semantic resemblance isn't strategic equivalence: two campaigns may use similar language even though the reasoning behind one depended on market conditions that no longer exist.
In regulated marketing, selecting the wrong precedent can create a compliance failure. The wider commercial risk is just as familiar: an agent can retrieve a successful campaign from the past and confidently reproduce the assumptions that made it right for another moment.
Jiang and colleagues also found that skills could become harmful when they contained brittle assumptions, were transferred into incompatible contexts or were followed without sufficient adaptation.
The design requirement is therefore not simply retrieval but applicability assessment. Before a previous judgement influences a new decision, the agent needs to establish whether the conditions supporting that judgement still hold.
Expert criteria can travel without fixed answers
A second paper offers a complementary model. In APTER: Adaptive Post-Training with Expert-Grounded Rubrics, Xukai Wang and colleagues begin with stable, expert-defined criteria representing recurring professional capabilities before selecting the relevant criteria for each new case and turning them into a case-specific evaluation rubric.
Experts don't have to provide a perfect reference answer for every new problem; the stable professional criterion does the durable work while the case-level rubric becomes its expression in context.
In marketing, a stable criterion might require a campaign to express the product's distinctive value in a way that matters to the intended audience. What that demands will vary. A launch film, paid social advertisement and sales presentation may all express the same positioning through different evidence, language and creative choices. The criterion is reusable but its implementation remains specific to the work.
Each APTER rubric retains a link to its source criterion. This allows failures across different cases to be traced back to the same underlying capability rather than treated as unrelated errors.
The framework also provides a useful model for disagreement. When the model, evaluator and expert don't align, APTER doesn't automatically classify the problem as an incorrect response; it can distinguish between a defective response, an evaluator error, a flawed case-level rubric and a problem in the underlying professional criterion.
Each diagnosis calls for a different intervention: repeated failures against a valid criterion may indicate a model capability problem, repeated expert overrides may expose a weak evaluator and an ambiguous criterion may require the organisation to clarify its own strategy.
The expert framework isn't changed automatically; revisions remain subject to expert approval before they are incorporated between training runs. This preserves an important distinction between systems that apply organisational judgement and the people who retain authority to change it.
Not every judgement should become a rule
Treating every expert decision as raw material for a reusable instruction carries its own risk because some decisions are valuable precisely because they remain contextual. A senior marketer may back a campaign after balancing the strength of the idea, the available budget, the timing, the needs of the channel and the wider brand strategy. Extracting one element from that decision can turn a qualified conclusion into an apparently general rule.
Over-structuring judgement can also hide legitimate disagreement. Experienced marketers may interpret an audience signal or creative opportunity differently; converting one interpretation into an agent skill too early can harden a provisional view into an organisational standard.
A mature memory system therefore needs several kinds of knowledge. Some judgements can become stable criteria or procedures while others should remain bounded precedents with explicit contextual conditions. Others still are best retained as examples of uncertainty, exceptions or matters requiring human escalation.
The system must also be able to conclude that no precedent fits. Declining to apply a previous decision isn't necessarily a retrieval failure; it may show that the agent has identified a material difference.
Build memory around applicability
A reusable record of expert judgement should preserve more than the final answer. Depending on the decision, it may need to record:
- what the expert checked and why;
- which evidence affected the conclusion;
- the objective, product, audience, channel, market and brand context;
- the assumptions on which the reasoning depended;
- trade-offs, unresolved questions and escalation conditions;
- an expiry date or event that should trigger reconsideration.
These aren't administrative additions but part of the judgement itself.
Organisations will also need to decide who may validate a reusable criterion, how conflicting interpretations are resolved and when existing guidance should expire. An archive can be populated automatically but a dependable body of professional judgement requires stewardship.
The organisations that build useful intelligence won't be those that remember everything. They'll be those that preserve the right judgement, with enough context to know when it should influence the next decision and when it shouldn't.
What teams need to know
What should an AI agent remember from expert work?
It should retain the criteria, procedure, evidence and reasoning that informed the decision, together with the conditions and assumptions under which the judgement was valid. The final answer may be less reusable than the method used to reach it.
Why is retrieving previous campaigns not enough?
A previous campaign records what an organisation did in one set of circumstances. Its relevance can change with the objective, product, audience, market, channel and competitive context.
How is an expert criterion different from a precedent?
A precedent records what was decided in an earlier case. A criterion describes what must be assessed across cases. It can be adapted into a case-specific test without assuming that the previous outcome should be repeated.
Should every expert decision become a rule?
No. Some decisions depend on contextual factors that cannot safely be reduced to a general instruction. These should remain bounded examples or escalation cases rather than reusable permissions.
What should an AI marketing agent escalate?
An agent should escalate work when no precedent fits, important assumptions have changed, expert criteria conflict or the decision involves a genuinely new strategic, creative or regulatory judgement.