All articles
Context Engineering

Context Engineering for Product Managers: Decide What Your AI Agents See

Context engineering for product managers: decide what your AI agents see, from discovery findings to acceptance criteria, with write, select, compress, isolate.

Context Engineering for Product Managers: Decide What Your AI Agents See

Product teams now hand real work to AI agents: research summaries, draft specs, user stories, code. When the output is wrong, the instinct is to rewrite the prompt. Most of the time the prompt was fine. The agent was looking at the wrong material, too much material, or material that contradicted itself.

The fix is a design discipline called context engineering: deciding, on purpose, what information goes into the model's context for each task, in what form, and what stays out. For a product manager this is not a technical side quest. You own the product decisions, the requirements and the acceptance criteria, which makes you the person best placed to decide what an agent should see.

What Is Context Engineering?

Context engineering is the practice of designing everything a model sees when it produces an answer: instructions, product context, requirements, examples, tool results and memory. Anthropic's engineering team describes it in Effective context engineering for AI agents as the set of strategies for "curating and maintaining the optimal set of tokens (information) during LLM inference."

The term caught on in mid 2025. Simon Willison collected the early definitions, including Andrej Karpathy's description of it as "the delicate art and science of filling the context window with just the right information." The important phrase is "just the right." Context engineering is as much about exclusion as inclusion.

How Is It Different From Prompt Engineering?

Prompt engineering is about how you phrase an instruction. Context engineering is about the whole information package that surrounds that instruction, across many turns and many tasks.

A useful way to see the split:

  • Prompt engineering asks: how should I word this request so the model understands it?
  • Context engineering asks: what does the model need to know, from which sources, at what level of detail, to do this task correctly, and what should it never see?

A perfectly worded prompt sitting on top of a stale requirement still produces the stale behavior. We cover phrasing and structure for PMs in our guide to prompt engineering for product managers. This article is about the layer underneath.

How Is It Different From Context Management?

The two terms get mixed up, so it helps to separate them by time horizon.

  • Context engineering is design work. You decide the shape of the package: which artifacts exist, which ones each kind of task receives, how they are summarized, and where they live.
  • Context management is the day-to-day operation of that design. You keep the right files loaded, refresh what changed and clear what went stale during a working session.

If you build with coding agents, our existing guide to context management for AI coding agents covers the operational side in depth. Here we stay at the design level, where product managers have the most influence.

Why the Context Window Is a Budget, Not a Bin

Every model has a context window, the total amount of text it can reference while generating a response. Claude's documentation on how the context window works defines it as all the text the model can reference "including the response itself," and lists what counts against it: the system prompt, every message, tool results, documents and tool definitions.

Windows are now large. Several current models accept up to 1M tokens. That tempts teams to treat the window as a bin: throw in the whole PRD, every interview transcript and the full backlog, and let the model sort it out. The same documentation warns against this directly: "more context isn't automatically better. As token count grows, accuracy and recall degrade, a phenomenon known as context rot."

Treat the window as a budget. Every token you add competes for the model's attention with every other token. Anthropic's guidance frames the goal as finding "the smallest possible set of high-signal tokens that maximize the likelihood of some desired outcome."

How Context Goes Wrong: Rot, Poisoning and Clash

Drew Breunig's essay on how long contexts fail and how to fix them names four failure patterns that map neatly onto product work:

  • Context poisoning: an error enters the context and keeps getting referenced. In product terms, a hallucinated requirement in an early draft gets quoted by every later task.
  • Context distraction: the context grows so long that the model leans on it and ignores what it already knows. Think of an agent copying the structure of an old spec instead of reasoning about the new one.
  • Context confusion: irrelevant material shapes the answer. Twenty unrelated user stories in the window nudge the agent toward features nobody asked for.
  • Context clash: two pieces of context contradict each other. Version 2 of the pricing rule and version 3 both sit in the window, and the agent picks one.

All four are product problems as much as model problems. A PM who knows which decision is current, which research is relevant and which draft was rejected can prevent most of them before the agent runs.

The Four Core Strategies: Write, Select, Compress, Isolate

LangChain's post on context engineering for agents groups the techniques into four strategies. Here is what each one means for a product team.

Write: save context outside the window

Writing means storing information outside the context window so the agent can use it later. For product work, this is your durable record: decision logs, approved specs, research summaries and agent memory that persists between sessions. If a decision only lives in a chat thread, it will be lost the next time that thread is cleared.

Select: pull in only what the task needs

Selecting means retrieving the relevant pieces into the window at the moment of need. Anthropic calls this "just in time" retrieval: agents keep lightweight references, such as file paths or links, and load the content when it matters. For a PM, selection is a mapping job: which artifacts belong to which kind of task.

Compress: keep the tokens that carry the decision

Compressing means retaining only the tokens required for the task. Summaries of interviews, a one-page decision brief instead of a 40-page discovery deck, and context compaction of long sessions all fall here. Anthropic notes that good compaction preserves things like "architectural decisions, unresolved bugs, and implementation details" while discarding the rest.

Isolate: split work so each agent sees less

Isolating means splitting context across separate agents or sessions. A research agent can read fifty sources and return a short, distilled summary to the agent writing the spec. Anthropic describes sub-agents returning condensed summaries "often 1,000-2,000 tokens," so the main agent never carries the raw research.

What Counts as Product Context?

Product context is the subset of information that describes what you are building and why. For agents, it usually includes:

  • The problem statement and target users
  • Validated discovery findings, with their sources
  • Strategic decisions and the alternatives that were rejected
  • Scope boundaries, including what is explicitly out of scope
  • Requirements, business rules and edge cases
  • Acceptance criteria that define done
  • Constraints: technical, legal, budget and timeline

Notice what is missing: raw transcripts, brainstorm notes, abandoned drafts and long chat histories. Those can be useful inputs to a human, but they are poor inputs to an agent unless someone has compressed them first.

What Should Agents See at Each Product Stage?

The right package changes as work moves from discovery to delivery. A simple mapping keeps you from dumping everything into every task.

StageGive the agentKeep out
DiscoveryResearch questions, interview summaries with sources, known assumptionsSolution ideas presented as facts
StrategyValidated findings, market evidence, goals and constraintsRaw transcripts, unranked idea lists
RequirementsApproved decisions, scope, business rules, edge casesRejected options without a "rejected" label
BacklogThe relevant requirement, acceptance criteria, dependenciesThe full PRD for every single ticket
BuildOne ticket, its acceptance criteria, related code and conventionsStrategy debates, unrelated tickets

The pattern is consistent: as work gets closer to code, context gets narrower and more precise. Our guide to writing PRDs an AI coding agent won't misread covers how to make the requirements layer precise enough to hand over.

How to Package Decisions So Agents Can Use Them

Decisions are the most valuable and most fragile part of product context. An agent that knows "we chose annual-only billing for the Team plan" but not why will happily reintroduce monthly billing the first time a ticket mentions it.

Package each decision with four fields:

  1. The decision, in one sentence.
  2. The reason, with a link to the evidence.
  3. What was rejected, so the agent does not propose it again.
  4. Status and date, so a superseded decision is clearly marked as superseded.

The status field is your main defense against context clash. If two versions of a rule exist, only one should be marked current, and only that one should be selected into an agent's window.

How to Package Specs and Acceptance Criteria

Specs and acceptance criteria are where context engineering pays off most, because they are what the agent acts on.

  • Make each requirement addressable. A short ID lets a ticket reference exactly one rule instead of pasting a whole section.
  • Keep acceptance criteria testable. "Fast checkout" is not usable context. "Checkout completes in three steps or fewer for a returning user" is.
  • Separate rules from rationale. Put the rule where the agent will read it and link the rationale for anyone who needs it.
  • Attach the acceptance criteria to the ticket. Do not rely on the agent finding them in a separate document.

When backlogs are built this way, each ticket becomes a self-contained context package. Our guide to structuring backlogs for AI software agents walks through the ticket format.

A Worked Example: Adding Seat Limits to a Team Plan

Consider a product manager at a small B2B SaaS company. The team is adding seat limits to its Team plan, and she plans to use three agents: one to draft the requirement, one to write user stories, and one coding agent to implement the change.

The first attempt. She pastes the full discovery notes, two earlier pricing drafts, a Slack export and the current PRD into one long session and asks for user stories. The stories come back mixing a five-seat limit from an old draft with a ten-seat limit from the current one. One story describes a "seat waitlist" that a customer once mentioned in an interview but the team rejected. That is context clash and context confusion in one output.

The redesign. She applies the four strategies.

  • Write. She records the decision once: "Team plan includes 10 seats. Additional seats are sold individually. Waitlist rejected: adds support load with no evidence of demand." The record has a status of current and a date.
  • Compress. She replaces the discovery notes with a half-page summary of the three findings that drove the decision, each linked to its source interview.
  • Select. For the story-writing agent, she provides only the decision record, the summary, the relevant PRD section and the acceptance criteria template. The old drafts stay out.
  • Isolate. The coding agent receives one story at a time with its acceptance criteria, for example: "Given a Team workspace with 10 active members, when the owner invites an 11th, then the invite is blocked and the owner sees the option to add a seat." It never sees the pricing debate.

The result. The stories now agree with each other, the waitlist is gone, and each coding task is small enough that review is quick. Nothing about the prompts changed much. What changed was what each agent was allowed to see.

How to Measure Whether Your Context Design Works

You do not need complex tooling to judge context quality. Watch for a few signals in review:

  • Contradictions between outputs. Two agents disagreeing on a rule usually means two versions of the rule are reachable.
  • Reintroduced rejected ideas. A sign that rejected options are in context without a clear label.
  • Long, generic outputs. Often a sign of distraction from oversized context.
  • Questions the agent should not need to ask. A sign that selection missed something the task needed.

Each signal points back to one of the four strategies, which tells you where to adjust.

Common Mistakes Product Teams Make

  • Treating a bigger window as a solution. A larger window raises the ceiling, not the quality of what you put under it.
  • Keeping decisions only in chat. Chats get cleared, compacted or forgotten. Decisions need a written home.
  • Passing the whole PRD to every ticket. The agent will act on sections that have nothing to do with the task.
  • Never retiring old material. Superseded specs that remain reachable are the most common source of context clash.
  • Skipping sources. Research without links cannot be checked, and unchecked claims become context poisoning later.

Context Engineering Checklist for Product Managers

Use this before you hand a task to any agent:

  • The task has a clear goal and a definition of done
  • Current decisions are recorded with reason, rejected options, status and date
  • Superseded decisions and drafts are marked or removed from reach
  • Discovery findings are summarized and linked to sources
  • The agent receives only the artifacts mapped to this task type
  • Acceptance criteria are testable and attached to the task
  • Long sessions are compacted or restarted with a clean summary
  • Large research jobs run in a separate agent that returns a short summary
  • You review outputs for contradictions and reintroduced rejected ideas

How Prodstack Fits

Prodstack is an AI product management operating system built around a 7-stage method: Discovery, Strategy, Prioritization, Roadmap, Requirements, Backlog and Growth. All stages share one product memory, so cited research, decisions, PRDs, user stories and acceptance criteria are available to later stages without re-pasting them. A Traceability Log records decisions as they are made, and agent-ready backlogs can be handed off to Jira and Claude Code, with bring-your-own MCP connectors such as Linear, GitHub, Jira and Notion available to the AI coach. If you want to see how a single shared product memory supports this kind of context design, read our overview.

Context engineering is a habit before it is a tool: decide what each agent should see, write it down once, and keep the stale material out of reach. If you want a workspace that keeps your product context organized across every stage, start your 7-day free trial.

Written by
The Prodstack Team
Product management research

The product team behind Prodstack writes practical, evidence-based guides on discovery, strategy, prioritization, requirements and growth.

Product discoveryPrioritizationRequirementsGrowth
Plan before you prompt

Validate the problem, define the requirements and hand your coding agent a clear backlog, all in one product thread.

No credit card required · 300,000 tokens · 4 documents
// Newsletter
Field notes, in your inbox.

Evidence‑driven thinking on discovery, prioritization, specs, and shipping — plus new articles the moment they drop. No noise.