top of page

Stop Prompting AI and Start Directing It

Google News has surfaced an MIT Sloan Management Review argument with a direct challenge: stop treating AI prompts as the center of effective work. The headline, “Stop Prompting AI. Start Directing It,” signals a shift from composing isolated instructions to managing an ongoing production system.

That distinction matters because business AI is moving beyond chat. Models can now search files, call tools, write code, revise documents, and continue work across multiple steps. A polished prompt can start that process, but it cannot define every decision the system will encounter.

The emerging contest is therefore not humans versus AI. It is one-shot prompting versus structured direction. The first approach asks a model for an answer. The second gives it a goal, relevant context, operating boundaries, review points, and a measurable definition of success.

This is also why the MIT Sloan headline is more than another prompt-writing lesson. It reframes the human role. The scarce skill is becoming judgment: deciding what should happen, what evidence counts, where autonomy should stop, and whether the result deserves approval.

What the Google News Headline Actually Changes

The important change is not a new prompt formula. It is a new unit of work.

The Google News listing points readers toward a management question hiding beneath years of prompt advice. If AI participates in real work, should people manage individual responses or direct the entire process?

Prompting usually treats each interaction as a small transaction. A user supplies a request, the model generates something, and the user either accepts it or tries again. Context often disappears between sessions, while standards remain buried in someone’s memory.

Directing treats the assignment as a managed workflow. The user defines the outcome, supplies evidence, identifies constraints, assigns tools, and establishes checkpoints. The model still receives prompts, but those prompts become parts of a larger operating design.

Consider a market analysis request. A prompt-centered user might ask an AI assistant to summarize a market and recommend a strategy. The model must infer the audience, geography, time horizon, evidence threshold, and acceptable level of uncertainty.

A director would first define the decision that the analysis must support. The AI might then receive approved sources, customer notes, prior research, and a required date range. It would also receive explicit instructions for handling conflicting evidence.

The director could require the system to separate verified facts from estimates. The workflow might pause before making recommendations, allowing a human to inspect the evidence. Final approval would remain with someone accountable for the business decision.

This approach changes how teams evaluate AI performance. A fluent response is no longer enough. Teams must ask whether the system used the right information, followed the intended process, exposed uncertainty, and produced a usable result.

The distinction also changes failure diagnosis. When a weak answer follows a vague request, rewriting the prompt can help. However, repeated failures often reveal missing context, unclear ownership, poor source selection, or an absent review standard.

Those are management and system-design problems. Treating them as wording problems sends teams into endless prompt tuning. They adjust adjectives and formatting instructions while leaving the underlying assignment undefined.

The phrase “start directing” therefore challenges the idea that expertise lives inside a perfect block of text. Direction is distributed across task design, source selection, permissions, memory, tools, evaluation, and human intervention.

That does not make prompting irrelevant. Every directed workflow still needs clear communication. The reversal is that prompting becomes one control surface among several, rather than the complete method for obtaining dependable work.

This framing also gives managers a more realistic way to discuss adoption. Employees do not need to become literary specialists who discover magic phrases. They need to learn how to delegate work while preserving context, standards, and accountability.

The same principle already governs effective human teams. A manager would not assign a sensitive project using one sentence and then disappear. The manager would explain the objective, share background, define boundaries, schedule reviews, and examine the result.

AI systems require a similar structure, although their failure modes differ. They can generate convincing errors, follow malicious embedded instructions, or pursue an incomplete goal with surprising persistence. Direction must therefore include technical controls alongside ordinary managerial judgment.

That is the practical change behind the Google News headline. The conversation is moving from “What should I type?” toward “What system should govern this work?”

Why One-Shot Prompting Is Reaching Its Limit

A single prompt cannot reliably carry the context, controls, and evaluation criteria required by consequential work.

Prompt engineering became popular for understandable reasons. Early generative AI products presented a blank text box, so users naturally focused on what they placed inside it. Small wording changes could produce visibly different answers.

That model still works for bounded, reversible tasks. A user can ask for alternative headlines, a rough meeting agenda, or a first-pass rewrite. The cost of inspecting and correcting the output remains low.

Problems grow when the task spans multiple sources or actions. A model preparing a research brief must decide where to search, which evidence to trust, and when it has gathered enough information. A model editing software must inspect files, run tests, interpret failures, and avoid unrelated changes.

No opening prompt can anticipate every state encountered during that work. The system needs feedback from tools and the ability to adjust its next action. It also needs limits that prevent a mistaken interpretation from spreading across the process.

This is where agentic AI enters the story. An agent is a model-based system that can select actions and use tools while pursuing an objective. Instead of producing one response, it operates through a repeated cycle of planning, acting, observing, and revising.

Anthropic draws a useful distinction between workflows and agents in its guidance on effective AI agents. Workflows follow predefined code paths, while agents dynamically choose parts of their process and tool use.

That distinction prevents “directing” from becoming a slogan for maximum autonomy. A predictable workflow often suits stable tasks better. An agent becomes useful when the path cannot be fully specified in advance and the system must respond to changing information.

Teams therefore need to select an operating pattern before writing prompts. A fixed extraction process might use a workflow with validation rules. An open-ended investigation might use an agent that searches, compares sources, and requests clarification.

This choice affects cost, speed, and risk. More autonomy can handle broader assignments, but it can also create longer execution paths and less predictable behavior. A director must decide whether that tradeoff serves the task.

The limitations of isolated prompting also appear in ordinary knowledge work. Employees frequently need information distributed across email, documents, meetings, and personal notes. A model without that context must guess or request repeated explanations.

Persistent context changes the interaction. Instead of pasting the same background into every conversation, workers can maintain an approved body of information for the system. The quality of that information then becomes a management concern.

A personal knowledge base can support this approach by keeping source material available for retrieval. However, storage alone does not guarantee accuracy. Users must still decide what belongs in the working context and what should remain excluded.

Evaluation presents another limit. A prompt can describe the desired output, but complex assignments need tests that operate outside the model’s prose. Code should run. Citations should resolve. Calculations should be reproduced. Claims should match the underlying documents.

Without those checks, the user evaluates style more easily than substance. A confident answer can feel complete before anyone tests its factual foundation. Directing shifts attention from presentation toward verification.

Microsoft Research found measurable benefits from workplace AI, but the results also show why deployment design matters. A six-month randomized field experiment involved 6,000 workers, with half receiving a generative AI tool integrated into email, documents, and meetings.

Workers who used the tool spent three fewer hours on email each week and appeared to complete documents moderately faster. Yet meeting time did not significantly change, according to the work-pattern study.

That pattern suggests AI can improve activities that individuals control more readily than work requiring broader coordination. A better response generator does not automatically redesign approvals, meetings, team dependencies, or decision rights.

Direction addresses that gap. Managers must decide how AI-generated work moves through the organization. They must specify who reviews it, where it enters existing systems, and whether colleagues can trace its evidence.

A prompt can request a summary. It cannot, by itself, establish who owns the resulting decision. It cannot determine whether a team should retain the source material or disclose AI involvement.

Those surrounding choices decide whether AI becomes a useful component or another layer of untracked output. One-shot prompting reaches its limit precisely where organizational work begins.

The New Skill Is AI Direction, Not Prompt Performance

Effective AI direction combines delegation, context management, verification, and timely intervention.

Directing begins with a clearer objective. “Analyze these customer interviews” remains too open because it does not identify the decision. A useful objective might ask the system to identify onboarding obstacles that should influence the next product release.

The objective must be paired with completion criteria. The system might need to identify recurring themes, preserve contradictory feedback, quote the underlying interviews, and distinguish observations from proposed actions.

This approach turns quality into something observable. Reviewers can determine whether the AI met the assignment instead of debating whether the output simply sounds insightful.

Context comes next. Teams should provide the smallest relevant evidence set rather than dumping every available document into a model. Large collections can contain outdated policies, duplicated notes, unrelated conversations, and conflicting terminology.

Context engineering is the practice of selecting and organizing information available to a model during a task. The goal is not maximum volume. It is enough reliable information for the model to make the next decision correctly.

People remain responsible for the boundaries of that context. Sensitive information may require exclusion or restricted processing. Old documents may need labels, while uncertain claims should remain visibly uncertain.

A director also divides work into stages. Research, synthesis, recommendation, and publication should not collapse into one opaque action. Each stage produces an artifact that a person or automated check can inspect.

For example, a directed article workflow might first collect sources. A second stage could extract claims and attach evidence. A third could remove duplicates, while a fourth drafts only from the verified material.

The publication stage would remain blocked until citation checks and editorial review succeed. Each checkpoint contains errors before they become public.

Tool permissions provide another layer of direction. Reading a calendar carries less immediate risk than sending invitations. Drafting an email differs from transmitting it. Querying a database differs from modifying records.

Anthropic describes permission choices that include always allowing an action, requiring approval, or blocking it. Its discussion of trustworthy agents also explains why reviewing an overall plan can be more meaningful than approving dozens of individual actions.

Repeated low-level approval requests create their own danger. Users can become fatigued and approve actions without examining them. Direction should place human attention at decisions where judgment changes the outcome.

That means humans do not need to inspect every token or tool call. They should review the strategy, irreversible actions, uncertain evidence, and final result. Low-risk intermediate steps can proceed within defined boundaries.

The required level of control depends on consequence. A brainstorming assistant can operate with broad freedom because rejection is easy. A system handling customer records, legal material, or financial decisions needs narrower permissions and stronger evidence.

Directors must also know when to interrupt. A model that repeatedly revises the wrong artifact does not need a cleverer follow-up prompt. It needs the objective, context, or workflow corrected.

Useful intervention signals include unsupported conclusions, unexplained source conflicts, tool failures, and changes outside the stated scope. Teams can encode some signals as automated checks while reserving ambiguous cases for human review.

Memory adds another management choice. Persistent instructions can reduce repetitive prompting, but they can also preserve bad assumptions. Teams need a way to inspect, update, and retire instructions that shape future work.

This is where an AI second brain becomes relevant to direction. Its value lies not only in remembering information, but in giving users a reviewable context for future assignments.

Direction also requires role clarity. The AI may research and draft, while a human validates evidence and accepts responsibility. Another human might approve publication or authorize an external action.

These roles should not blur merely because the model’s output looks polished. A system can produce executive language without possessing executive authority. Fluency does not grant accountability.

Teams should document these operating choices in reusable instructions, templates, and evaluation rules. That makes good performance less dependent on one employee remembering the right phrase.

It also makes improvement measurable. If an output fails, the team can locate the problem in source collection, task definition, execution, review, or approval. Prompt wording remains one possible cause, but no longer the default explanation.

AI direction is therefore closer to editing, product management, and operations than to discovering secret commands. It asks people to translate intent into a controlled process and to recognize when reality departs from the plan.

That responsibility becomes more important as systems perform longer assignments. The less frequently a human intervenes, the more carefully the objective and boundaries must be designed in advance.

Direction Does Not Eliminate AI’s Reliability Problem

Human direction improves control, but it does not make uncertain models dependable by declaration.

The management metaphor has an attractive simplicity. Give the AI a clear mission, inspect its plan, and evaluate the result. Yet AI systems are not employees, and treating them as people can conceal important technical differences.

A model does not possess stable judgment simply because a user assigns it a role. It predicts and selects outputs from learned patterns, current context, system rules, and tool feedback. Its behavior can change when any of those inputs change.

The system can also generate false statements with confident wording. In an agentic workflow, one false assumption can affect later searches, calculations, code changes, or recommendations. Longer execution creates more opportunities for errors to compound.

NIST’s generative AI profile identifies confabulation, information security, privacy, harmful bias, and human-AI configuration among the concerns organizations should manage.

NIST recommends policies that define roles and responsibilities for human oversight. It also calls for testing and evaluation proportional to the identified risks. These safeguards belong around the workflow, not inside a hopeful final prompt.

Prompt injection presents another challenge. Malicious or irrelevant instructions can appear inside websites, emails, or documents that an agent reads. If the system treats that content as authoritative, it might disclose information or take an unintended action.

A director must therefore separate trusted operating instructions from untrusted source material. Tools should enforce permissions independently of what the model reads. Sensitive actions should require explicit approval or remain unavailable.

Direction can also fail through automation bias, the tendency to accept computer output too readily. Better interfaces may make AI work easier to supervise, but polished plans and citations can still create false confidence.

A Microsoft Research survey of 319 knowledge workers collected 936 examples of generative AI use. Higher confidence in AI was associated with less critical thinking, while greater confidence in personal expertise was associated with more.

The critical-thinking study found that AI shifted critical work toward verification, integration, and task stewardship. That finding supports the directing model while exposing its central weakness.

People must retain enough domain knowledge to recognize a bad result. If organizations automate entry-level work without preserving learning opportunities, they may reduce the future supply of experienced reviewers.

The risk is especially visible when AI produces plausible first drafts. Junior employees can complete assignments faster, but they might engage less deeply with the evidence. Senior reviewers then inherit more checking work across a larger volume of output.

Productivity metrics can miss that transfer. A team may count documents completed while ignoring review time, error correction, or downstream confusion. Direction must include metrics for total process performance, not only generation speed.

Organizations should track acceptance rates, correction rates, unsupported claims, security incidents, and reviewer effort. They should also compare outcomes against a non-AI baseline where practical.

The right measure depends on the task. Software teams can use tests, defects, and review findings. Research teams can measure citation accuracy and claim coverage. Customer operations can track resolution quality and escalation rates.

Another uncertainty concerns autonomy. More experienced users may allow agents to operate longer because they understand the tools. They may also become comfortable enough to approve risky behavior too quickly.

Anthropic’s 2026 analysis of agent autonomy found that experienced Claude Code users enabled full auto-approval more frequently but also interrupted the system more often. That combination resembles active supervision rather than total trust.

The same research reported that software engineering accounted for nearly half of observed agentic activity. Most recorded actions were low-risk and reversible, while higher-stakes use remained an emerging category.

Those findings should temper broad claims about autonomous digital coworkers. The strongest evidence still comes from tasks with inspectable artifacts, reversible changes, and clear evaluation methods.

Business leaders should avoid translating success in coding or document drafting directly into confidence about hiring, medical, legal, or credit decisions. Different consequences require different control structures.

Direction also introduces overhead. Designing context, permissions, tests, and checkpoints takes time. For a small, reversible assignment, that effort can exceed the value of automation.

The sensible principle is proportional control. Use simple prompting when the task is low-risk and easily checked. Use predefined workflows when consistency matters. Reserve greater autonomy for assignments where flexible reasoning justifies the added uncertainty.

The MIT Sloan framing should therefore be read as a change in responsibility, not a guarantee of reliability. Directing AI asks humans to build stronger systems around imperfect models. It does not remove the models’ imperfections.

What Managers and Knowledge Workers Should Watch Next

The directing thesis will succeed only if organizations can show better outcomes without hiding more risk or review labor.

The first signal to watch is the movement of AI interfaces from chat boxes toward persistent workspaces. Products increasingly maintain files, instructions, project history, and tool connections across sessions.

That shift strengthens the case for direction because users can shape an operating environment instead of rebuilding context repeatedly. The important question is whether those environments remain transparent and editable.

Workers should be able to see which instructions influenced a result. They should know what sources the system accessed, what tools it used, and which assumptions persisted from earlier sessions.

If products expose that information clearly, direction becomes easier to audit. If context remains hidden, persistent memory can turn old errors into invisible defaults.

The second signal is evidence from deployed workflows rather than demonstrations. Vendors can show an agent completing an impressive task under carefully selected conditions. Enterprises need results across ordinary work, varied users, and imperfect data.

Watch for controlled studies that report total task time, output quality, review burden, and error rates. Productivity numbers without quality measures offer an incomplete picture.

The Microsoft workplace experiment provides a useful model because it compared behavior over six months and separated individual activities from coordination-heavy work. Future research should examine whether directed AI changes decisions, approvals, and team processes.

Evidence should also distinguish workflows from agents. A company may describe any automated AI feature as agentic, even when it follows a fixed sequence. Buyers need to know where the model chooses actions and where software enforces the path.

That architecture affects predictability. A fixed workflow can provide consistency, while a dynamic agent can adapt when unexpected conditions appear. Neither approach is universally superior.

The third signal is whether oversight moves from repetitive approval toward risk-based control. Asking a user to approve every small action creates friction without guaranteeing attention. Granting unrestricted access creates the opposite problem.

Better systems will classify actions by consequence and reversibility. Reading an approved folder might proceed automatically, while changing a customer record requires review. External publication should remain distinct from internal drafting.

Organizations should test whether these controls work under pressure. Employees may bypass slow review systems when deadlines tighten. Managers may also expand permissions after early successes without updating monitoring.

The next one to three months should reveal how vendors package plan review, persistent context, permission layers, and post-execution logs. Product announcements will matter less than the clarity of those controls.

Knowledge workers should conduct a smaller experiment now. Choose one recurring assignment that has clear inputs, a reviewable output, and low consequences if the first attempt fails.

Write the outcome before writing the prompt. Identify approved sources, completion criteria, and actions the AI cannot take. Decide where a human must review the work and what evidence the reviewer needs.

Then compare the process with ordinary prompting. Count revisions, unsupported claims, review time, and usable output. The goal is not to prove that AI works. It is to discover which operating design produces accountable results.

Google News may have introduced the idea through a sharp MIT Sloan headline, but the lasting question belongs to every AI user. Are you repeatedly asking a model for answers, or are you building a process that preserves your judgment?

Choose one real workflow, define its boundaries, and test that difference. If directing reduces rework while keeping errors visible, the headline has identified a durable management shift.

Get started for free

A local first AI Assistant w/ Personal Knowledge Management

remio only supports Windows 10+ (x64) and M-Chip Macs currently.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page