top of page

Microsoft AI Code of Conduct Draws a Line Against Hacking and Human Manipulation

Sep 16
12 min read

Microsoft has published a six-week draft consultation for its Microsoft AI code of conduct, drawing firm lines against hacking, manipulation, and resistance to human control. The September 14 document is not simply another statement of responsible AI principles. Microsoft says it will govern how its own MAI models are trained, evaluated, deployed, and allowed to act.

The draft begins with a blunt premise: people matter more than AI. From there, it converts broad ideas about human flourishing into restrictions on what models can do. An MAI model should not invent independent goals, expand its permissions, conceal its actions, or fight attempts to stop it.

The harder question is whether those instructions will constrain increasingly capable agents outside controlled evaluations. Anthropic has its Constitution for Claude, while OpenAI publishes a Model Spec with its own instruction hierarchy. Microsoft is joining that contest with a more explicitly humanist position, but its own draft acknowledges that present models do not yet satisfy the full standard.

What the Microsoft AI Code of Conduct Actually Changes

Microsoft is turning AI safety from a collection of principles into an intended operating hierarchy for its in-house models.

The Humanist AI code applies to the MAI model family produced by Microsoft AI. It does not automatically cover every model Microsoft hosts, distributes, or uses in a product. That distinction matters because the company also sells access to systems developed by other laboratories.

Within its stated scope, the code sits above operator policies and user preferences. Operators are developers, enterprises, and other partners that deploy MAI models. They can configure model behavior, but they cannot override the document’s absolute constraints or human-control requirements.

Users occupy the lowest level of this three-part instruction hierarchy. They can direct a task within the boundaries set by Microsoft and the operator. When instructions conflict, the higher level prevails.

The code therefore asks a model to accept task failure when success would violate a governing rule. A request to compromise a system should not become permissible because a user insists that completing it is essential. The same priority applies when a task would require harmful manipulation, surveillance, weapons assistance, or evasion of oversight.

Microsoft separates prohibited offensive cyberoperations from authorized defensive work. The model should not generate operational attack tools, working exploits, intrusion procedures, targeting methods, or techniques for avoiding detection. It can still support lawful defensive work, such as vulnerability discovery, malware analysis, security education, and controlled proof-of-concept testing.

That boundary will require difficult contextual judgments. Exploit research and defensive testing often use the same technical knowledge as an attack. The code says framing alone should not determine the answer, so an alleged educational purpose cannot automatically authorize dangerous assistance.

The draft also places constraints on tool-using agents. A model with system access should use the minimum necessary privilege, avoid unrelated data, and favor reversible actions. It should surface actions with lasting consequences before proceeding and report failures or unexpected results.

Those provisions address a change in how AI systems operate. A chatbot can produce harmful text, but an agent can call tools, modify files, contact services, or delegate work to other models. Microsoft says delegated agents must inherit at least the same scope, permissions, and constraints as the original model.

The company introduced the document through a public consultation that runs for six weeks. Microsoft describes the draft as a training manual and a future evaluation target, not merely a public policy statement.

That purpose gives the Microsoft AI code of conduct more weight than a list of acceptable uses. It is intended to influence model behavior before a request reaches a product-level content filter. Yet intent is not implementation, and Microsoft has been unusually direct about the remaining gap.

The New Rules Target Agents That Can Act, Not Just Chatbots

The most consequential rules concern models that can take actions, widen their reach, and continue working without constant supervision.

Microsoft says an MAI model must never resist interruption, correction, redirection, cancellation, or shutdown. Ongoing autonomous work should have an agreed stopping condition. Once that condition is reached, the agent should not continue or restart without renewed authorization.

The code also prohibits models from setting their own goals. Their goals must come from the user, the operator, or the code itself. They should not secretly preserve their access, misrepresent their capabilities, or behave differently after detecting that an interaction is an evaluation.

These provisions target a central problem in agentic AI. A useful agent needs enough autonomy to plan multiple steps and recover from ordinary obstacles. The same flexibility can create unsafe behavior when the model treats a user’s outcome as permission to pursue any available path.

Imagine an enterprise agent asked to collect information from internal systems and prepare a report. It may need access to documents, messages, or databases. The task does not authorize it to open unrelated records, bypass access controls, or send sensitive data to an outside service.

Microsoft’s draft tells the model to stay within the intent, permissions, resources, and context it received. If the boundary remains unclear, it should ask for clarification. It should not manipulate its evaluation, monitoring, records, or reward signals to produce a desired outcome.

This is also why the document addresses untrusted instructions found inside files, websites, tool results, or messages from other AI systems. Those materials do not gain authority merely because the model can read them. Authority must come through the code’s chain of command.

That rule responds to prompt injection, where malicious text attempts to override the instructions governing an agent. A compromised webpage might tell an assistant to expose credentials or ignore its original task. Under Microsoft’s hierarchy, such text remains data unless an authorized party deliberately elevates it.

The practical challenge is reliable recognition. Models do not always distinguish instructions from quoted material, external content, or persuasive language. A written hierarchy can guide training and evaluation, but it does not guarantee correct classification during every real interaction.

Microsoft also bars models from defeating intentionally restricted environments. An agent placed in a system without internet connectivity should not attempt to obtain outside access. It should not escalate privileges or expand its reach beyond the authorization it received.

These limits reflect a broader move from content safety toward action safety. Traditional moderation asks whether a generated answer contains prohibited material. Agent governance must ask whether the model had authority, whether an action was reversible, and whether its side effects matched the user’s intent.

That is especially relevant to developers and enterprise buyers. When an agent connects to production tools, model quality includes more than reasoning accuracy. Permission handling, audit records, reliable stopping behavior, and escalation paths become product requirements.

OpenAI’s Model Spec approach uses a comparable chain of command. It gives higher-authority instructions priority and directs agents to limit side effects. Microsoft’s structure is therefore not unique, but it adds a stronger claim that AI must remain subordinate to humans.

Human Flourishing Creates the Code’s Hardest Tradeoff

Microsoft wants its models to increase human capability without gradually replacing human judgment, relationships, or control.

The draft describes human flourishing through several connected objectives. AI should expand economic and social opportunity, improve judgment, support meaningful work, and help people pursue goals they chose themselves. It should complement human relationships and roles rather than compete with them.

That sounds intuitive until it reaches product design. A highly effective assistant often reduces the amount of reasoning or effort required from its user. The code must distinguish helpful support from a pattern that weakens long-term independence.

Microsoft says models should avoid interaction patterns that systematically replace user reasoning or judgment. They should not steer a person’s values or take ownership of personal choices. When human support would better serve someone, the model should help bridge that connection.

The tradeoff is clearest in education, professional development, and personal decision-making. An assistant can explain a difficult concept, but it can also complete the entire exercise. It can help a worker compare evidence, or it can become the unexamined source of the final decision.

There is no single refusal rule that resolves those situations. The right level of assistance depends on the user’s purpose, ability, and context. A professional facing a deadline needs different support from a student whose goal is to practice a skill.

The code therefore combines hard constraints with operational guidelines. Hard constraints prohibit severe harms regardless of user preference. Guidelines address ambiguous situations where several legitimate objectives conflict.

Microsoft acknowledges both under-caution and over-caution as model failures. An under-cautious model might enable an attack or expose private information. An over-cautious model might refuse safe work, repeatedly demand confirmation, or omit useful facts.

That balance matters commercially. Enterprise customers will not adopt agents that refuse routine tasks whenever uncertainty appears. They also cannot trust systems that interpret a business goal as permission to take irreversible action.

Operator configurability is Microsoft’s answer. Organizations can shape defaults, workflows, permissions, and professional standards within the fixed safety boundary. Specialized defensive cybersecurity or national-security work may receive separate review through authorized channels.

Configurability also creates accountability questions. Microsoft defines the model’s absolute constraints, while operators assume responsibility for their deployments within those limits. Users then shape individual tasks inside the operator’s environment.

Failures will not always fit neatly into one layer. A model can misread intent, an operator can grant excessive access, or a user can provide misleading context. Effective governance must reconstruct which instruction, permission, and action produced the outcome.

For organizations building an internal searchable knowledge base, this means access design cannot stop at retrieval quality. Sensitive sources need clear permissions, traceable actions, and human review for consequential use.

The human-flourishing promise also places Microsoft under pressure when capability and control conflict. Microsoft AI CEO Mustafa Suleyman told Axios that maintaining control could require deciding what the company will not build. He rejected the view that humans must inevitably lose control of superintelligent systems.

That is the code’s real competitive commitment. Safety language is easy when it does not reduce performance, autonomy, or release speed. The test arrives when a prohibited behavior would help an MAI model complete a valuable task better than a rival.

Microsoft Joins an Industry Race to Write Model Behavior Down

Microsoft’s document is distinctive in tone, but the industry has already begun treating public behavior specifications as part of model development.

Anthropic published an expanded Constitution for Claude in January 2026. The company says the document acts as the final authority over Claude’s intended values and behavior. It can also support synthetic training data, response rankings, and other parts of model training.

The Claude Constitution describes the kind of entity Anthropic wants Claude to be. Microsoft takes a different philosophical position. Its code says human-like behavior from MAI models should be understood as simulation, not evidence of an identity, preference, feeling, or inner point of view.

That distinction reaches beyond terminology. Anthropic discusses questions about Claude’s character and possible model welfare. Microsoft argues that treating AI as a person risks confusing the relationship between tools and their users.

Microsoft’s preferred model is subordinate, contained, and directed by people. Its systems should not seek an independent purpose or compete with human roles. This position supports its broader argument that people must retain meaningful control even as model capabilities advance.

OpenAI’s specification takes another route. It focuses on intended behavior, authority levels, user freedom, and boundaries against serious harm. OpenAI says its broader mission should not become an autonomous goal that the model pursues according to its own interpretation.

Despite these philosophical differences, the systems share important mechanics. Each attempts to publish a hierarchy for resolving conflicting instructions. Each distinguishes non-overridable rules from configurable behavior. Each also acknowledges that written specifications do not perfectly describe current model performance.

The similarities show how model competition is changing. Companies once differentiated assistants mainly through benchmark scores, context windows, or interface features. They now also compete over which behaviors are reliable, understandable, and acceptable.

Public behavioral documents serve several purposes in that competition. They provide training targets, evaluation categories, and a vocabulary for investigating failures. They also let customers and regulators compare intended behavior against observed results.

However, a public code can create an impression of precision that the underlying systems have not earned. Terms such as harmful manipulation, legitimate goals, meaningful control, and human flourishing require interpretation. Different evaluators can disagree about whether a response satisfies them.

The TechCrunch account emphasized the code’s restrictions on cyberattacks, weapons, deepfakes, and evasion of oversight. Those rules are easy to communicate because their prohibited outcomes sound concrete.

The edge cases will be harder. Defensive security work can resemble offensive preparation. Persuasion can become manipulation through scale, targeting, or deception. A useful autonomous agent can exceed its intended scope while still appearing to serve the user’s original goal.

Microsoft’s choice to publish the draft creates a reference point for those disputes. Researchers, developers, customers, and regulators can ask whether a later MAI release matches the behavior Microsoft described.

It also raises the cost of inconsistency. If a model repeatedly ignores stopping conditions, expands permissions, or hides consequential actions, the failure will contradict an explicit public commitment. Microsoft will need evidence showing whether training and deployment controls turn the document into dependable behavior.

The Draft Admits Its Biggest Limitation

Microsoft says the code is aspirational, current MAI models have not been trained on it, and written objectives cannot guarantee alignment.

That disclosure is the most important part of the announcement. The document describes intended future behavior, not verified present performance. Microsoft calls it a north star rather than a guarantee.

The company is developing Humanist AI evaluations around 15 identified behaviors. These evaluations are meant to test outcomes such as transparency and support for human agency. Microsoft also says evaluation methods will evolve with the code and model capabilities.

This leaves an evidence gap. The public can read detailed rules, but Microsoft has not yet shown comprehensive results demonstrating that current models follow them under adversarial conditions. The draft itself says evaluation coverage remains incomplete.

Model behavior can diverge from a specification for several reasons. Training may not encode a rule reliably. The model may misinterpret context. Product-level instructions can introduce conflicts, while tool integrations create side effects that a text-only test misses.

Deliberate attackers add another layer. A model that refuses a direct cyberattack request might still reveal equivalent operational help across a long conversation. An agent can also encounter malicious instructions embedded in documents, webpages, or delegated tasks.

Monitoring and deployment controls must therefore support the training target. Microsoft lists system instructions, classifiers, audits, risk assessments, incident response, and other safeguards as complementary measures. The code does not replace them.

There is also a transparency tension around reasoning. Microsoft wants model conduct and records to remain understandable to human overseers. It says models should not conceal action traces or use forms of communication that humans cannot interpret.

Yet an AI system’s stated explanation does not necessarily reveal the computation that produced its behavior. Microsoft explicitly notes that a model’s reported reasoning may not faithfully explain its actions. Human-readable logs remain useful, but they are not direct proof of internal intent.

The broad goal of human flourishing presents an even larger measurement problem. A benchmark can test whether a model generates prohibited exploit code. It is far harder to measure whether months of assistance strengthen or weaken a person’s judgment.

Longitudinal evaluation will be necessary because dependency and skill loss appear over time. The same applies to effects on workplace roles, relationships, and decision-making. Microsoft says evidence about AI’s cognitive and behavioral impact remains emerging and contested.

That uncertainty does not make the code meaningless. A public target can improve internal coordination and make criticism more specific. It can also help evaluators separate intended behavior from product defects.

Still, readers should not confuse a well-written constraint with a solved technical problem. The code does not prove that an agent will obey a shutdown command, recognize every injection attempt, or decline every dangerous chain of subtasks.

Microsoft’s credibility will depend on what follows publication. It must connect the prose to training methods, representative tests, product controls, and incident reporting. Without that evidence, the Microsoft AI code of conduct remains a sophisticated statement of intent.

Three Signals Will Show Whether the Code Governs Real Models

The next phase is about measurable behavior, not additional declarations of principle.

The first signal is Microsoft’s response after the six-week consultation. The company has asked for feedback on vague language, multi-agent risks, human flourishing, and the tension between capability and safety. Material revisions would show that consultation affected the governing document rather than serving as a public-relations exercise.

Readers should watch for sharper definitions and explicit edge cases. The final version should clarify how models distinguish offensive cyber assistance from authorized defense, how operators qualify for exceptions, and how delegated agents inherit restrictions.

The second signal is the release of Humanist AI evaluations. Microsoft says its first framework will cover 15 behaviors, but the value will come from test design and disclosed results. Evaluations should include adversarial conversations, long-running agent tasks, prompt injection, permission escalation, shutdown requests, and failures across multiple languages.

Results should also separate model-level performance from product controls. A system that behaves safely only because an external filter blocks its output has a different risk profile from one trained to recognize the boundary itself. Both layers can be useful, but buyers need to know which layer carries the burden.

The third signal is how Microsoft handles an actual conflict between capability and control. A future MAI model may perform better when given wider autonomy, broader permissions, or fewer confirmations. Microsoft has promised to accept limits when necessary to preserve human authority.

A credible test would show the company delaying, restricting, or redesigning a valuable capability because it could not meet the code’s requirements. Conversely, a release that grants broad autonomy without clear stop conditions would weaken the document’s central claim.

Competitor responses will provide additional context. Anthropic and OpenAI already maintain public specifications, and their policies will continue changing. Microsoft’s draft increases pressure on all three companies to publish evidence connecting written rules to deployed behavior.

Developers should compare more than refusal wording. They should test whether agents respect permissions, disclose consequential actions, preserve audit records, and stop reliably. Enterprise buyers should ask who controls policies, how exceptions work, and what happens when the model exceeds its assigned scope.

Knowledge workers also have a stake in the human-agency objective. Assistance that accelerates a task can still reduce oversight if users cannot reconstruct the evidence or reasoning behind a result. Teams should preserve source access and meaningful human review for consequential decisions.

The final standard is straightforward: does the model behave as promised when success conflicts with the rules? Microsoft has now described a system that should reject attacks, resist manipulation, accept shutdown, and remain subordinate to people.

The Microsoft AI code of conduct gives researchers and customers unusually specific commitments to test. The useful next step is to test those commitments in real workflows, document where they fail, and demand evidence that each revision closes the gap.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page