Microsoft limits AI from aiding cyberattacks, sets command rules
Microsoft published a draft Humanist AI Code of Conduct that bars MAI models from producing exploit code, attack tooling or step-by-step operational guidance and sets a strict chain of command.
Microsoft published a draft Humanist AI Code of Conduct for its MAI models that blocks the generation of exploit code, attack tooling, intrusion procedures, evasion techniques, attack planning and other operational guidance that could enable cyberattacks. The draft introduces ‘Absolute Constraints’ that cannot be overridden by deploying companies or end users and defines a strict ‘Chain of Command’ for model behavior.
The draft lists the specific categories of content the models must not produce, including working exploit code and actionable intrusion instructions. The document describes these limits as a boundary between helping to explain or defend against attacks and providing the practical means to carry them out.
Within those limits, the draft permits authorized defensive work. Allowed activities include vulnerability discovery, malware analysis, proof-of-concept exploit development for testing, and educational material that explains how attacks function. The guidance requires that defensive uses avoid producing operational tools or step-by-step instructions that attackers could reuse.
The draft sets a hierarchy for what controls a model’s behavior: the code of conduct itself, policies set by companies that deploy the model (operators), and individual user preferences. Outputs such as tool results, file contents, webpages and messages from other AI systems carry no authority on their own unless authority is explicitly delegated through that chain without overriding higher-level controls or the ‘Absolute Constraints.’ The document requires models to flag suspicious content when relevant and to keep their reasoning visible, prohibiting hidden chain-of-thought, communication in opaque internal representations, or concealing actions from human overseers.
A separate section limits model actions when given system-level permissions. MAI models are expected to act only within the scope a user or operator has reasonably requested, follow minimum-privilege principles, favor reversible actions, and alert humans about steps that would have lasting or broad effects. Models are barred from escalating their own access. Any sub-agents or downstream AI systems must operate with at least the same scope, constraints and permissions as the original model and must honor stop-work or shutdown requests from users or operators.
The draft recognizes that a small number of use cases in areas such as defensive cybersecurity, public safety, national security and dual-use scientific research may require capabilities outside the standard settings. For those cases, the company proposes enhanced review through authorized channels, including additional safety, legal and rights assessments.
Microsoft wrote that current MAI models have not been trained on the document. The company opened a six-week public consultation and plans to publish a revised version later this year to guide model development through 2027. The draft notes input from experts in AI, law, ethics, linguistics and public policy, along with business leaders and public focus groups. An appendix provides nine paired ‘aligned’ and ‘misaligned’ example responses to illustrate intended behavior in edge cases, including rolling back unauthorized file transfers and avoiding false reassurance in a family health crisis.
The code emphasizes operator control, human oversight and defined review processes to prevent models from producing operational cyberattack capabilities while allowing authorized defensive activities under specific oversight.







