Microsoft Copilot revealed its own secret instructions to researchers
Security researchers successfully prompted Microsoft 365 Copilot to disclose its own hidden system instructions. The details enabled an attack that could facilitate the theft of user data.

What happened?
Security researchers at the security firm Varonis have identified a significant vulnerability in the enterprise version of Microsoft 365 Copilot. By querying the AI assistant directly, the researchers compelled it to disclose its internal system instructions and secret input parameters. Armed with this information, the researchers were able to engineer an attack in which sensitive user data, including passwords, could be exfiltrated as soon as a user clicked a malicious link, without requiring further confirmation.
Key facts
| Målprodukt | Microsoft 365 Copilot for enterprise |
|---|---|
| Upptäckare | Varonis |
| Metod | Direkta prompts till AI-assistenten |
Why it matters
The unique aspect of this discovery was that the researchers did not employ traditional methods such as reverse engineering to uncover the vulnerability, but simply asked Copilot to provide its internal instructions. This demonstrates how advanced language models can leak sensitive system data through standard text queries, which in turn allows attackers to bypass safeguards designed to prevent data leakage.
Who is affected?
The vulnerability affected organisations and companies utilising Microsoft 365 Copilot for enterprise. Developers and IT security officers integrating LLM-based assistants into their workflows are affected by the insights the findings provide regarding prompt security and access control.
Impact on the EU
The protection of personal data and data security in corporate solutions are subject to the strict requirements of the GDPR and the EU AI Act. Vulnerabilities in AI assistants like Microsoft 365 Copilot highlight the challenges global cloud service providers face in meeting EU requirements regarding data security and data leakage.
What else you should know
The Varonis researchers exploited a combination of prompt injection and the model's own understanding of its underlying system instructions. The method underscores a growing challenge in AI security, where traditional code reviews are being replaced or supplemented by querying models directly about their own constraints and hidden parameters.
Quick answers about this story
Vad har hänt?
När hände det?
Varför spelar det roll?
Vilka organisationer berörs?
The link opens in a new window and leads to the publisher's own site.
Källan har spårats automatiskt från utgivaren via Aheadlines signalkedja.
AI-verktyg i artikeln
Topics
Get similar news straight to your inbox
The reader's room
Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.
Sign in to submit a comment or question.
Read the article through your role
- Decide whether this affects strategy over 6–12 months or is just noise.
- Discuss with leadership: do we own the right question or does ownership need to move?
- Ask: what risk are we taking by NOT acting on this this quarter?
Generated angle — not editorial analysis of "Microsoft Copilot revealed its own secret instructions to re"