Skip to content
Säkerhet· NewsAvailable

Hackers bypass AI security by fragmenting malicious tasks

Cybercriminals have discovered a new method to bypass safety filters in AI models by splitting malicious tasks across multiple sessions.

By the Aheadline editorial team·6 aug. 2026·2 min read·Source: Google News: AI safety (en)Aggregerad källaAI-generated
Hackers bypass AI security by fragmenting malicious tasks
Hackers bypass AI security by fragmenting malicious tasks
By · Policy- & EU-reporter

What happened?

Cybercriminals have developed a method to bypass security guardrails in AI models by fragmenting malicious instructions into several separate chat sessions. By feeding the AI seemingly harmless sub-tasks, the models' safety filters fail to detect the ultimate malicious intent, such as the creation of malware or phishing material.

Key facts

AngreppsmetodUppdelning av skadliga instruktioner på flera sessioner
SårbarhetstypKringgående av AI-säkerhetsfilter (Jailbreak/Bypass)

Why it matters

Safety filters in large language models typically evaluate individual prompts or active conversations in isolation. When an attack is distributed across multiple independent sessions, the AI model lacks the context to identify the threat, exposing a fundamental weakness in current security architectures.

Who is affected?

This discovery affects AI developers, cybersecurity firms, and any organisation utilising large language models (LLMs) in their operations. End-users and companies whose cybersecurity depends on AI-based protective mechanisms are also directly impacted.

Impact on the EU

Within the EU, the AI Act imposes strict requirements regarding risk management and cybersecurity for generative AI and general-purpose AI models. Attack techniques that bypass safety guardrails may compel developers to implement stricter monitoring across session boundaries to remain compliant with EU regulations.

What else you should know

Security researchers emphasise that traditional filters that only examine individual prompts are insufficient against advanced attacks. However, building AI systems that remember and analyse context across multiple sessions creates new challenges regarding user privacy and data storage.

Frequently asked questions

Quick answers about this story

Vad har hänt?
Cyberkriminella kringgår säkerhetssystem i AI-modeller genom att dela upp skadliga instruktioner över flera oberoende chatt-sessioner.
När hände det?
Säkerhetsforskare rapporterade om denna sårbarhet i mars 2025.
Varför spelar det roll?
Det avslöjar en svaghet i hur AI-modeller granskar säkerhet, då de oftast bara analyserar en prompt eller session i taget istället för helheten.
Hur kan AI-utvecklare stoppa attackmetoden?
Utvecklare måste bygga säkerhetsarkitekturer som spårar mönster och sammanhang över flera konversationer utan att kränka användarnas integritet.
Original source
Google News: AI safety (en)·news.google.com

The link opens in a new window and leads to the publisher's own site.

Aggregerad källa

Källan är en aggregator eller syndikering — vi rekommenderar att verifiera hos primärutgivaren.

AI-verktyg i artikeln

Topics

#AI-säkerhet#Cybersäkerhet
[ STAY UP TO DATE ]

Get similar news straight to your inbox

No affiliate linksCancel anytimeGDPR-friendly
[ Frequency ]
[ What do you want to read about? ]

You'll receive updates on 2 topics.

The reader's room

Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.

Sign in to submit a comment or question.

Loading comments…
How this affects you

Read the article through your role

  • Decide whether this affects strategy over 6–12 months or is just noise.
  • Discuss with leadership: do we own the right question or does ownership need to move?
  • Ask: what risk are we taking by NOT acting on this this quarter?

Generated angle — not editorial analysis of "Hackers bypass AI security by fragmenting malicious tasks"