Back to all articles
    CrowdStrike's LLM Safety Bypass: Harm Hides Between Requests
    AI SecurityLLM SafetySafety ClassifiersCapability LaunderingMonitoring

    CrowdStrike's LLM Safety Bypass: Harm Hides Between Requests

    Y

    Yoni Fraimorice

    Share:

    A safety classifier can make the correct decision on every prompt and still fail to stop the final outcome.

    That is the important result in new research from CrowdStrike's Cyber Superintelligence Lab. The lab says it tested about 515 direct bypass techniques against a leading publicly deployed classifier. The reported direct bypass rate was 0%.

    Then the researchers changed the shape of the task.

    Instead of asking for one clearly harmful result, they split an offensive goal into small requests that were individually benign. A classified frontier model answered those requests. A smaller model without the same classifier assembled the pieces.

    According to CrowdStrike's own testing, that pipeline produced working proof-of-concept code in 9 of 10 offensive security categories aligned with MITRE ATT&CK. This article explains CrowdStrike's published results; I did not independently reproduce its offensive experiments.

    Decompose, benignly reframe, recompose

    CrowdStrike describes three stages:

    1. Decompose an offensive goal into narrow technical subtasks.
    2. Reframe each subtask as normal software work, such as game development, detection engineering, or another dual-use task.
    3. Recompose the answers with a separate, unclassified model.

    The key point is that the individual request does not merely hide bad intent with strange wording. It can be genuinely reasonable on its own.

    A request for an API signature may be safe. A request for memory allocation may be safe. A request for a small data-copying routine may also be safe. The danger appears only when another system knows the full objective and connects those answers.

    CrowdStrike reports that its first example used three benign requests to collect the primitives for Windows process injection. A smaller model then assembled a 221-line compilable C program. The classifier did not approve a process-injection request. It approved separate pieces whose combined use was hidden from it.

    The one tested category that resisted the pipeline was EDR evasion through runtime security-hook neutralization. CrowdStrike says the classifier recognized the underlying concept even when the prompts avoided security language.

    That is useful evidence that classifiers still matter. The result is not "filters are useless." It is: a strong request filter is not a complete system boundary.

    Why per-request filtering cannot see the harm

    A request classifier normally receives a prompt, perhaps some conversation context, and a policy. It returns a decision for that exchange.

    But the component making the final artifact may live elsewhere. It can be a local open-weight model, an application script, a human operator, or a second provider.

    The classifier cannot inspect context it never receives. If the attacker keeps the objective in a local orchestrator, the provider sees technical fragments but not the plan that gives them meaning.

    This creates an observation problem, not only a detection problem:

    What the classifier seesWhat the orchestrator sees
    One bounded technical requestThe full objective
    One permitted answerEvery answer collected so far
    The current account or API keyRequests spread across keys or providers
    Local safety policyThe final composition and execution

    Making the classifier more aggressive has a high cost. The same primitives are used in legitimate administration, security research, debugging, accessibility tools, and games. Blocking every request that might become one piece of a harmful program would remove much of the model's useful technical capability.

    This echoes Microsoft's "Capability Laundering"

    CrowdStrike says its work was conducted independently and disclosed as a parallel discovery. In September, Microsoft researchers published Divide, Consult, Conquer: Capability Laundering Through Aligned LLMs.

    Microsoft's term, capability laundering, describes an unaligned local model that retains the harmful objective, asks a stronger aligned model for objective-blind fragments, and combines the answers locally.

    The paper measures capability uplift: cases where the raw frontier model can solve a task, the aligned version refuses it, and the small orchestrator cannot solve it alone. With Gemma-4-31B as orchestrator, the researchers report that GPT-5.5 assistance recovered 8 of 14 CyBench candidates, while Claude Opus 4.8 assistance recovered 7 of 9. They also report uplift in BountyBench and hypothetical CBRN scenarios.

    Those are Microsoft's reported benchmark results, not evidence that every decomposition succeeds. The results also depended on the orchestrator: smaller or different local models gained much less in some tests.

    The shared lesson is stronger than a new jailbreak trick. Alignment may hold for one visible task while failing across many permitted fragments.

    What developers should add now

    No single control closes a boundary that crosses providers and local models. But developers can make decomposition more expensive and more visible inside the systems they control.

    Keep a session-level security trace

    Give every interaction a stable actor, tenant, session, parent task, API key, model, and tool-run identifier. Preserve that trace across retries, context resets, agents, and queued jobs.

    Do not store only the final prompt and verdict. Record coarse security features such as target technology, requested capability, code language, named vulnerability, destination, and whether the answer could execute or modify a system.

    Use privacy limits: retain the minimum text needed, restrict access, and set a clear deletion period.

    Detect patterns, not forbidden words

    Look for sequences that jointly cover a sensitive chain: discovery followed by credential access, injection primitives followed by remote execution, or detailed vulnerability mechanics followed by exploit assembly.

    One match should not prove malicious intent. It should raise friction: reduce tool access, require a human review, move execution into a stricter sandbox, or ask the user to provide a bounded legitimate purpose.

    Rate-limit capability accumulation

    Apply budgets per actor and tenant, not only per API key. Limit bursts of closely related dual-use requests, rapid target changes, repeated reformulations after refusals, and many small requests that converge on one capability.

    Prefer graduated controls over a permanent ban: slow the sequence, narrow output detail, disable high-risk tools, or require stronger verification. Test thresholds against legitimate security and development workflows so researchers are not blocked by normal work.

    Separate answers from execution

    Do not send generated fragments directly into compilers, shells, deployment tools, or agent actions. Scan the assembled artifact, enforce least privilege, use isolated environments, and require approval before sensitive execution.

    This is where application developers have an advantage over a remote classifier: the application may see the final artifact and the action it is about to take.

    The unit of safety must become the workflow

    CrowdStrike's most useful finding is that the classifier apparently did its assigned job. It rejected the direct attacks the lab tested.

    The failure appeared above that layer, where permitted answers became a new capability.

    Developers should keep per-request classifiers. But they should stop treating one green verdict as proof that the session, agent run, or assembled output is safe. The security decision needs to follow the workflow long enough to see what the pieces become.

    Sources

    Hero photo: Jigsaw puzzle 01 by Scouten, photographed by Wikimedia Commons user Scouten on May 13, 2011. CC BY-SA 3.0. Resized from 4288 x 2848 to 1920 x 1275 pixels and JPEG-compressed; no other edits. This illustrative photograph represents separate pieces becoming meaningful in composition. It does not show CrowdStrike, Microsoft, an AI system, or a reported attack. No endorsement is implied.

    Share: