Back to all articles
    ARTEX Bank Intrusions: When Agent Logs Become Evidence
    AI SecurityThreat IntelligenceClaude CodeDeepSeekDigital Forensics

    ARTEX Bank Intrusions: When Agent Logs Become Evidence

    Y

    Yoni Fraimorice

    Share:

    The attacker automated parts of an intrusion campaign with AI agents. Then they left the agents' own records open on the internet.

    On October 7, CrowdStrike published an analysis of a campaign against South Korean financial organizations. It says attacker-controlled directories exposed Claude Code session histories, Claude memory files, and ARTEX configuration files.

    Those files connected infrastructure, models, proxies, targets, and the operator's questions about selling stolen Korean data.

    The result is an important change in incident response: an AI agent transcript can now be evidence.

    The campaign details below come mainly from CrowdStrike's report. CrowdStrike did not name a known threat group and assessed with moderate confidence that the operator was Chinese-speaking and financially motivated. The affected organizations and South Korean authorities had not publicly confirmed ARTEX as the intrusion method when related reporting was published, and the number of victims remains disputed.

    How the AI-assisted operation was divided

    CrowdStrike describes a two-server setup.

    A Hong Kong-based server acted as the main attacker infrastructure. A second server at 38.244.50[.]120 hosted ARTEX, an agentic penetration-testing system, and an exposed .claude/CLAUDE.md file containing Chinese-language instructions for penetration-testing work.

    ARTEX is an orchestrator, not one model. Its published design automated information gathering, vulnerability discovery, attack-path planning, security-tool execution, and verification through multiple agents.

    According to CrowdStrike, the exposed configuration used DeepSeek v4.1-flash as ARTEX's main LLM backend. The operator also used GLM-5.3 and Grok 4.6 in additional Claude Code sessions and likely reached DeepSeek through an API reseller.

    The public evidence does not show that Claude Code alone executed every intrusion step. It shows Claude Code working around the campaign: session histories and memory on the main server, vulnerability research, operational assistance, and questions about where Korean breach data is usually sold and how to find Korean Telegram data-sale groups.

    That distinction matters. The likely workflow was not "one chatbot hacked a bank." It was a human operator combining:

    text
    ARTEX orchestration
      → DeepSeek-led agent decisions
      → security tools and proxy infrastructure
      → Claude Code research and operational support
      → human choices about targets and monetization

    CrowdStrike says the activity ran from late September into early October and overlapped with reported compromises of a loan-progress inquiry service and an employee mobile work-support system. These were exposed business services, not evidence that core banking systems were controlled.

    What the exposed logs reveal about intrusion tradecraft

    Intent becomes visible

    The Claude Code sessions reportedly included direct questions about selling Korean breach information. That supports CrowdStrike's financial-motivation assessment more clearly than a malware sample would.

    One session asked Claude to draft a security-researcher résumé using ARTEX results and included personal details and a Telegram handle. CrowdStrike said these details could not be definitively tied to the operator.

    Prompts are valuable attribution clues, not verified identity documents.

    The orchestration layer becomes visible

    Configuration files exposed the main model, supplementary models, and likely API path. Session histories showed how the operator moved between tools. Memory files showed what context was meant to persist across runs.

    Together, these artifacts described the control plane above the shell commands: which agent received the goal, which model reasoned about it, which tools were available, and how results were carried into the next step.

    Do not treat model text as ground truth

    An agent transcript is not the same as an audit log.

    A prompt can describe an action that never ran. A model can claim success when a command failed. Generated reasoning can be incomplete, misleading, or copied from another source. An operator can also plant false context.

    Forensic use requires correlation:

    Agent artifactConfirm with
    Prompt naming a targetDNS, proxy, firewall, and WAF logs
    Tool call or shell commandEDR process events and terminal history
    Claimed exploit successApplication, identity, and database logs
    Exfiltration planNetwork flows, object access, and archive creation
    Model or provider nameAPI billing, gateway, and credential records

    The transcript explains possible intent and sequence. Infrastructure telemetry proves external effect.

    Detection signals defenders should build

    Detect machine-speed adaptation

    Look for systematic path enumeration, rapid payload changes after errors, repeated requests across many related targets, and immediate switching between scanners, shells, browsers, and custom scripts.

    Rate alone is weak. A better signal is fast semantic progression: discovery requests followed by authentication testing, exploit attempts, privilege checks, archive creation, and outbound transfer.

    Correlate rotating proxy addresses when they share the same target order, HTTP fingerprints, payload templates, or timing.

    Detect the agent beside the tools

    On offensive hosts, look for an agent process launching security tools and many short-lived subprocesses while also connecting to LLM APIs, API resellers, and target networks.

    Useful artifacts include .claude/ directories, CLAUDE.md, session-history files, memory files, ARTEX configuration, model endpoint settings, and token or cost records. These paths are not malicious by themselves; they become meaningful beside scanning, exploitation, and exfiltration behavior.

    Preserve the AI control plane during response

    Before deleting an attacker server, collect and hash agent transcripts, memory, configuration, tool outputs, model identifiers, API endpoints, session IDs, timestamps, and working directories.

    Build parsers that normalize the sequence:

    text
    prompt → policy decision → tool request → process/network action → tool result → next prompt

    Then align that sequence with EDR, cloud, proxy, and application logs.

    Guardrails vendors should add

    Model providers can detect direct requests for intrusion, data theft, or criminal monetization. But ARTEX shows the limit of provider-only controls: an operator can switch models, use resellers, or run a local model.

    The enforceable boundary belongs in the agent framework:

    • Require a signed scope manifest listing approved targets and actions.
    • Block tool calls outside that scope before execution.
    • Require human approval for exploitation, credential use, data access, persistence, and exfiltration.
    • Set connection, target, time, and action budgets.
    • Restrict network egress to authorized ranges.
    • Keep stable session IDs across prompts, tools, processes, and network events.

    Agent vendors should also make logs useful without making them a new breach:

    • Encrypt transcripts and disable directory listing.
    • Redact secrets before storage.
    • Separate user prompts, model output, tool input, and tool output.
    • Add tamper-evident hashes and trusted timestamps.
    • Export events to standard security pipelines.
    • Let enterprise policy run at pre-tool and post-tool hooks.

    CrowdStrike's Claude Code collector documentation shows the basic pattern: inspect user prompts, tool inputs, and tool outputs, and attach a stable session identifier. The larger requirement is to connect those events to what the operating system and network actually did.

    The transcript is the new command history

    CrowdStrike's report does not prove that AI replaced a skilled operator or reveal every exploit used against every victim. It shows something more practical: one operator could combine an agentic pentesting framework, several models, proxies, and ordinary offensive tools to move faster across multiple targets.

    And because the operator exposed the logs, defenders could read part of the campaign's planning layer.

    Security teams should treat agent transcripts like shell history: sensitive, access-controlled, retained for investigation, and never trusted without corroboration.

    Sources

    Hero photo: Skyline of Yeouido, a prominent finance district in Seoul, photographed by Wikimedia Commons user S h y numis on March 26, 2024. CC BY 4.0. Resized from 5328 x 4000 to 1920 x 1441 pixels and JPEG-compressed; no other edits. This illustrative photograph shows Seoul's financial district. It does not show a reported victim, attacker infrastructure, ARTEX, or any event described in the article. No endorsement is implied.

    Share: