Skip to main content

Product evaluation··Sevion Xia

RST QRadar AI Copilot: a hands-on evaluation

A phishing invoice, a night of lateral movement and a screen of Offenses nobody can read: how IBM QRadar and RST QRadar AI Copilot together turn it back into one complete story.

Version 1.1 · September 2026 · Test environment: QRadar 7.6 CE (Community Edition)

Summary

Over one afternoon and into the night, an external attacker got into the internal network through a phishing invoice and went all the way to moving 2.3 GB of finance data out of the country, then wiping the traces. IBM QRadar faithfully correlated the part of this attack chain it could reach into 7 OPEN Offenses. But when those 7 Offenses land in front of an analyst, they are just 7 rule names. Which one comes first, whether any is a false positive, and what actually happened behind them all still have to be dug out by hand.

We replayed this intrusion on a real QRadar 7.6 CE (the attack traffic comes from the public rst-attack-sim) and let RST QRadar AI Copilot take over the work that follows detection:

  • Triage: the 7 Offenses are clustered, rated, ranked and given response advice end to end in about 20 seconds. Priorities: 1 critical, 2 high, 3 medium, 1 low. None was misjudged as a false positive.
  • Investigation: an autonomous investigation on the top-ranked lateral-movement Offense, with the model writing its own AQL to collect evidence from Ariel. 4 retrieval rounds, about 56 seconds, reconstructing the full chain "phishing document → encoded PowerShell → LSASS dump → password spraying → PsExec lateral movement → robocopy exfiltration", with a timeline, MITRE ATT&CK, affected assets and graded response advice.
  • Guardrails: a false-positive verdict only stands with counter-evidence; an entity that appears in the conclusion but not in any evidence row is marked "unverified"; when the model tries to query with a masked placeholder, it is stopped and the server restores the real value before running the query.

In one sentence: QRadar detects, and Copilot explains what actually happened behind each Offense in about a minute, including the larger half of the context that QRadar rules will not connect for you.


At a glance: what RST QRadar AI Copilot is

RST QRadar AI Copilot is an AI security operations assistant that sits on top of IBM QRadar. It connects to QRadar read-only over the REST API and brings the work analysts do after detection into one interface: the security overview, smart query (natural language to AQL), batch triage, rule copilot and alert investigation, plus management features such as analysis history, platform health and system settings.

After sign-in the default landing page is the security overview, which shows on one screen the current alert volume, severity distribution, alert timeline, the rules with the most alerts and the latest alerts:

RST QRadar AI Copilot · Security overview

Smart query is the entry point analysts use most. Say in one sentence what you want to see, and Copilot returns the AQL together with the results. Next to it sits a row of ready-made questions such as "Is anyone attacking us?", "Has an account been compromised?" and "Is anyone operating in the middle of the night?":

RST QRadar AI Copilot · Smart query home

The sections below use one real intrusion to show how far the triage and investigation behind this interface actually go.


1. Prologue: an invoice, and a screen of Offenses nobody can read

At three in the afternoon, j.chen in finance receives an "overdue invoice" email with a .docm attachment. He opens it and enables macros. In that moment, on an endpoint called WS-FIN-07, Word launches an encoded PowerShell that sends its first heartbeat to an address abroad.

Over the next few hours the attacker does not stop. They dump LSASS on that machine to harvest credentials, use the credentials to spray passwords against the domain controller DC01, move laterally with svc_backup over PsExec to the file server FILE01 and to DC01, brute-force their way into srv-app01 and log in with a key, create a sysadm account and a cron backdoor there, and finally copy 2.3 GB out through the Finance share with robocopy. On the way out they clear the audit log on DC01 and wipe /var/log on srv-app01. Meanwhile an unrelated attacker at 91.240.118.66 is scanning the DMZ and brute-forcing the oracle account on srv-app01.

QRadar watches the whole time. It correlates and de-noises this flood of events, and by the next morning the Console quietly shows 7 OPEN Offenses:

QRadar Offenses list

This is the screen an analyst sees on opening the Console in the morning: 7 rule names. The noise reduction is done, but the time-consuming second half has only just begun:

  1. Which one first? Magnitude ordering does not understand business context.
  2. Is it a false positive? Scanners, ops scripts and batch jobs produce similar alerts every day.
  3. What actually happened? A "Multiple Login Failures" Offense might just be an expired password, or the prelude to a successful lateral move. Telling them apart means hand-writing a dozen AQL queries to dig through process creation, outbound connections and post-login account activity in Ariel.

That second half is where RST QRadar AI Copilot works. It does not replace QRadar's detection. It takes over the triage and investigation that come after it.


2. Two tracks: what the attacker did, and what QRadar and Copilot saw

The intrusion is clearest as two parallel tracks. On the left, what the attacker did step by step. On the right, what that step left (or failed to leave) in QRadar, and what Copilot eventually filled in.

Attack chain

TimeAttacker actionNative QRadar detectionContext Copilot adds
T0External port scan (45.142.212.100)✅ Offense: Excessive Firewall DeniesTriaged as a port scan, medium, pushed to the end of the queue
T+1j.chen opens the phishing .docm; the macro launches powershell -enc, beaconing to C2 185.220.101.4:443⛔ No Offense (not covered by CE rules)Investigation round 5 digs winword → powershell -enc out of entry host WS-FIN-07
T+2LSASS dumped for credentials⛔ No OffenseCollected in the same batch in round 5, marked as credential theft
T+3Password spraying against DC01 with the credentials✅ Offense: Multiple Login Failures (DC01/srv-app01)Raised to critical · lateral movement, ranked first
T+4svc_backup moves laterally over PsExec to FILE01/DC01✅ Offense: Login Failures Followed By SuccessRound 6: svc_backup PsExec execution + robocopy reading the Finance directory
T+5Brute force + key login to srv-app01 (deploy)✅ Offense: Multiple Login Failures (several)Round 2: deploy publickey login, sudo escalation
T+6New sysadm account + cron backdoor⛔ No OffenseRound 7: persistence, new account + crontab backdoor
T+72.3 GB exfiltrated through the Finance share with robocopy⛔ No OffenseCollected in the same batch in round 6, marked as data exfiltration
T+8DC01 audit log cleared, /var/log on srv-app01 wiped⛔ No OffenseMarked as anti-forensics / trace removal in the evidence chain
(parallel) Incident B91.240.118.66 scans the DMZ and brute-forces oracle on srv-app01✅ Offense: Firewall Denies + Login FailuresTriaged as a separate low-severity scan/brute force, kept apart from the main chain

The key rows in this table are the ⛔ rows: phishing, the C2 beacon, the LSASS dump, persistence, exfiltration and trace removal. The most damaging steps of the attack chain produce no Offense of their own under the CE default rules (the dashed stages in the attack-chain diagram). They sit in Ariel as raw events, and someone has to pull them out before they can be seen. That is the gap between detecting and understanding.

Open the top-ranked Offense and QRadar gives this summary:

QRadar Offense details (lateral movement)

"Login Failures Followed By Success … preceded by Multiple Login Failures …": the rule name is accurate, but the story behind it has to be dug out. There is also a classic trap: this Offense shows the collector IP as its source and root as the user name, a subject mismatch common in Linux DSM collection (the subject lands on the syslog collector). Copilot falls back to the real log-source host as the subject, so triage does not grab the wrong target.

The raw attack events are all in Log Activity, though. Take the separate incident B, where external IP 91.240.118.66 brute-forces SSH on the oracle account of srv-app01 and finally logs in. In Log Activity it is a clear run of device events: 11 "User failed to login to SSH" (SSH Login Failed) followed by one "Accepted Password for User" (Host Login Succeeded), source ports climbing from 52000, log source Linux OS @ srv-app01:

QRadar Log Activity · oracle on srv-app01 brute-forced over SSH, then logged in

Double-click an event and expand the raw payload, and the hard evidence is right there. Below is the raw payload of a Windows 7045 event on DC01: Service Name: PSEXESVC, Service File Name: %SystemRoot%\PSEXESVC.exe. That is the proof that the attacker used PsExec to move laterally and install a service on the domain controller. Raw payloads like this sit behind the "lateral movement" verdict in the triage table, and a QRadar rule name will not read them for you:

QRadar event details · PsExec service installed on DC01 (Windows 7045, PSEXESVC)

QRadar's capability at this layer comes from its rules and log-source configuration: collect the logs from seven devices, then reduce the noise into Offenses with correlation rules:

QRadar detection rules

QRadar log source configuration

Up to this point detection is clean and effective. The next two sections show how Copilot turns these 7 Offenses from "a screen of rule names" into "a plan you can act on".


3. Triage: a screen of Offenses put in order in about 20 seconds

The analyst clicks "Start triage" in the product. Copilot pulls every OPEN Offense in QRadar for the given time window, clusters them by (rule, subject), and has the model rate, rank, check for false positives and suggest a response. Measured on a live system:

About 20 seconds from click to a fully rendered result table: 7 Offenses → 7 clusters, all scored, no degradation. Severity: 1 critical · 2 high · 3 medium · 1 low, none judged a false positive.

Product · Batch triage results (full page, with the RST logo and module sidebar)

Output (the triage queue in priority order, matching the screenshot above):

PrioritySeverityAttack intentSubject
#1CriticalBrute forceWindowsAuthServer @ DC01 / Linux OS @ srv-app01
#2HighBrute forceSource IP 10.10.20.57
#3HighBrute forceUsername oracle
#4MediumBrute forceUsername root
#5MediumPort scanSource IP 45.142.212.100
#6MediumBrute forceUsername probe9
#7LowPort scanSource IP 91.240.118.66

Copilot rates the cluster hitting DC01 / srv-app01 as the only critical and ranks it first (the same lateral movement the next section confirms in depth), and pushes the two port scans to the end. That matches the real threat level instead of copying magnitude. Every cluster carries a severity, an attack-intent label, an alert count and a time span, and opening one shows a response suggestion and the original alert IDs, so the analyst knows at a glance what to look at first and why. The unreadable screen from the prologue becomes, about 20 seconds later, an ordered plan with reasons.


4. Investigation: the most urgent one told as a complete story in about a minute

The analyst selects the top-ranked lateral-movement Offense and clicks "Investigate".

AI investigation flow

This is not handing logs to a model and asking for a summary. It is an autonomous retrieval loop: the model reads the Offense, writes a read-only AQL query to collect evidence from Ariel, reads the (masked) results, then decides what to query next, like an analyst who never tires and works through the same checklist every time.

The retrieval trace from this investigation on a live system (4 rounds, about 56 seconds, verdict critical, no degradation; every AQL query was generated by the model on the spot, not from a template; results with 0 hits are kept as evidence too):

RoundWhat the model queriedHitsWhat it revealed
1INOFFENSE(5): the Offense's own events20Baseline: login failures and successes for several accounts on DC01/FILE01/srv-app01
2Activity + payload of the deploy account on srv-app016deploy logs in with publickey after 8 failed SSH password attempts, then escalates with sudo
3Aggregated outbound sessions for the target hosts / ports0Negative check: no abnormal direct outbound connections
4Lateral footprint + payload of source sourceip='10.10.x.x' (the masked IP is restored server-side before querying)11New sysadm backdoor → that account logs in from external 185.220.x.x 48 seconds later → packs /etc/shadow and /root/.ssh → clears logs, stops rsyslog

These steps are not random log browsing. They follow a minimum evidence checklist (the Offense's own events, account activity after a successful login, the subject's outbound traffic, exfiltration and trace removal). On the surface the Offense only says "many login failures on DC01/srv-app01", yet the model followed deploy → sysadm → external callback → exfiltration → trace removal and strung together a complete, successful intrusion: exactly the ⛔ stages in the two-track table that QRadar could not alert on by itself.

In this investigation the product did several things that "feeding logs to a large model" cannot:

  • Writes correct AQL on its own: QIDNAME / CATEGORYNAME / LOGSOURCENAME / INOFFENSE / UTF8(payload) ILIKE are all used correctly, with no query templates supplied by a person;
  • Follows the trail under masking: the model only ever sees 10.10.x.x; the server quietly restores the real value it has seen in the evidence before querying (round 4), so real IPs never reach the model and the investigation is not interrupted;
  • Does negative checks: the 0 hits in round 3 are written into the evidence to rule things out, not ignored as "nothing found";
  • Read-only with an allowlist throughout: every AQL query passes the log-source allowlist and anything outside it is refused, so the model cannot touch data it should not just because it "wants to look".

Every step's AQL, hit count and duration are visible to the analyst (below is "View generated AQL" in smart query, full page with the logo and module sidebar; the investigation loop uses the same read-only retrieval):

Product · Smart query and generated AQL (full page, with the RST logo and module sidebar)

At the end the verdict comes in one piece: summary, a complete timestamped timeline, the attack chain, MITRE ATT&CK techniques, affected assets and graded response advice. Below is the product's in-depth investigation report on the whole main attack chain, tracing from the phishing entry on WS-FIN-07 (Invoice_Q3_2026.docm → WINWORD → powershell -enc) to svc_backup moving laterally through PSEXESVC and robocopy reading the Finance directory for exfiltration:

Product · AI investigation report (lateral movement · critical, with timeline / attack chain / MITRE / affected assets / response advice)

Every investigation is archived automatically in analysis history, where it can be traced and re-checked. For the same intrusion, the product can start from the lateral-movement Offense or from the whole main attack chain, and the conclusions corroborate each other:

Product · Analysis history (full page, with the RST logo and module sidebar; every investigation and triage is archived)

Back to the prologue: QRadar gives the analyst the rule name "many login failures on DC01/srv-app01"; Copilot gives them the whole story from j.chen opening a phishing document to finance data leaving the country, with the evidence source for every step. A picture that used to take a senior analyst a dozen hand-written AQL queries and half an hour now takes about a minute, and it can be traced and audited.


5. Guardrails: why these conclusions are fit to show an analyst

On a real intrusion chain, an AI that is "clever but unreliable" brings new risks: inventing an account that does not exist, or letting a real lateral move through as a false positive, can mislead the response. The product builds three guardrails into triage and investigation:

  1. A false positive needs counter-evidence. Before the model may call an Offense a false positive, it checks the time window for privileged logins, account changes, direct external connections, audit clearing and suspicious processes. If any one is present, the false-positive verdict is overruled and escalated. Better to escalate a real false positive for human review than to let a real threat through.
  2. Entities must come verbatim from the evidence. A user name or IP that appears in the conclusion but in no evidence row is removed from the affected assets and marked "unverified", so the model cannot hallucinate accounts.
  3. Masked values are never query conditions. The model only sees masked rows (such as 10.10.x.x). When it tries to query with a placeholder, the server restores it with the real value seen during evidence collection (one match → replace, several candidates → an IN list), so masked values do not pollute queries and the investigation is not interrupted. This is what let the investigation above trace from the Offense all the way to the entry host.

All three guardrails serve one purpose: conclusions that can be checked, rather than "it's true because the AI said so".


6. Value

The most expensive part of a SOC was never buying the SIEM. It is the analyst time spent understanding each Offense after the SIEM raises it. That time has three old problems: it is slow, inconsistent and dependent on senior people. The product targets all three.

Slow → fast. Ordering a screen of Offenses and suggesting responses goes from "open each one and rely on experience" to about 20 seconds; a thorough investigation of one Offense goes from half an hour of a senior analyst hand-writing a dozen AQL queries to about 1 minute.

Inconsistent → reproducible and auditable. Every investigation follows the same minimum evidence checklist, so it will not remember the entry host today and forget exfiltration tomorrow. False-positive verdicts need counter-evidence (see section 5), and every conclusion carries its AQL evidence source.

Dependent on seniors → levelled. "Knowing which dozen AQL queries to run" is the hardest senior experience to copy, and the product builds it into the retrieval loop. A first-line analyst clicks "Investigate" and gets an evidence trail close to a senior's.

StepManualRST QRadar AI Copilot
PrioritiseOpen each by magnitude or experience; easily drowned in noiseAbout 20 seconds to priorities that understand intent, plus response advice
Judge false positivesExperience-based; a miss lets a real threat throughCounter-evidence overrules hasty false positives; verdicts are auditable
See the full pictureA dozen hand-written AQL queries, 30–60 minutes, varies by person4 autonomous retrieval rounds, about 1 minute, with timeline + MITRE + evidence
CoverageTends to stop at the Offense surface and miss the entry host / exfiltrationActively traces back to the entry host and exfiltration path via the minimum checklist
Data safety—Read-only + log-source allowlist + outbound masking; can connect to production QRadar

The product does not replace analysts and does not respond automatically: response decisions stay with people. What it does is explain "what actually happened behind this Offense" in about a minute, with evidence, so analysts spend their time deciding rather than collecting evidence. The same SOC team can watch more Offenses and miss fewer of the real ones.


7. Scope and notes

  • The evaluation ran on QRadar 7.6 CE; commercial editions ship more rules, so the detection layer only gets stronger, while the triage and investigation layers stay the same.
  • Against QRadar the product only runs read-only retrieval plus three limited writes (notes / close / assign for follow-up). It does not block automatically or change rules; response decisions stay with people.
  • Every Ariel query is enforced against the log-source allowlist; referencing a log source outside it is refused outright.
  • The evaluation data is generated by the public rst-attack-sim, so anyone can reproduce it in their own authorised lab.

8. Conclusion

A phishing invoice, a night of lateral movement, a screen of Offenses nobody can read: every SOC meets this day. On that day IBM QRadar does what it does best, reducing a flood of events to 7 Offenses. RST QRadar AI Copilot adds the two things most missing after QRadar detects: triage and ranking in seconds that understands intent, and an autonomous, traceable investigation with guardrails.

It turns a "many login failures" Offense into the complete story from phishing to exfiltration in about a minute, without inventing accounts for you or letting a real threat through as a false positive. Detection goes to QRadar, understanding goes to Copilot, and decisions stay with people.

To reproduce this evaluation: drive the attack chain in with rst-attack-sim, connect RST QRadar AI Copilot, and you get the triage and investigation results in this article.