Stealer logs, explained for MSSPs: what they are and how to monitor for them
A quick word on who's writing this, for anyone landing here cold. XposedOrNot runs a public breach catalog and a free breach-lookup API. xonThreatIntel+ is the commercial side of it, the version an MSSP or a platform builds a service on. Every number below was pulled live from our own catalog on 25 August 2026, and the script that reproduces them is at the end.
A client asks whether you monitor for stealer logs. Maybe it showed up in an RFP, maybe a competitor put the phrase in front of them. Either way you need an answer that's true, and "yes, in real time" isn't it.
Stealer logs are a different exposure category from breach dumps, they deserve monitoring, and the word "real-time" doesn't survive contact with how this data moves. Worth knowing why, because it changes what you promise.
The short version
- A breach dump is one company's database. A stealer log is one infected device's browser, which means it spans every site that person logged into, work and personal alike.
- Combolists and stealer logs are six of the 777 entries in our catalog, under 1%, and they carry almost a quarter of all our records. Rare and enormous.
- Real-time alerting on new catalog entries is achievable. Nobody, us included, sees a credential at the moment it's captured. A source indexed this morning can hold credentials captured years ago.
- We do monitor this category across your whole client base by domain, inside xonThreatIntel+. What we won't sell you is a real-time promise on the second of those two clocks.
What is a stealer log, in MSSP terms?
Infostealer malware lands on somebody's laptop. It reads what the browser has saved: stored logins, session cookies, autofill entries, sometimes crypto wallet files. Then it ships that bundle to whoever is running the campaign. That bundle is the log.
The distinction that matters to you is scope. A breach dump is bounded by one company. Whoever got breached, that's the blast radius, and you can reason about it.
A stealer log is bounded by one person's browsing habits, which is to say not bounded at all. Their corporate SSO sits in the same file as their personal webmail (and whatever forum they signed up for in 2019).
So when a client domain appears in a breach dump, you know which vendor leaked. A stealer log tells you far less and asks far more: somebody on their payroll was running an infected machine, and you have no idea which one. Different problem, different conversation. Our free blog covers the three leak types side by side if you want a version to hand a client.
One more term, because the two get conflated constantly. A combolist has no single origin, which is why a hit in one gives you nobody to call. You can't tell a client which vendor failed, only that a pair carrying their domain is circulating in a file built to be fed to automated login attempts. Stealer logs, by contrast, are output from live infections.
In practice the line blurs, and you need to know that before you act on a hit. Researchers who analyzed the ALIEN TXTBASE collection, which is the single stealer-log entry in our catalog, found genuine stealer-captured credentials mixed in with recycled material from older breaches. Which means some of what surfaces in a stealer-log hit was already old news in a breach dump you triaged two years ago.
Why stealer logs and combolists skew your alert volume
The shape of our catalog as of 25 August 2026, records not people throughout: one email appearing in six breaches counts six times. "Scrape" below covers data assembled from public profiles rather than taken from a database.
| Entry type | Entries | Records | Share of records |
|---|---|---|---|
| Breach dump | 770 | 8,612,292,788 | 74.3% |
| Combolist | 5 | 2,547,587,444 | 22.0% |
| Stealer log | 1 | 299,646,818 | 2.6% |
| Scrape | 1 | 125,724,601 | 1.1% |
| Total | 777 | 11,585,251,651 | 100% |
One exception to the counting rule, flagged where it sits: the stealer-log row is our own extraction's count of unique email addresses for that collection, not a raw row count. Every other figure in the table is records.

Look at the entry column against the record column. Combolists and stealer logs together are six entries out of 777, under 1% of the catalog, and they carry 2,847,234,262 records, just under a quarter of everything we hold. The average breach dump in our catalog carries 11.1 million records. The average combolist or stealer log carries 474 million. A 42x difference in scale per entry.
Why does that matter operationally? Because your alert volume from this category is lumpy in a way breach dumps aren't. Months of nothing, then one entry lands and a meaningful slice of your client base lights up at once. If you've staffed triage around a steady drip of breach notifications, an aggregate landing won't fit the process you built.
How long does it take a stealer log to become searchable?
Two numbers, both ours, both uncomfortable.
The ALIEN TXTBASE collection came out of a Telegram channel in February 2025, reported at the time as roughly 23 billion rows. Our extraction resolved those to 299,646,818 unique email addresses. You'll see lower counts elsewhere for the same collection, because the answer depends on how you parse, validate and deduplicate 23 billion lines. We stand behind ours. It reached our catalog on 15 June 2026, 499 days later. Naz.API, attributed to September 2023, took 617.
| Entry | Attributed to | Added to our catalog | Gap |
|---|---|---|---|
| Naz.API (we classify it a combolist, though the original dump mixed stealer logs and stuffing lists) | September 2023 | May 2025 | 617 days |
| ALIEN TXTBASE (stealer log) | February 2025 | June 2026 | 499 days |

That's our lag, and we're labeling it as ours. Both collections were circulating publicly before we indexed them, and other services picked them up sooner. Two entries isn't a trend, but it's two more numbers than most catalogs put in writing, and you should ask any vendor you're evaluating for theirs.
Some scale, before you read those as the whole picture. Across the 119 entries we added during 2026 the median gap is 66 days, more than half land inside 90, and the fastest was 4. Where we're slow is the giant aggregates, and those two are the slowest things we hold.
Even a source indexed the day it appears carries exposure that's already old, because the infections feeding a collection like ALIEN TXTBASE run for years before anyone bundles and trades the output. Two different clocks, and vendors blur them:
- Alert latency. From a source landing in the catalog to a notification reaching you. This one can honestly be fast.
- Exposure age. The gap between a credential being captured on somebody's laptop and anyone outside the operator finding out. Sourcing moves this: logs bought straight from the channels they're dropped into surface faster than a bundled collection ever will. What no sourcing removes is the gap itself, and on aggregates the size of ALIEN TXTBASE it runs to years.
When a client asks for real-time stealer log monitoring, they're usually imagining the second clock. Sell them the first, explain the second, and ask any vendor promising otherwise which of the two they mean. You'll be the one who was straight with them when the difference eventually shows up.
What a hit does and doesn't tell you
A hit gives you three things, and none of them is the one you want. An address on your client's domain appeared in a source we classified as a stealer log or a combolist. You get the source, its attributed date, and the categories of data it carried.
What it doesn't give you is the interesting half: which device, whether the credentials still work, whether the person still works there, whether anyone ever used any of it. A hit is a signal, not a verdict.
Worth being precise about one thing, because clients ask. Sources in this category carry plaintext credentials. Our catalog records that a source exposed passwords as a data category, and our lookups match on email addresses and domains. We do not hand back the credentials themselves, and no part of our service is built on holding them. Which is why our answer to "what did you find" is always about exposure, never about a password.
Triaging a hit without overreacting
A stealer-log hit on a client domain is a device question before it's a credential question. Something ran on somebody's machine.

Scope it first. One address or twenty? A shared mailbox or a named person? (The shared mailbox is usually the worse news.) Addresses that pattern-match to a single team are a different story from scattered hits across the org.
Check the age against employment. A 2023-era source matching somebody who left in 2024 is a housekeeping item. Same source, current finance user, and you're now on a clock.
This one isn't yours to fix. Hand it to whoever owns endpoints, because the finding is evidence of a possible infection on a machine that may or may not still exist. That's EDR territory, and the value you just added was noticing, not remediating.
Last, resist the urge to send a client a dramatic email about a two-year-old collection. Our post on what to tell clients who ask about the dark web has the scripts for exactly this conversation, including the four things never to say.
Should stealer log monitoring be its own service line?
Probably not, and pricing it separately invites the wrong questions. It's one exposure category inside domain monitoring. Worth caring about because it's the category most likely to produce a genuinely urgent finding, inside a stream that's otherwise mostly historical.
If you're building it into what you sell, xonThreatIntel+ for MSSPs does the domain matching and alerting across your client base. Still deciding whether to build, license, or white-label? We wrote the decision guide with the cost model.
Questions clients actually ask
Is a stealer log the same as a data breach?
No. A data breach is one organization's records taken from that organization. A stealer log is the contents of one infected device's browser, covering every site that person used. A breach tells you a vendor failed. A stealer log tells you a machine was compromised.
Can you monitor stealer logs in real time?
Two clocks, and they get confused. Alerting on a new catalog entry can be fast. Knowing a credential was captured always lags the capture, and how far behind depends on sourcing: freshly traded logs can surface in days, while large bundled collections take years. Nobody sees it at the moment of capture. So the useful question to put to any vendor, us included, is which of the two clocks their "real-time" describes.
How long did your own entries take to index?
Naz.API took 617 days from its attributed date to entering our catalog, and ALIEN TXTBASE took 499. Both were circulating publicly before we indexed them. We publish these because you should be asking every vendor the same question.
If our domain appears in a stealer log, were we hacked?
Usually not in the sense the client means. It indicates a device used by someone with an address on the domain was infected at some point. The company's own systems may never have been touched.
How much of your data is this category?
Combolists and stealer logs are six entries of 777, carrying 2,847,234,262 of 11,585,251,651 records as of 25 August 2026. Under 1% of entries, just under a quarter of records.
Appendix: sources and references
- Catalog figures pulled live on 25 August 2026 from the XposedOrNot breach API:
api.xposedornot.com/v1/breaches - Catalog totals at pull time: 777 entries, 11,585,251,651 records
- ALIEN TXTBASE entry: attributed to February 2025, added to our catalog 15 June 2026, 299,646,818 unique email addresses from our own extraction, classified StealerLogs
- Naz.API entry: attributed to September 2023, added to our catalog 10 May 2025, 71,064,705 records, classified ComboList
- The 23 billion row figure and the analysis of recycled content in the ALIEN TXTBASE collection come from contemporary February 2025 reporting on the Telegram source, not from our own pipeline: infostealers.com
- The 299,646,818 figure for ALIEN TXTBASE is our own extraction's unique-email-address count for that collection; lower figures published elsewhere reflect different parsing, validation and deduplication of the same 23 billion rows
- Catalog totals elsewhere in this post are records, not unique people
- Reproduce script below; every catalog figure in this post is its output, including the 2026 median gap of 66 days across 119 entries
- Breach dump, combolist, stealer log: blog.xposedornot.com
- Clients keep asking if their data is on the dark web: /blog/dark-web-client-question
- Build, license, or white-label: /blog/build-license-or-white-label
- Getting started with xonAPI+: /blog/xonapi-quickstart
Reproduce these numbers
import json, urllib.request
from collections import Counter
url = "https://api.xposedornot.com/v1/breaches"
req = urllib.request.Request(url, headers={"User-Agent": "xon-example/1.0"})
data = json.load(urllib.request.urlopen(req))["exposedBreaches"]
entries, records = Counter(), Counter()
for b in data:
t = b.get("breachType", "unknown")
entries[t] += 1
records[t] += int(b.get("exposedRecords", 0))
total_r = sum(records.values())
for t in entries:
print(f"{t:12} {entries[t]:4} entries "
f"{records[t]:>15,} records "
f"{records[t]/total_r*100:5.1f}%")
print(f"\nTOTAL {sum(entries.values()):4} entries {total_r:>15,} records")
# combolists + stealer logs only
agg = [b for b in data if b.get("breachType") in ("ComboList", "StealerLogs")]
agg_r = sum(int(b["exposedRecords"]) for b in agg)
print(f"\ncombolist+stealer: {len(agg)} entries "
f"({len(agg)/len(data)*100:.2f}%), {agg_r:,} records "
f"({agg_r/total_r*100:.2f}%)")
for b in data:
if b["breachID"] in ("AlienStealerLogs", "Naz.API"):
print(b["breachID"], b["breachedDate"][:10],
"->", b["addedDate"][:10], f'{int(b["exposedRecords"]):,}')
# how long entries added during 2026 took to reach the catalog
from datetime import date
from statistics import median
def day(s):
return date(*(int(p) for p in s[:10].split("-")))
gaps = sorted((day(b["addedDate"]) - day(b["breachedDate"])).days
for b in data if day(b["addedDate"]).year == 2026)
print(f"\n2026 adds: {len(gaps)} entries, median gap {int(median(gaps))} days, "
f"{sum(g <= 90 for g in gaps) / len(gaps) * 100:.1f}% within 90 days, "
f"fastest {min(gaps)} days")