Numbers below match Model Card v3.1.2 (external). Four independent measurement perspectives calibrate from benign only. The operator controls one parameter, α (false positive budget). Full operating points and limitations are on the model card.
CICIDS-2017 (multi-scale)
5-fold CV · benign calibration only · independently reproduced AUC to 4 decimals
| Day | Flows | Attack % | AUC | F1 @ α=0.05 | F1 @ α=0.005 |
|---|---|---|---|---|---|
| Friday | 500,000 | 46.9% | 0.9990 | 0.972 | 0.992 |
| Wednesday | 496,641 | 35.7% | 0.9986 | 0.957 | 0.995 |
| Tuesday | 322,078 | 2.2% | 0.9985 | 0.470 | 0.898 |
| Thursday | 362,076 | 20.4% | 0.9984 | 0.910 | 0.969 |
Tuesday shows α at work: low attack prevalence with a 5% FPR budget hurts F1 while AUC stays high. Tightening α recovers F1 without changing the instrument.
MAWIFlow (real backbone)
| Year | Method | AUC | F1 @ α=0.05 | Peak F1 |
|---|---|---|---|---|
| 2016 | Flow-only | 0.954 | 0.608 | 0.634 (α=0.07) |
| 2021 | Multi-scale | 0.924 | 0.521 | 0.547 (α=0.07) |
| 2021 | Flow-only | 0.860 | 0.517 | 0.525 (α=0.07) |
| 2011 | Multi-scale | 0.892 | n/a | n/a |
Multi-scale on 2016 is not stable and is not reported. Do not tighten α on MAWI: at α=0.005, 2016 F1 falls to 0.150 and 2021 to 0.118.
CSE-CIC-IDS-2018 (10 daily captures)
500K random sample per day · attack-weighted F1 across days: 0.876
| Day | AUC | F1 | Attack types |
|---|---|---|---|
| Fri 16-02 | 1.000 | 0.948 | DoS Hulk |
| Tue 20-02 | 1.000 | 0.773 | DDoS-LOIC |
| Wed 21-02 | 1.000 | 0.929 | DDoS-HOIC |
| Wed 14-02 | 0.993 | 0.781 | FTP/SSH Brute Force |
| Thu 15-02 | 0.993 | 0.348 | DoS GoldenEye/Slowloris |
| Fri 02-03 | 0.994 | 0.592 | Botnet Ares |
| Wed 28-02 | 0.995 | 0.318 | Infiltration |
| Thu 01-03 | 0.991 | 0.287 | Infiltration/NMAP |
| Fri 23-02 | 0.997 | 0.003 | Web Attack (20 flows in sample) |
| Thu 22-02 | 0.998 | 0.002 | Web Attack (18 flows in sample) |
Other domains (flow-level)
| Dataset | AUC | F1 | FPR |
|---|---|---|---|
| NSL-KDD Train | 0.993 | 0.953 | 2.5% |
| UNSW Train | 0.993 | 0.961 | 2.3% |
| UNSW Test | 0.980 | 0.924 | 2.7% |
| NSL-KDD Test | 0.969 | 0.878 | 2.6% |
| TON-IoT | 0.948 | 0.646 | 3.6% |
| DoH | 0.938 | 0.681 | 5.0% |
| IoT-DIAD | 0.864 | 0.504 | 3.0% |
BCCC-IoT withheld pending end-to-end reproduction of a previously published figure.
Operational profile
| Labels for detection | None (benign calibration only) |
| Cold-start calibration | ~60 seconds |
| Total footprint | ~50 MB |
| Measurement perspectives | 4 |
| Operator parameters | One (α) |
| Hardware | CPU only, no GPU |
| Air-gapped | Yes |
| Encoding throughput | 15,000+ to 70,000+ flows/s |
| Detection throughput | 1,300 to 2,100 flows/s |
| Network required | None |
Disclosure
This page publishes performance only. Clustering recipes, composition formulas, threshold math, named scale internals, and the internal model card are not on the public website, not in downloadable public cards, and not in browser-reachable demo assets.