Comprehensive engineering evaluation of parsing accuracy, browser streaming performance, forensic correctness, local-first reliability, and workflow usability across standardized MBOX email archives.
These values describe the planned benchmark corpus and evaluation methodology. Accuracy and performance results are published after measurement and validation.
Planned across 5 dataset size classes (Tiny to Very Large) strictly bounded below the 1 GB browser Web Worker memory ceiling.
Real-time observations aggregated from local browser storage when Benchmark Mode is enabled. No data leaves your local machine.
Planned dataset sizing and record counts across 5 memory classes (Tiny to < 1 GB Very Large).
| Dataset ID / Name | Size Class | Size (MB) | Planned Records | Attachments | Methodological Description |
|---|---|---|---|---|---|
| Dataset-01 (Tiny) | TINY | 18 MB | 384 | 61 | Small individual mailbox for quick unit and smoke verification. |
| Dataset-02 (Tiny) | TINY | 24 MB | 512 | 83 | Personal archive sample with standard MIME headers. |
| Dataset-03 (Small) | SMALL | 48 MB | 1,042 | 145 | Corporate team correspondence with mixed HTML/plain body structures. |
| Dataset-04 (Small) | SMALL | 85 MB | 1,890 | 310 | Multi-language character sets and RFC2047 encoded headers. |
| Dataset-05 (Medium) | MEDIUM | 142 MB | 3,150 | 520 | Heavy attachment corpus with embedded images and PDFs. |
| Dataset-06 (Medium) | MEDIUM | 215 MB | 4,820 | 780 | Mailing list archive with complex thread hierarchies. |
| Dataset-07 (Medium) | MEDIUM | 284 MB | 6,100 | 940 | Forensic incident investigation sample containing simulated IOCs. |
| Dataset-08 (Large) | LARGE | 412 MB | 6,241 | 1,105 | Enterprise mailbox export with deep forwarding chains and DKIM/SPF headers. |
| Dataset-09 (Large) | LARGE | 540 MB | 8,120 | 1,420 | High-volume transactional system notification logs. |
| Dataset-10 (Very Large) | VERY LARGE | 810 MB | 11,425 | 1,890 | Multi-year executive correspondence archive. |
| Dataset-11 (Very Large) | VERY LARGE | 986 MB | 4,820 | 2,450 | Maximum benchmark boundary sample (< 1 GB strict limit) for memory stability evaluation. |
| CUMULATIVE PLANNED CORPUS | 3,561 MB (~3.56 GB) | 48,504 Records | 12,504 Payload Blobs | 100% Client-Side Ingestion | |
Academic definitions and exact formulas governing statistical aggregation, percentile latency calculations, and verification criteria.
Formula: (Successfully Parsed Messages / Total Detected Messages) × 100
Proportion of positive identifications that were actually correct.
Proportion of actual positive conditions that were correctly identified.
50th percentile value of sorted duration sample measurements.
95th percentile value below which 95% of sample latencies fall.
Benchmarking observers exist strictly outside hot parsing loops and search indexing pathways. When Benchmark Mode is disabled, a null-object observer (NoopBenchmarkObserver) intercepts calls with zero CPU cycles or IndexedDB writes.
To prevent database write lock contention during keystroke debouncing, query timing samples are accumulated in an in-memory ring buffer (SearchBenchmarkBuffer) and bulk-added in coarse checkpoints of 50 samples or on clean session teardown.
No personal mailbox filenames, email subjects, sender/recipient addresses, or attachment content hashes are ever written to benchmark export tables. Datasets are attributed strictly by anonymized IDs and byte size classes.
Participants can use the dedicated Usability Study Portal to record task measurements and Likert evaluations anonymously.