Engineering Evaluation Protocol

MBOX VIEWER TESTING & EVALUATION

Comprehensive engineering evaluation of parsing accuracy, browser streaming performance, forensic correctness, local-first reliability, and workflow usability across standardized MBOX email archives.

Zero-Touch Runtime
Passive Observational Subsystem
Forensic Correctness
Ground-Truth Validated Metrics
Strict Auditability
Target vs. Measured Transparency
EVALUATION PROTOCOL STATUS
TESTING IN PROGRESS

These values describe the planned benchmark corpus and evaluation methodology. Accuracy and performance results are published after measurement and validation.

Strict State: PLANNED / MEASURED
Planned Evaluation CorpusTARGET METHODOLOGY
MBOX Datasets
11
Email Records
48,504
Cumulative Size
3.56 GB
Participants
15

Planned across 5 dataset size classes (Tiny to Very Large) strictly bounded below the 1 GB browser Web Worker memory ceiling.

Measured So Far (Local Runtime)AWAITING DATA
Tested Datasets
0
Parsed Records
0
Ingested Size
0 MB
Participants
0

Real-time observations aggregated from local browser storage when Benchmark Mode is enabled. No data leaves your local machine.

STANDARDIZED EVALUATION CORPUS (11 DATASETS)

Planned dataset sizing and record counts across 5 memory classes (Tiny to < 1 GB Very Large).

Dataset ID / NameSize ClassSize (MB)Planned RecordsAttachmentsMethodological Description
Dataset-01 (Tiny)TINY18 MB38461Small individual mailbox for quick unit and smoke verification.
Dataset-02 (Tiny)TINY24 MB51283Personal archive sample with standard MIME headers.
Dataset-03 (Small)SMALL48 MB1,042145Corporate team correspondence with mixed HTML/plain body structures.
Dataset-04 (Small)SMALL85 MB1,890310Multi-language character sets and RFC2047 encoded headers.
Dataset-05 (Medium)MEDIUM142 MB3,150520Heavy attachment corpus with embedded images and PDFs.
Dataset-06 (Medium)MEDIUM215 MB4,820780Mailing list archive with complex thread hierarchies.
Dataset-07 (Medium)MEDIUM284 MB6,100940Forensic incident investigation sample containing simulated IOCs.
Dataset-08 (Large)LARGE412 MB6,2411,105Enterprise mailbox export with deep forwarding chains and DKIM/SPF headers.
Dataset-09 (Large)LARGE540 MB8,1201,420High-volume transactional system notification logs.
Dataset-10 (Very Large)VERY LARGE810 MB11,4251,890Multi-year executive correspondence archive.
Dataset-11 (Very Large)VERY LARGE986 MB4,8202,450Maximum benchmark boundary sample (< 1 GB strict limit) for memory stability evaluation.
CUMULATIVE PLANNED CORPUS3,561 MB (~3.56 GB)48,504 Records12,504 Payload Blobs100% Client-Side Ingestion

EVALUATION PROTOCOL & MATHEMATICAL METHODOLOGY

Academic definitions and exact formulas governing statistical aggregation, percentile latency calculations, and verification criteria.

Statistical Formulas

Parsing Success Rate(Successful / Total) * 100

Formula: (Successfully Parsed Messages / Total Detected Messages) × 100

PrecisionTP / (TP + FP)

Proportion of positive identifications that were actually correct.

RecallTP / (TP + FN)

Proportion of actual positive conditions that were correctly identified.

Median Latencysorted[middle]

50th percentile value of sorted duration sample measurements.

P95 Latencysorted[ceil(0.95 * N) - 1]

95th percentile value below which 95% of sample latencies fall.

Observational Protocol

Zero Core Execution Impact

Benchmarking observers exist strictly outside hot parsing loops and search indexing pathways. When Benchmark Mode is disabled, a null-object observer (NoopBenchmarkObserver) intercepts calls with zero CPU cycles or IndexedDB writes.

Search Latency Buffering

To prevent database write lock contention during keystroke debouncing, query timing samples are accumulated in an in-memory ring buffer (SearchBenchmarkBuffer) and bulk-added in coarse checkpoints of 50 samples or on clean session teardown.

Anonymous Dataset Attribution

No personal mailbox filenames, email subjects, sender/recipient addresses, or attachment content hashes are ever written to benchmark export tables. Datasets are attributed strictly by anonymized IDs and byte size classes.

Conducting an evaluation study?

Participants can use the dedicated Usability Study Portal to record task measurements and Likert evaluations anonymously.

Open Usability Portal →