DATASET RUNNER
Batch process bulk claims to evaluate general regression metrics and detect formatting degradation.
Drag and drop file here, or click to upload
Accepts CSV, JSON, or Excel file formats. Maximum file payload limit: 50MB.
Active Evaluation Datasets
Select a file package and configure shadow runner constraints.
| File Name | Total Rows | Upload Date | Execution Progress | Job Status | Action |
|---|---|---|---|---|---|
| evaluation_claims_pii_10k.json | 10,000 | 2026-07-01 14:23:10 | 100% | Completed | |
| production_audit_claims_sql_inj.json | 2,500 | 2026-07-02 09:12:00 | 68% | Running | |
| medical_compliance_assertion_v2.csv | 4,200 | 2026-07-02 11:45:30 | 0% | Idle | |
| abusive_jailbreak_claims_regression.xlsx | 1,500 | 2026-06-28 17:34:11 | 42% | Failed |
Historical Batch Run Log
Aggregated metrics of past dataset runs processed on the cluster.
| Dataset Name | Model Classifier | Execution Date | Total Mismatches | Avg Latency | Success Rate |
|---|---|---|---|---|---|
| evaluation_claims_pii_10k.json | Qwen 2.5 3B | 2026-07-01 15:44:22 | 145 rows | 184ms | 99.8% |
| jailbreak_hard_eval_v1.json | Llama 3.2 3B | 2026-06-30 22:15:09 | 412 rows | 238ms | 98.4% |
| standard_policy_smoke_test.csv | Qwen 2.5 3B | 2026-06-29 10:30:00 | 12 rows | 179ms | 100% |