Evidence record 18694 · automatically gathered

SCOUT: Symmetric Consensus Outlier Detection for Failure Localization in LLM Pre-Training

In LLM pre-training, synchronization propagates rank-local stalls, slowdowns, and numerical errors into job-wide symptoms, obscuring their origin. Existing diagnosis often relies on in-process monitors that cannot report after the trainer blocks or terminates, or on post-mortem logs that preserve only synchronized symptoms; offline health tests lose the workload and operating conditions that triggered the failure. We present SCOUT, a unified runtime failure-localization framework built on one de

Record details

Published: 11 August 2026
Source: arXiv cs.LG
Category: Research
Topics: Jobs & economy · Healthcare
Retrieved: 12 August 2026

source-onlyevidence status

These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.

How to cite this record

ethics.ai (11 August 2026), “SCOUT: Symmetric Consensus Outlier Detection for Failure Localization in LLM Pre-Training,” evidence record 18694, https://ethics.ai/record/18694 (originally published by arXiv cs.LG).

JSON

Use and limitations

This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.