An experimental reliability-engineering pipeline built around the Jev decision model passed 80 of 105 diagnoses in a set of Kubernetes fault scenarios, according to results published by SREGym. The 76.2% pass rate came from five runs across each of 21 faults in the historical September 4 SREGym-Lite cohort. Median diagnosis time was 14.6 seconds.

The system differs from a conventional language-model agent. A programmatic collector first reads Kubernetes objects, events, recent pod logs and resource use. It groups observations by component and summarizes suspected failures. Jev then chooses among options supplied by the pipeline: which component appears most relevant, whether it is a cause or a downstream victim, and which evidence best supports the diagnosis. The pipeline, rather than Jev, assembles the final report.

That constrained design worked particularly well when the collector surfaced a clear mismatch. In one social-network scenario, newly created nginx-thrift pods had a 16Mi memory limit even though the deployment template specified 256Mi. The collected evidence also identified a matching admission webhook. Jev selected the affected component and supporting observation, and all five attempts passed the evaluation rubric.

The failed cases illustrate the cost of narrowing the evidence too aggressively. In an Astronomy Shop test, Jev correctly focused on a saturated frontend proxy but blamed a low CPU limit rather than the expensive request-filter rule that triggered the load. Each attempt scored 0.67, below the study's 0.70 passing threshold. In a hotel-reservation scenario, the diagnosis recognized an overloaded rate service but missed the reinforcing interaction between queues, timeouts and retries because the pipeline never supplied enough measurements about that loop.

SREGym says the Jev pipeline ran roughly seven times faster and cost about 200 times less per diagnosis than the comparison language-model agent, while its 76.2% result was close to the comparison system's 77.8%. Those figures come from the project's own controlled evaluation and should not be read as a broad production benchmark.

The experiment's main lesson is architectural: a fast selector can be useful only when the surrounding collector offers the right facts and answer choices. The authors suggest it could serve as a first diagnostic pass, with a more flexible language-model agent handling incidents that require wider investigation or reasoning across multiple services.

The consistency of the outcomes was also notable: each fault either passed all five attempts or failed all five, and 18 faults received the same score in every run. That pattern suggests the supplied representation of cluster state strongly shaped the result. Coarse summaries can conceal a decisive detail, while exhaustive configuration dumps can bury it in noise.