CODE REVIEW / SEPTEMBER 8, 2026

Across two model configurations, Looply found 15 more reference defects.

How many known defects did each workflow find? We compared the same models and reasoning efforts on 34 PRs. Sol medium found 32 of 41 reference defects versus 17; Astra low found 28 versus 13.

34 real PRs · 41 reference findings · Golden 2.0

BENCHMARK RESULTS

The same 41 reference defects. More of them found.

Compare the same model and reasoning effort across two review workflows. Each pair adds 15 reference matches; these are separate runs on the same cohort, not 30 distinct additional defects.

Recall · 34 real PRs · 41 reference findings · Golden 2.0
5.6-Sol mediumWithout Looply
41.5%

17 / 41 reference matches · Precision 14.4% · TP / FP / FN: 17 / 101 / 24

5.6-Sol mediumWith Looply
78.0%+36.6 pp

32 / 41 reference matches · Precision 25.4% · TP / FP / FN: 32 / 94 / 9

6-Astra lowWithout Looply
31.7%

13 / 41 reference matches · Precision 14.3% · TP / FP / FN: 13 / 78 / 28

6-Astra lowWith Looply
68.3%+36.6 pp

28 / 41 reference matches · Precision 28.6% · TP / FP / FN: 28 / 70 / 13

6-Astra xhighWith Looply
82.9%

34 / 41 reference matches · Precision 21.9% · TP / FP / FN: 34 / 121 / 7

Recall uses a 0–100% scale. Percentage-point gains use the underlying counts, before rounding.

6-Astra xhigh with Looply matched 34 of 41 reference defects (82.9% recall), the highest coverage measured here. Precision was 21.9%, below Astra low’s 28.6%. There is no without-Looply xhigh run in this study.

Precision here measures agreement with an incomplete reference set—not the proportion of all valid bugs found. Additional valid bugs may be scored as unmatched. Recall is the primary measure in this comparison.

Without Looply uses Matt Pocock’s code-review skill: Standards and Spec outputs are concatenated, without merging or reranking. Matt Pocock / code-review ↗

THE DATASET

Real changes, with a history of corrections.

Looply collected real pull requests and their correction traces from codebases serving different business needs. This study evaluates a frozen Python subset, python-bug-20260902: 34 PRs with 41 reference findings. The corpus includes business systems such as ERPNext, document workflows such as Paperless, and other application codebases.

Dataset available on request

Request dataset access ↗

Source: Looply evaluation snapshot, September 8, 2026.

FROM FINDING TO FAILURE

What those additional matches look like.

Three selected Sol medium cases: Looply matched the reference finding with high judge confidence; the baseline’s Standards and Spec axes reported no findings. These are archived review results, not newly reproduced failures.

ERPNext #58260 ↗

Legacy scrap can be counted again after an upgrade.

  1. Job Card → Scrap
  2. Legacy entry → blank type
  3. Quantity not deducted

The migration types Job Card scrap but leaves historical Stock Entry types blank. Accounting keys include both item and type, so previously used scrap is not deducted and can be included in a later manufacture entry.

co_by_product_patch.py → manufacturing.py

Matched: GT-58260-01 · Looply PR totals: 1 TP / 2 FP / 0 FN

Paperless #13659 ↗

A template passes validation, then fails during execution.

  1. Compile succeeds
  2. Unknown name at render
  3. Title stays unchanged

A misspelled variable such as create_yaer is valid Jinja syntax. Compile-only validation accepts it, but the execution context cannot resolve it. Rendering fails before the workflow can update the document title.

templating/workflows.py → workflows/mutations.py

Matched: GT-13659-01 · Looply PR totals: 2 TP / 1 FP / 0 FN

Paperless #13739 ↗

One send error can stop every later heartbeat.

  1. send() raises
  2. Loop exits; error suppressed
  3. Heartbeats stop

Exception suppression wraps the entire heartbeat loop. A send error exits that loop before being swallowed, ending the background task without a retry or active connection close. The WebSocket can remain open without heartbeat traffic.

consumers.py · _heartbeat_loop

Matched: GT-13739-01 · Looply PR totals: 1 TP / 0 FP / 0 FN

HOW TO READ THE RESULTS

Reference coverage, with the scoring made explicit.

Recall
TP / (TP + FN): the share of the 41 reference defects matched by the review.
Precision
TP / (TP + FP): the share of reported findings matched to a reference defect.
Matching
Final outputs are micro-averaged across all 34 PRs. An independent gpt-5.6-luna judge at high effort performs one-to-one semantic matching against Golden 2.0.
Interpretation
The reference set is not an exhaustive inventory of valid issues. An unmatched finding (benchmark FP) is not necessarily an incorrect suggestion. All 34 samples contain known defects, so this study does not estimate false alarms on clean PRs.
Study boundaries and scoring notes
  • Matching model and effort does not isolate Looply as a single experimental variable. Review objectives, tools, context, environments and resource use differ. This is a workflow comparison, not a strict ablation, and repeated-run variance was not measured.
  • Sol Looply results combine historical successful a6 samples; Astra results combine independent runs and supplemental runs. All five reported configurations completed 34/34 PRs. Unfinished Terra and Luna configurations are excluded.
  • Wagtail #13878 has a disputed reachability premise in its reference finding. The supplied Final scores are preserved.
  • For baseline Sol on Open edX #38848, one finding matches in the combined judge output but not in the independent Spec judgment. The combined result has 17 TP; the union of independently matched axis references has 16. We retain the original combined score.
  • An ERPNext #58394 correction affects the historical Sol Hypothesis score (33/41 corrected to 32/41). This page uses Final outputs only. The Paperless template example highlights the undefined-variable path; its separate exception-path finding has an unverified external premise.
  • The baseline concatenates Standards then Spec, retaining style and code-smell candidates and cross-axis duplicates. One-to-one matching can count only one claim per reference. The Sol baseline used an isolated macOS sandbox and restricted model-service network access; the imported run records R1 provenance, not full R3 controlled-run certification.

The review workflow changes what gets found.

On this cohort, Looply increased reference coverage across both matched configurations. The concrete value is visible in the paths a review follows—from migration to accounting, from validation to execution, and from an exception to a task’s lifetime.

← Back to deck