How well does the bill check-up catch problems?
We ran the detector on 51 made-up households (Maria plus 50 synthetic ones). Each synthetic household got two hidden problems and the same set of traps that should not raise an alarm, like a summer electricity bill.
- 96/105
- Problems caught
- 0
- False alarms
- 100%
- Precision
- 91%
- Recall
01By problem type
| Type | Hidden | Caught | Missed | False alarms | Precision | Recall |
|---|---|---|---|---|---|---|
| Price jump | 22 | 19 | 3 (3 below threshold) | 0 | 100% | 86% |
| Duplicate charge | 11 | 8 | 3 (3 below threshold) | 0 | 100% | 73% |
| Creeping fee | 14 | 14 | 0 | 0 | 100% | 100% |
| New subscription | 15 | 15 | 0 | 0 | 100% | 100% |
| Promo ended | 21 | 18 | 3 (3 below threshold) | 0 | 100% | 86% |
| Unusual charge | 7 | 7 | 0 | 0 | 100% | 100% |
| Bill ≠ payment | 15 | 15 | 0 | 0 | 100% | 100% |
02Traps that should stay quiet
- Maria: summer rise every year (seasonal)no alarms · 1 tested · 1 shown as “probably fine”
- Maria: normal grocery variationno alarms · 1 tested
- Maria: flat rentno alarms · 1 tested
- Maria: flat phone billno alarms · 1 tested
- Seasonal electricityno alarms · 50 tested · 49 shown as “probably fine”
- Variable water billno alarms · 50 tested
- Grocery variationno alarms · 50 tested · 2 shown as “probably fine”
- One-off large purchaseno alarms · 50 tested
- Annual renewalno alarms · 50 tested
- New subscription with emailno alarms · 23 tested
03How to read this
- An alarm is a High or Medium finding. Low findings are shown as “probably fine” and are not counted.
- Some hidden problems are deliberately smaller than the rules’ thresholds (for example an 8% price rise or a repeat charge 7 days later). Missing those is expected, and they are counted separately.
- The households are generated by the same team that wrote the rules, with a fixed seed (4242). This checks that the rules behave as designed. It is not a measure of accuracy on real bank data.
- Run it yourself with
npm run eval:detector.