← All incidents

minorotherunverifiedcollected

A pricing rule computed in three places was fixed in one, and sign-ups stopped (unverified report)

In the author's company's service, a member trying to complete sign-up on the web was stopped by the final check, because the amount authorized on their card did not match the contract total. The authorization was about 10% short. The member tried twice, was stopped both times, and still could not sign up after starting over a while later.

Observed
Severity score
3/10
Blast radius
uptime
Tags
#ai-review#duplicated-logic#proration#payments#partial-fix#consistency-test

Cause

Per the author, the same rule, first-month proration, was computed separately in three places: the API that issues the card authorization, the web member screen, and the LINE member screen. A change about a month earlier made the web side charge the full amount for plans set not to prorate, but the API ignored that setting and always prorated a mid-month start. The API side had been deliberately left alone in an earlier change, without being recorded as debt. The web change had passed AI review and checks in the staging environment.

Consequence

One member could not complete sign-up on the web. The authorization was released when the check failed, so no money was held. The mismatch went unnoticed for 19 days after reaching production; hitting it required four conditions at once, which in production applied only to three plans at one gym, about one person a month.

Fix

A scheduled check picked up the failure log in about 30 minutes. An AI session investigated production read-only, found the cause, wrote a failing test first and then the fix; reproduction in staging was done in about three hours and the fix shipped the same day. It was a change of a few dozen lines aligning the API with the web side. The author lists four preventive measures: consistency tests comparing the calculations, filing deferred changes as tickets, listing every implementation before changing a business rule, and revising the test matrix.

What happened

In the author’s company’s service, a member signing up on the web sees the amount on a confirmation screen, an authorization for that amount is placed on their card when they enter it, and the final confirm button recomputes the contract total and compares it with the authorization before charging. One afternoon, for a plan at one gym, that comparison failed and sign-up was stopped. The authorization was about 10% less than the contract total.

The chaos on the ground

The member tried twice and was stopped both times, and still could not sign up after starting over a while later. The authorization was released when the check failed, so no money was held. A scheduled check picked up the failure log in about 30 minutes.

The mismatch turned out to have been in production for 19 days. Hitting it took four conditions at once: a plan set to start on the 1st of each month without proration, a start date other than the 1st, a sign-up with a card authorization, and a sign-up screen that allowed a mid-month start date. The general sign-up screen limited such plans to the 1st; only a screen built for a particular gym let members pick freely. In production that meant three plans at one gym. The two people who signed up for them in that period both chose the 1st; the third was the first to choose another day.

Root cause

Per the author, the sign-up amount was computed separately in three places: the API that issues the card authorization, the web member screen, and the LINE member screen. About a month earlier the web side had been changed to follow the plan setting and charge the full amount for plans that do not prorate. The API ignored the setting and always prorated a mid-month start.

The API side was not simply forgotten. In an earlier change, the authorization estimate had been deliberately left as it was, with a comment saying the other places used the same unconditional proration and leaving it kept the diff audit simple. That decision was never recorded as debt. The web change passed AI review and checks in staging, but staging was tested on the general sign-up screen, and the screen that could hit the mismatch was not in the test matrix.

The fix

The same day, an AI session investigated production read-only, identified the cause, wrote a failing test first and then the fix. Comparing the full test suite before and after and reproducing the failure and the fix in staging took about three hours. The fix, a few dozen lines aligning the API with the web side, went to production that day. When the first report came in, an engineer suspected it must be happening more often and had another session measure it; that corrected two points, the date the mismatch began and the range of affected plans.

The author lists four preventive measures: consistency tests that compare the results wherever the same rule is computed, filing a deliberate deferral as debt along with its assumptions, listing every implementation before changing a business rule, and adding screen differences and the combinations the final check protects to the test matrix. The check itself had been added about a month earlier; before that, the same sign-ups had gone through silently.

How many places hold the same rule is outside the scope of a diff review. AI review can work correctly inside the diff and still not see the side that was never changed.

Sources