Menu

OTP Testing in End-to-End Suites: Reading the Code

OTP testing means getting the one-time code out of a test mailbox and back into the form. This covers parsing, freshness, retries and the human fallback.

Published

  • end-to-end tests
  • one-time codes
  • automation

OTP testing is the small, stubborn part of an end-to-end suite that reads a one-time code from a mailbox and types it back into a form. It looks trivial until it is unreliable, and it becomes unreliable for reasons that have nothing to do with the code being wrong: a slow message, an older message still sitting in the inbox, a resend that was rate-limited. This article covers how to make that step dependable and when to stop automating it.

What an end-to-end suite needs from a code message

A code message has one job: carry a short secret from your application to a place the test can read it. Everything else about it — layout, branding, footer, language — is decoration from the suite’s point of view.

That asymmetry is what makes the step so easy to get wrong. Because the test only needs a handful of characters, it is tempting to scrape them loosely, from the first message that looks close enough. The result is a test that usually passes and occasionally reads the wrong message, which is the most expensive kind of flake: it is rare enough to be dismissed and real enough to hide a genuine bug.

Which message is the right one?

Identity, not recency alone. The safest rule is a combination: the message must be addressed to the address this case created, and it must be the newest matching message in that inbox.

Both halves earn their place. Matching on the recipient removes cross-test contamination, because a message sent to another case’s address cannot satisfy the assertion however new it is. Preferring the newest among matches handles the case where the same flow was triggered twice, once by a retry or an earlier attempt.

Subject matching is a useful narrowing step when a mailbox legitimately contains more than one kind of mail — a welcome message and a code, for example. Match on a stable part of the subject rather than the whole line, because the decorative parts change: the same subject with a different greeting should still match.

Why do tests that read codes go flaky?

Four causes account for most of it, and each has a different fix.

  • A fixed wait. Sleeping for a constant number of seconds is a guess, and guesses are either too short when the message is slow or wastefully long when it is fast. Poll until the message appears, with a ceiling that fails the test rather than hanging the run.
  • A stale message. If a case reuses an address, or if a previous attempt left a message behind, a fresh lookup can satisfy itself with old content. Give each case its own address, or at minimum record what was there before the flow started and ignore anything older.
  • A resend limit. Code flows normally allow only a few resends in a short window, which is a deliberate protection rather than a defect. A test that retries by asking for another code will eventually be refused and fail for the wrong reason; retry by re-reading the mailbox instead of by pressing the button.
  • Parsing that is too eager. A body may legitimately contain several numbers — a reference, a timestamp, a price — and a parser that grabs the first digit run will sometimes grab the wrong one. Anchor the extraction to the wording around the code rather than to the digits alone.

When must a human stay in the loop?

Some steps should not be automated, and pretending otherwise produces tests nobody trusts.

A human is the right tool when the flow needs a device that a pipeline does not have, when the code arrives through a channel that cannot be read programmatically, or when the assertion is really a judgement about whether the message reads well. There is also a simpler case: some providers actively discourage automated signup, and a suite that works around that is not testing the product, it is testing the workaround.

The useful pattern is an explicit manual gate that fails loudly instead of silently passing. A step that says “a person must check this before the run continues” is honest; a step that tries and occasionally succeeds is not.

A throwaway inbox for each run

The cleanest fixture for this work is a mailbox created for the case that needs it and discarded afterwards. Because nothing else has ever been sent there, the newest message is almost certainly the right one, and the freshness problem largely disappears.

The temp mail page creates an address on demand: choose a mail domain, optionally give it a prefix, and an inbox appears that lists what arrives. Codes are surfaced separately so they can be copied without reading the whole body, which is convenient when a person is doing the reading and useful to imitate when software is.

These addresses are scaffolding for a test run. They stand for nobody, they are discarded with the run, and they must never be recorded as a real identity or used as anybody’s contact point.

For developers: parsing, freshness and retry ceilings

Five decisions make the difference between a suite that reads codes and a suite that is trusted about it.

Give each case its own address, and make the address carry the case identity, so a message can be attributed by inspection rather than by timing.

Poll with a ceiling, and make the ceiling part of the failure message. When it trips, the report should say which address was read, how many messages were in it and what subjects they had. Without that, the next person re-runs the pipeline instead of diagnosing it.

Parse defensively. Take the newest match on the recipient and the expected subject, then extract the code from the context around it. If no code is found, fail with the body rather than with a generic timeout.

Never disable a rate limit to make a test pass. The limit is part of the behaviour under test, and a suite that switches it off is no longer describing the product your users meet. Handle the limit instead: choose an address that has not been used recently, and read the existing message rather than requesting a new one.

Keep the human path alive. For the cases where automation genuinely cannot read the code, document the manual step and make it visible in the run output, so that a skipped check is never mistaken for a passed one.

The verification flow testing article covers the flow around the code, and how temporary mail works explains why a slow or duplicated message is normal rather than broken.

Next steps

Find the code step in your suite and check the two things most often missing: whether the address is unique to the case, and whether the failure message would let somebody else diagnose the run without repeating it. Then run the same case with an address created fresh from the temp mail page and see whether the flakiness you have been tolerating disappears.

Keep reading

Temp Mail (Disposable Email / 10 Minute Mail) guides