Menu

Temporary Mailbox API for Automated Tests: Create, Poll, Verify

Use a temporary mailbox API in automated tests: create an inbox over HTTP, poll for the confirmation mail, extract the code and clean up in CI.

Published

  • automation
  • test mailbox
  • continuous integration

A confirmation mail is the part of a signup flow that a browser test cannot drive on its own. Something has to own an address, receive the message and hand the code back to the test. A temporary mailbox API does exactly that: the suite asks a service to create an inbox over HTTP, reads what arrives, and discards the inbox when the case is done. There is no browser, no shared human inbox and no manual copy and paste.

This article is about the API side of that arrangement. It is not about pointing your application at a local receiving endpoint — that is how to capture email in CI, and the two solve different problems. Local capture proves what your application sends. A mailbox API proves what actually arrives, over a real delivery path, at an address the test owns.

Why use a mailbox API in tests?

Because the alternative is either a real human inbox or nothing.

A shared inbox is a poor fixture. Several runs read the same mailbox, messages from an earlier run are still sitting there, and the address accumulates traffic that no test asked for. Every assertion then has to guess which message belongs to the current case, and guessing is where flakiness is born.

Harvesting a provider’s web interface with a browser is the second bad option. It makes the test depend on markup that changes without notice, on session state, and on a login the suite has to babysit. The moment the provider restyles a button, a green suite turns red for no reason connected to the product.

A mailbox API removes both problems. The address is created for the case, it is empty by construction, and it is read through a stable interface that the test can call directly. The suite no longer cares how the provider looks, only that the contract holds: create, receive, read, delete.

There is also a privacy argument. An address created through an API stands for nobody. It is not a person’s mailbox, it is not a place a real message could land, and it is thrown away with the run.

Which endpoints does a test actually need?

A mailbox API can expose dozens of routes, but a test client needs four operations, and it helps to name them the way the suite will use them.

Create returns an address and a handle. The address is what the application under test is told to send to. The handle, often a token or an identifier, is what the test uses to ask about that inbox in every later call. The suite should treat the pair as one object and never reconstruct the handle from the address, because providers are free to make the two unrelated.

List returns summaries rather than bodies: one entry per message, with an identifier, the sender, the subject and the arrival time. This is the call a polling loop should use, because it is cheap and it is enough to answer the only question that matters early on, which is whether anything has arrived yet.

Read returns one message in full, including the text and the HTML parts. This is where the code lives, and this is the call that should be made only after list has reported a match.

Clear removes the messages from an inbox or deletes the inbox entirely. A test needs it for two reasons: to reset between attempts without creating a new address, and to clean up when the case finishes.

Some services add a wait or long-poll endpoint that holds the connection until a message arrives or a timeout elapses. It is convenient, but a client should still be able to fall back to list, because the wait call is the part most likely to be rate-limited.

Polling for the code without flakiness

The single most common mistake in this kind of test is a fixed sleep. A constant number of seconds is a guess: too short when the delivery is slow, wastefully long when it is fast, and wrong in both directions on a loaded CI runner. Replace it with a loop that calls list, checks for a match and returns as soon as one is found, with a ceiling that fails the test rather than hanging the job.

Matching is the second half of the problem. The message the test wants is the one addressed to the address the case created and, if the inbox can hold more than one kind of mail, the one whose subject carries a stable fragment. Prefer the newest match, so a duplicated delivery from a retry does not confuse the read. Never take the first message unconditionally; on a reused address that is exactly how an old code gets validated.

Extraction should be anchored. A body may contain a reference number, a timestamp and a price, and a parser that grabs the first run of digits will sometimes grab one of those instead of the code. Look for the wording that introduces the code, then read the code from its neighbourhood, and fail with the body attached when nothing matches.

Finally, respect the resend limit. A code flow usually allows only a few sends in a short window, and that limit is part of the behaviour under test. A test that presses the button again to get a fresh code will eventually be refused and fail for the wrong reason. Retry by reading the mailbox again, not by triggering another message.

Wiring it into an end-to-end or CI suite

The clean shape is a fixture. Before the flow starts, the fixture creates an inbox and returns its address. The test drives the application using that address. After the application confirms it has sent something, the assertion reads the mailbox and extracts the code. When the case ends, the fixture deletes the inbox.

Keep the client small and injectable. One module wraps the four calls; the test depends on that module, never on raw HTTP scattered through the suite. That makes it possible to swap in a fake for unit tests and to point the same suite at a different provider without rewriting assertions.

In CI, credentials belong in the job’s secret store, never in the repository and never in a log line. Give each job or each parallel worker its own inbox, and prefix generated addresses with something that identifies the run, so a stray message can be attributed by inspection. Set the client’s timeout below the job’s own timeout, so a stuck poll fails with a clear message instead of an abrupt job cancellation.

Retry the read, not the whole flow. If the code has not arrived yet, wait and read again; re-running the signup would produce a second message and, with it, a second candidate for the assertion. And keep the mailbox API out of production flows: it is test infrastructure, and a suite should never be able to send to a real customer from it.

The flow around the code, rather than the mechanics of reading it, is covered in email verification flow testing, and the parsing step specifically is the subject of OTP testing in end-to-end suites.

Isolation and cleanup

One address per case is the rule that prevents most cross-test failures. It removes the need to reason about which message belongs to whom, and it makes the freshness question disappear, because the inbox has only ever received the traffic of one case.

Cleanup should be explicit and unconditional. Delete the inbox in a teardown that runs whether the case passed or failed, not only on the happy path. Relying on the provider’s time-to-live alone is a mistake: the message may linger long enough to be read by a later run on the same machine, and that lifetime is a convenience, not a guarantee.

If cleanup fails, log it and let the suite finish. A cleanup error is worth knowing about, but it is not the same as a product defect, and failing the run for it teaches the team to ignore teardown failures. Treat anything the address received as test data: it exists for one assertion, it should not be exported or shared, and it should never be treated as anyone’s contact point.

Limits and caveats

A mailbox API is still a third-party dependency, and its limits become your limits. Per-minute rate limits can refuse a burst of creates from a large parallel run. Quotas cap how many addresses exist at once. Messages can be delayed, and a delayed message looks exactly like a missing one until it arrives.

Disposable domains are also widely blocked. A provider’s domain may be refused by the very signup form you are testing, which turns a legitimate test into a confusing failure. When that happens, the answer is not to special-case the provider but to understand whether the product under test intentionally rejects disposable addresses, and to test that behaviour deliberately instead.

The honest summary is that the mailbox API is the right tool for asserting what arrives. For asserting what your application emits, a local receiving endpoint is faster and has no quota. Most mature suites use both: local capture for the bulk of assertions, and a mailbox API only where the real delivery path is the thing under test.

Next steps

Find a code-reading test that polls a shared inbox and replace it with a fixture that creates a fresh address from a mailbox API. Log the address with the run, delete it in teardown, and watch how much of the flakiness you had accepted simply stops happening. When you need a real inbox by hand, the temp mail page creates one in a moment, and the addresses it hands out are scaffolding for a run, never a real identity.

Keep reading

Temp Mail (Disposable Email / 10 Minute Mail) guides