To capture email in tests is to stop treating mail as something that leaves the machine. Instead of letting a build send through a public provider and then polling an inbox in the outside world, the application is pointed at a receiving endpoint on the build host, and the test reads what was handed over. The result is faster, offline-capable and immune to the provider incidents that otherwise turn a green suite red.
Why a pipeline should not depend on a mail provider
A test that sends real mail through a real service has imported everything fragile about that service into its own result. Authentication can expire, quotas can be exhausted, the provider can be briefly unavailable, and none of those outcomes say anything about your code. Worse, they are indistinguishable from the failures you actually want to catch, so the team learns to re-run the pipeline instead of reading it.
There is a second cost that is easy to overlook: a pipeline that sends real mail has to be told where to send it. If that destination is a real address, every run delivers test traffic to somebody. Pointing the build at a local endpoint removes the question entirely, because nothing leaves the network.
How does a local capture endpoint work?
The mechanism is the same one any mail server uses. Your application is configured with a mail server host and port for the duration of the test run. That host is the loopback interface and the port is whichever one the harness chose for this run, so no configuration value has to be asserted as a fact about the world.
Once the application hands the message over, the capture endpoint accepts it, keeps it in memory or writes it to a file, and makes it available to the test. Nothing is relayed anywhere.
That shape brings three properties that make assertions reliable:
- The message exists before the test looks for it, because the handover is synchronous, so there is no polling loop to get wrong.
- The content is exact, including headers and encoding, because nothing rewrote it in transit.
- The message can be re-read as many times as the test needs, because it is stored rather than consumed.
What is actually inside a captured message?
More than the body, and the extra parts are where the useful assertions live. A captured message carries the envelope information from the handover together with the message itself, which means a test can check who the message claims to be from, which address it was sent to, the subject line, the content type, and the body in whichever form your application produced it.
That matters because the delivery is not the thing under test. The content is. A suite that only verifies a message was accepted passes happily while the template renders a blank name or the link points at the wrong environment.
Should a failing test keep the raw message?
Yes, and it should keep it deliberately rather than by accident. When an assertion about the body fails, the single most useful artefact is the message as it was actually produced. Without it, the next step is usually to reproduce the failure locally by hand, which is precisely the manual work the capture endpoint was meant to remove.
Two habits make the artefact useful. Attach it to the failing run rather than to a shared location, so that concurrent runs cannot overwrite each other. And treat it as test data when it is stored: a captured message may contain addresses and generated values from the run, and it should be retained on the same short schedule as the rest of the run’s output. Addresses used this way exist for assertions, never as real identities.
When you need a real mailbox instead
A local capture cannot answer questions that involve the outside world. If the test must prove that a message survives an actual delivery path, or that a third party’s system reacts to it, then something has to leave the machine.
For those cases a disposable inbox is often enough: an address created for the run, read once, discarded. The temp mail page creates one on request, and because nobody registered it, the inbox starts empty and contains nothing but the traffic the test triggered. The distinction is worth keeping clean in your head: local capture is for assertions about what your application sends, and a real mailbox is for assertions about what arrives.
For developers: ports, parallelism and assertions
Four decisions determine whether this arrangement stays boring.
Binding the capture endpoint to the loopback interface only is the first. It limits exposure and makes it obvious that the endpoint is not a mail service for anything outside the run. Let the port come from the environment rather than a constant, so two runs on one machine cannot collide.
Isolating runs is the second. Give each run its own endpoint process, or at least its own storage, and give each case its own recipient address. A shared queue read by several tests in parallel produces the most irritating class of failure there is: a test that passes alone and fails in a full pipeline.
Asserting content rather than elapsed time is the third. Because the handover is synchronous, there is nothing to wait for; a sleep in this kind of test is a sign that something is being tested through the wrong interface.
Failing loudly is the fourth. When an assertion about a message fails, print or attach the part that was compared. A test that reports only a mismatch, without showing the message it read, sends the next person back to manual reproduction.
If your capture path feeds tests that read short codes rather than bodies, OTP testing in end-to-end suites covers the parsing side, and the catch-all mailbox approach covers what to do when an environment genuinely needs a whole domain rather than a single endpoint.
Next steps
Find the one test in your suite that sends mail through a provider, and count how often it has failed for reasons unrelated to the change under test. Then give it a local endpoint for a single run and compare. Keep the public provider for whatever genuinely needs the open internet; the rest belongs on the machine that runs the tests.