An ASP.NET Core endpoint can exceed its configured request timeout and still return 200 OK. The timeout policy may be installed correctly. The request token may even be cancelled. If the application ignores that token, its work can continue.
This is a useful failure to reproduce before putting a timeout on an expensive endpoint. A configuration value says when the application should stop spending effort on a request. It does not prove that the database call, HTTP request, retry delay, or computation underneath it will stop.
We will compare two endpoints that call the same service. One forwards cancellation; the other deliberately discards it. Then we will separate a server timeout from a disconnected client and inspect what actually happens to the work.
The example was built with .NET SDK 10.0.101 and tested on ASP.NET Core 10.0.1 over real loopback HTTP connections. These are the recorded test versions, not a claim that they are the newest servicing releases. The service uses Task.Delay as controlled simulated I/O; it does not exercise a database provider.
Cancellation is a signal that needs a receiver
A cancellation token communicates a request to stop. The operation receiving it decides where and how to observe that request. An async API may observe it while waiting; a computation may check it between units of work. .NET describes this as cooperative cancellation.
For the application below, the chain is small enough to audit:
request timeout or detected disconnect
-> HttpContext.RequestAborted
-> endpoint CancellationToken parameter
-> ReportService.GenerateAsync(..., token)
-> Task.Delay(..., token)
Minimal APIs bind a CancellationToken parameter to the request’s abort token. That makes the endpoint parameter useful without constructing a new token source in every handler. See Minimal API binding for special types.
The fragile part is the next method call. A service can accept the token and still omit it when calling its dependency. A repository can offer an optional token that every caller leaves at its default. The method signatures look cancellation-aware, but the operation doing the waiting never receives the signal.
In this example the token parameter is required all the way down. That does not make misuse impossible, but it makes the deliberate CancellationToken.None substitution visible in a code review.
Build a small app with one intentionally broken route
Create an empty .NET 10 web app:
dotnet new web -n TimeoutLab --framework net10.0
cd TimeoutLab
Replace Program.cs with the following complete program. The synthetic report always contains 42 rows; only the waiting and cancellation behavior is under test.
var builder = WebApplication.CreateBuilder(args);
builder.Services.AddRequestTimeouts(options =>
options.AddPolicy("interactive", TimeSpan.FromMilliseconds(250)));
builder.Services.AddScoped<ReportService>();
var app = builder.Build();
app.UseRouting();
app.UseRequestTimeouts();
app.MapGet("/health", () => Results.Ok(new { status = "ready" }));
app.MapGet("/reports", GenerateReport).WithRequestTimeout("interactive");
app.MapGet("/without-timeout", GenerateReport);
// Deliberately broken: the endpoint has a policy but discards its token.
app.MapGet("/ignores-token", async (
string operationId,
int delayMs,
ReportService reports,
HttpContext context) =>
{
if (delayMs is < 0 or > 5000)
return Results.BadRequest("delayMs must be between 0 and 5000.");
var report = await reports.GenerateAsync(
operationId, delayMs, CancellationToken.None);
return Results.Ok(new
{
report,
cancellationObserved = context.RequestAborted.IsCancellationRequested
});
}).WithRequestTimeout("interactive");
app.Run();
static async Task<IResult> GenerateReport(
string operationId,
int delayMs,
ReportService reports,
CancellationToken cancellationToken)
{
if (delayMs is < 0 or > 5000)
return Results.BadRequest("delayMs must be between 0 and 5000.");
var report = await reports.GenerateAsync(
operationId, delayMs, cancellationToken);
return Results.Ok(report);
}
public sealed class ReportService(ILogger<ReportService> logger)
{
public async Task<Report> GenerateAsync(
string operationId,
int delayMs,
CancellationToken cancellationToken)
{
logger.LogInformation("work.started {OperationId}", operationId);
try
{
// Controlled stand-in for an async dependency, not a database test.
await Task.Delay(delayMs, cancellationToken);
logger.LogInformation("work.completed {OperationId}", operationId);
return new Report(operationId, 42);
}
catch (OperationCanceledException) when (cancellationToken.IsCancellationRequested)
{
logger.LogInformation("work.cancelled {OperationId}", operationId);
throw;
}
finally
{
logger.LogInformation("work.exited {OperationId}", operationId);
}
}
}
public sealed record Report(string OperationId, int Rows);
Start it without attaching a debugger:
dotnet run --no-launch-profile --urls http://127.0.0.1:5087
The timeout middleware skips its timeout handling when a debugger is attached. A Release build with an attached debugger is therefore not the same test condition. The request timeout documentation also specifies that explicit routing must precede UseRequestTimeouts, and that a policy or default timeout must be configured before limits apply.
Here the interactive policy is attached to two routes. The third route, /without-timeout, deliberately has no server deadline. It gives us a control case and a way to test client disconnection without confusing it with the timer.
The 250 ms timeout and five-second input cap are laboratory values. The intentionally broken route and the unbounded-by-policy control route are for a local demonstration, not endpoints to expose as a production reporting service.
Run the same slow operation through both routes
A quick request should complete normally:
curl -i 'http://127.0.0.1:5087/reports?operationId=fast&delayMs=20'
Now ask the cooperative route for two seconds of work:
curl -i 'http://127.0.0.1:5087/reports?operationId=slow&delayMs=2000'
In the test run this returned 504 Gateway Timeout. The application log contained the following sequence:
work.started slow
work.cancelled slow
work.exited slow
There was no work.completed slow event. That observation matters more than the response code alone: the service reached its cancellation path and left its scope.
Next, call the route that drops the token:
curl -i 'http://127.0.0.1:5087/ignores-token?operationId=ignored&delayMs=600'
This returned 200 OK, with the body:
{"report":{"operationId":"ignored","rows":42},"cancellationObserved":true}
The timer had expired, as the boolean demonstrates. Nevertheless, the service logged completion because its delay received CancellationToken.None. The endpoint then returned a normal result.
Finally, the control route can finish a 600 ms operation without triggering the named policy:
curl -i 'http://127.0.0.1:5087/without-timeout?operationId=no-policy&delayMs=600'
It also returned 200 OK. Registering the middleware had not silently assigned the named policy to every route.
Why the middleware does not force a 504 response
The distinction is visible in the middleware source for the tested runtime. It substitutes a linked request token and awaits the downstream pipeline. Its timeout response path catches an OperationCanceledException under specific cancellation conditions; it does not replace every late successful result with an error.
The cooperative route causes its delay to throw when the token is cancelled. The service logs the event and rethrows. The broken route produces no cancellation exception from the delay, so the normal result can flow back through the middleware.
Two other conditions affect that response path. The original request-abort token must not already be cancelled, and the response must not have started. If headers have already been sent, the middleware cannot clear the response and substitute a new 504 status. This source-level behavior explains why a buffered JSON endpoint and a streaming response need different tests.
This also makes exception handling part of timeout correctness. A broad catch inside application code can consume a cancellation exception and return a fallback success. A custom exception wrapper can translate it into an unrelated error. Either change may alter the result observed by the caller while leaving the policy registration untouched.
A disconnected client is a different event
A browser closing a connection and the server reaching its own deadline are separate reasons to stop work. At the endpoint, however, cancellation of the linked request token alone does not identify which reason occurred.
ASP.NET Core exposes RequestAborted for request abortion, and Microsoft recommends passing it to long-running operations. Connection abortion must first be detected by the server; a user’s decision to leave a page is not an instantaneous signal across every intermediary. See the HttpContext guidance.
The disconnect test used /without-timeout with a five-second delay. It waited for the service’s start event, then reset the local TCP connection. The service logged cancellation and exit without logging completion. Because that route had no server deadline, this result could not be attributed to the 250 ms policy.
There is no useful assertion that the disconnected caller receives a 504: the connection is already gone. The assertion belongs on the server side, where we can observe whether unwanted work stopped.
Behind a proxy, repeat this test through the actual network path. A proxy that stops waiting for an upstream response may behave differently from a direct client connection. This laboratory test verifies local Kestrel behavior; it does not establish the disconnect behavior of an ingress controller or load balancer.
Carry the token to the operation that owns the wait
Replacing Task.Delay with a real dependency is where the example becomes an integration task. Inspect every await in that path: connection acquisition, command execution, response-body reading, retry backoff, and any queue admission before the dependency call.
EF Core accepts cancellation tokens on its asynchronous operations and passes them to the database provider. A provider may or may not honor them. Microsoft explicitly documents that limitation in EF Core’s async guidance. A cancelled fake delay is therefore insufficient evidence that a particular SQL query stops consuming database resources.
For a database-backed endpoint, run a controlled slow query against the actual provider. Cancel the request, observe the client-side outcome, and inspect the database side for continuing work. Also verify connection cleanup and pool recovery under repeated cancellation. Those checks address a different boundary from merely observing an exception in the application.
An HTTP client needs the same scrutiny. Passing the token when sending a request is only part of the path if the application separately streams or deserializes the body. Keep the intended cancellation scope visible at each step, and check the client library’s contract for the exact overload being used.
A timeout on a wait has another subtle limit. Task.WaitAsync returns a task representing the wait completing, timing out, or being cancelled. It does not add cancellation support to the underlying operation. See the WaitAsync API. If the operation continues after its caller stops waiting, somebody still owns its eventual result, exception, and resources.
The deadline is a budget for the remaining path
The timer in this example starts in the timeout middleware. It does not account for time the request spent at a client, gateway, or earlier admission point.
Suppose a caller allows one second overall, but 400 ms has already been spent before the endpoint starts its dependency call. Giving that dependency another full second changes the effective end-to-end budget. Repeating the same mistake on a retry compounds it.
Choose where the request budget begins, then account for work already spent. If a dependency needs a smaller local limit, link that local cancellation source to the request token instead of replacing the request’s cancellation path. Dispose of the source when its owner finishes.
For retries, include backoff and the next attempt in the remaining budget. Do not reset a full request-sized timeout for every attempt and call the resulting behavior bounded. The useful question is whether another attempt has enough time to do meaningful work and return a result.
A token communicates cancellation, not a timestamp or an explanation. If business logic needs remaining time or a reason code, represent that separately and define how it is propagated. Avoid inferring an exact chronology from whichever token flag happens to be checked first during a race.
Stopping a request does not reverse a completed write
Our example is intentionally free of external side effects. A reporting delay can be abandoned without reconciling a write. A payment submission, an email send, or a committed database transaction is different.
Imagine an external operation succeeds just before the response connection disappears. The server may have completed the effect while the caller cannot tell whether it happened. Reporting cancellation does not settle that uncertainty, and blindly replaying the entire request can repeat the effect.
Define the recoverable business operation before choosing the retry policy. Stable operation identifiers, durable outcome records, and idempotency can make retries safe enough for that operation. The earlier article on idempotent .NET RabbitMQ consumers explores a related repeated-delivery problem.
If work must outlive the request, give it an explicit durable handoff and a worker that owns its lifetime. A detached task launched from the endpoint is not such a handoff: it can retain request-scoped dependencies and still disappear when the process exits.
What the verification established
The local integration suite built the exact program above and started Kestrel on an ephemeral loopback port. Seven tests passed:
- A short cooperative request returned the expected report.
- A slow cooperative request returned 504 and logged cancellation and scope exit.
- The token-discarding route returned 200 after its request token had been cancelled.
- A route without the named policy completed beyond 250 ms.
- A TCP reset cancelled work on the route without a server deadline.
- Out-of-range delays returned 400 before entering the service.
- A timed-out request did not cancel a concurrent successful request.
The assertions checked response bodies and application events, not an exact millisecond cutoff. Local scheduling and timer delivery make a stopwatch threshold a poor substitute for observing the cancellation path. The build completed without warnings or errors.
For a production endpoint, keep that two-sided view: what did the caller observe, and what happened to the work after the caller stopped waiting? A timeout setting is useful only when the path underneath it has an answer to both.
What do you think?