Middleware¶
Middleware wraps the request/response cycle — it runs before a handler sees the request and after it produces a response. Use it for cross-cutting concerns: CORS, compression, security headers, logging.
Adding middleware¶
app.add_middleware() accepts middleware in two forms — a configured
instance, or a class together with its keyword options:
from veloce import CORSMiddleware, Veloce
app = Veloce()
# Instance form — build the middleware, then add it.
app.add_middleware(
CORSMiddleware(
allow_origins=["*"],
allow_methods=["GET", "POST", "PUT", "DELETE"],
)
)
# Class form — pass the class and its options; Veloce constructs it.
app.add_middleware(CORSMiddleware, allow_origins=["*"])
Middleware can also be passed when constructing the app, via the
middleware=[...] argument to Veloce(...).
Veloce middleware vs ASGI middleware¶
add_middleware accepts two distinct shapes, and it tells them apart by
what the class subclasses:
- A
Middlewaresubclass (or instance) is Veloce-native middleware. It definesprocess_request(request)and/orprocess_response(request, response)and runs inside Veloce's own pipeline, with access to the parsedRequestand per-route exclusion. Every built-in in the table below is this shape. - Any other class is treated as a standard ASGI middleware: Veloce
constructs it as
MiddlewareClass(app, **options)when the ASGI stack is assembled, so it wraps the whole application at the scope/receive/send level. This is the seam for third-party ASGI middleware — tracing, profiling, observability — that expects to wrap an ASGI app.
from veloce import Middleware, Request, Response, Veloce
app = Veloce()
# Veloce-native: split request/response hooks.
class StampMiddleware(Middleware):
async def process_response(self, request: Request, response: Response) -> Response:
response.headers["X-Stamped"] = "1"
return response
app.add_middleware(StampMiddleware)
# ASGI: a class taking (app, **options); Veloce passes the wrapped app in.
# `SomeTracingMiddleware` here stands in for any third-party ASGI middleware.
app.add_middleware(SomeTracingMiddleware, service_name="api")
Note
Veloce-native middleware runs against the parsed Request/Response,
so it is the right place for almost everything. Reach for an ASGI
middleware class only when you are plugging in a third-party component
that is already written to the ASGI interface.
BaseHTTPMiddleware goes through add_http_middleware
A BaseHTTPMiddleware subclass is a
dispatch-shape middleware, not an ASGI app. Passing one to
add_middleware raises TypeError — register it with
add_http_middleware instead.
CORS preflight and Private Network Access¶
CORSMiddleware answers a preflight (OPTIONS with an Origin) with a
204. A preflight whose Origin is not in the allow-list, or whose
Access-Control-Request-Method is not in allow_methods, gets a
diagnostic 400 instead of a silently-blocked 204 so the rejection is
visible to developers.
Set allow_private_network=True to participate in
Private Network Access:
when a preflight carries Access-Control-Request-Private-Network: true,
the response echoes Access-Control-Allow-Private-Network: true. The grant
is opt-in and never emitted unless configured.
app.add_middleware(
CORSMiddleware(
allow_origins=["https://app.example.com"],
allow_private_network=True,
)
)
For the full parameter table — allow_origins, allow_origin_regex,
allow_methods, allow_headers, allow_credentials, expose_headers,
max_age — and the credentials/wildcard rule, see CORS.
Built-in middleware¶
Veloce ships the following middleware, all importable from the top-level
veloce package:
| Middleware | Purpose |
|---|---|
CORSMiddleware |
Cross-Origin Resource Sharing |
CompressionMiddleware |
Response compression, negotiated across zstd / brotli / gzip |
GZipMiddleware |
Response compression, gzip only |
CSRFMiddleware |
Double-submit-cookie CSRF protection |
SessionMiddleware |
Signed, timestamped session cookies |
ServerSessionMiddleware |
Server-side sessions; the cookie carries only an opaque id |
TrustedHostMiddleware |
Host-header allow-list |
HTTPSRedirectMiddleware |
Redirect plain HTTP to HTTPS |
SecurityHeadersMiddleware |
Attach common hardening response headers to every response |
CSPMiddleware |
Content-Security-Policy with a per-request nonce and report-only support |
ConditionalGetMiddleware |
Emit 304 Not Modified for satisfied GET/HEAD preconditions |
RateLimitMiddleware |
Per-client rate limiter with a selectable algorithm and backend |
WebSocketOriginMiddleware |
Reject cross-site WebSocket handshakes (CSWSH) |
LoggingMiddleware |
Structured request/response access logging |
RequestIDMiddleware |
Assign a unique request ID and echo it in the response |
ProxyFix |
Honour X-Forwarded-* from trusted proxies |
The base classes Middleware and BaseHTTPMiddleware
are also exported, along with the rotate_csrf_token helper used with
CSRFMiddleware.
SessionMiddleware and ServerSessionMiddleware have a dedicated guide —
see Sessions. Cookie attributes are constructor arguments
(secure=, httponly=, samesite=, cookie_name=), not app.config keys.
SESSION_COOKIE_* keys are retired and stop startup
Setting SESSION_COOKIE_NAME, SESSION_COOKIE_SECURE,
SESSION_COOKIE_HTTPONLY or SESSION_COOKIE_SAMESITE now raises
AuditFailed at startup rather than being ignored. Pass the value to the
middleware instead — see Sessions.
Trusted hosts¶
TrustedHostMiddleware validates the Host header against an allow-list and
rejects anything else with a 400 — defence against Host-header injection. It
takes a single positional allowed_hosts list and supports literal names, the
catch-all *, and subdomain wildcards like *.example.com (which match
api.example.com but never the bare example.com):
from veloce import TrustedHostMiddleware, Veloce
app = Veloce()
app.add_middleware(
TrustedHostMiddleware(allowed_hosts=["example.com", "*.example.com"])
)
No www_redirect
Unlike some other frameworks, TrustedHostMiddleware does not redirect a
bare apex host to its www. form — there is no www_redirect option. The
middleware only allows or rejects; to canonicalise a host, add an explicit
redirect in a handler or a before_request hook.
Content-Security-Policy with a nonce¶
CSPMiddleware emits a Content-Security-Policy (and/or
Content-Security-Policy-Report-Only) header, optionally with a fresh
per-request nonce. Pass policy as a string template containing the
literal {nonce} placeholder, or as a directive mapping where the
'nonce' source is substituted with the generated nonce:
from veloce import CSPMiddleware
app.add_middleware(
CSPMiddleware(
policy={"default-src": "'self'", "script-src": ["'self'", "'nonce'"]},
report_only_policy="default-src 'self'",
)
)
Read the nonce inside a handler with
csp_nonce(request), or write
{{ csp_nonce }} straight into a template, and place it on the matching
<script>/<style> tags as nonce="...":
The nonce is materialised lazily on first read, so a request that never embeds one pays no extra cost.
A missing nonce fails silently
A <script> whose nonce does not match the policy is refused by the
browser with no server-side signal - no error in the response, the log,
or the test suite, since TestClient does not enforce CSP. If inline
scripts stop running, check the rendered nonce="" attribute first. A static, nonce-free policy can stay on
SecurityHeadersMiddleware; use CSPMiddleware when you need a nonce or a
report-only policy.
Conditional GET¶
ConditionalGetMiddleware evaluates If-None-Match / If-Modified-Since
against a buffered GET/HEAD response and downgrades a matching request
to 304 Not Modified with an empty body
(RFC 9110 §13). With
auto_etag (the default) it also synthesises a weak ETag for a buffered,
non-empty 200 that lacks one. Register it after GZipMiddleware so a
synthesised ETag reflects the compressed bytes:
from veloce import ConditionalGetMiddleware, GZipMiddleware
app.add_middleware(GZipMiddleware())
app.add_middleware(ConditionalGetMiddleware())
StreamingResponse bodies are not buffered for ETag synthesis.
Choosing a content coding¶
GZipMiddleware offers gzip and nothing else, so a browser sending
Accept-Encoding: gzip, deflate, br, zstd — which every current browser does —
is served the oldest coding it offered. CompressionMiddleware offers the
newer ones too and picks per response:
from veloce import CompressionMiddleware, Veloce
app = Veloce()
app.add_middleware(CompressionMiddleware())
Brotli and zstd each need a package, so install what you want to offer:
A coding whose package is missing is simply not offered — the middleware still serves gzip. Asking for only a missing one raises at startup, naming the package, rather than quietly serving every response uncompressed.
On an 18 KB JSON response, measured on the machine this was written on:
| Coding | Level | Size | vs gzip | Encode |
|---|---|---|---|---|
| gzip | 6 | 1,627 B | — | 149 µs |
| br | 4 | 735 B | 2.2x smaller | 329 µs |
| zstd | 3 | 834 B | 2.0x smaller | 34 µs |
Your payloads will differ — compression ratio is a property of the data — so treat this as the shape of the trade-off, not a promise.
Brotli's default quality is not a serving default
brotli defaults to quality 11. On the response above that costs 81x
the CPU of quality 4 to save a further 1.8% of size — a setting for assets
compressed once at build time, not for a response being produced now. The
middleware defaults br to quality 4. Override per coding with levels=,
whose scales differ (gzip 1–9, brotli 0–11, zstd 1–22):
Which coding is chosen follows the client first. Accept-Encoding q-values
rank the candidates (RFC 9110 §12.5.3); among equally-weighted ones the
algorithms order decides, so a deployment states its own preference for the
clients that express none:
# Prefer speed: zstd first, gzip as the fallback. No brotli offered at all.
app.add_middleware(CompressionMiddleware(algorithms=("zstd", "gzip")))
q=0 is a refusal, and a request refusing everything on offer is served
uncompressed. Every response carries Vary: Accept-Encoding whether or not it
was compressed, so a cache never serves one client's coding to another.
Added in version 0.18.0
CompressionMiddleware. GZipMiddleware is unchanged — it now subclasses
it with the coding pinned to gzip, so existing stacks behave exactly as
before.
Streaming compression¶
GZipMiddleware also compresses streaming responses chunk-by-chunk
through a single deflate stream, so a long-running streamed body no longer
has to be buffered to be compressed. Chunks at or above
min_stream_chunk_offload bytes (32 KiB by default) are offloaded to the
thread pool; latency-sensitive types (text/event-stream by default, via
latency_sensitive_types) are passed through uncompressed so server-sent
events are never merged or delayed.
When compression uses the thread pool¶
The same threshold decides both halves of the middleware: a buffered body below it is compressed inline, and one at or above it is offloaded, exactly as a streamed chunk is.
Offloading is not free. It costs a handoff per response, and under load every compressing request queues on the same pool. For a small body that handoff costs more than the compression, so compressing inline is faster even though it holds the event loop — the hold is brief, and the pool contention it avoids is not. Past roughly 48 KiB the balance inverts: zlib releases the GIL for long enough that the pool genuinely parallelises, and keeping the loop free wins. The default sits below that crossover so the offload engages before inline becomes the slower choice.
Lower min_stream_chunk_offload to push more work to the pool, or raise it to
keep more inline:
from veloce import GZipMiddleware, Veloce
app = Veloce()
app.add_middleware(GZipMiddleware(min_stream_chunk_offload=8 * 1024))
Rate limiting¶
RateLimitMiddleware limits requests per client. Used with no arguments it runs
a process-local sliding-log limiter — max_requests per window_seconds:
from veloce import RateLimitMiddleware
app.add_middleware(RateLimitMiddleware(max_requests=100, window_seconds=60))
Per-client state is process-local and size-bounded: max_keys (default
100_000) caps how many client keys are tracked at once, so a caller cycling
source addresses cannot grow the limiter's memory without limit. When the cap is
reached, expired entries are dropped first and then the oldest by arrival — never
the client that triggered the eviction, which would otherwise be a way to flush
another client's counter. The same knob applies to the default backend on the
strategy= form below; if you pass your own backend, set its max_keys
instead.
Pass a strategy to choose the algorithm, and a backend to choose where the
per-client state lives:
from veloce import RateLimitMiddleware, TokenBucket
app.add_middleware(RateLimitMiddleware(strategy=TokenBucket(rate=100, per=60, burst=20)))
| Strategy | Behavior |
|---|---|
FixedWindow |
limit per fixed window; cheapest, but allows a burst at the boundary |
SlidingWindow |
limit per rolling window; smooths the boundary burst with two counters |
TokenBucket |
refills rate per per seconds, allowing a burst up to burst (default rate); burst=1 is a strict leaky bucket |
The default InMemoryRateLimitBackend counts per process. For one limit shared
across every worker and host, use RedisRateLimitBackend (see below).
To give a specific route its own limit, decorate its handler with rate_limit.
The limit lives on the handler, so there is no route string to mistype:
from veloce import RateLimitMiddleware, TokenBucket, rate_limit
app.add_middleware(RateLimitMiddleware(strategy=TokenBucket(rate=1000, per=60)))
@app.post("/login")
@rate_limit(TokenBucket(rate=5, per=60)) # stricter, just for this route
async def login(request):
...
Put @rate_limit below the route decorator so the route registers the tagged
handler. A decorated route gets its own per-client counter, independent of the
default budget.
The tag is honoured whichever way the middleware was constructed. It works the
same above a bare max_requests/window_seconds limiter as it does above a
strategy=:
from veloce import FixedWindow, RateLimitMiddleware, rate_limit
app.add_middleware(RateLimitMiddleware(max_requests=1000, window_seconds=60))
@app.post("/login")
@rate_limit(FixedWindow(limit=5, window=60))
async def login(request):
...
A thousand requests a minute for the app, five a minute for /login, counted
separately — spending the login budget leaves the rest of the app untouched, and
vice versa.
Changed in version 0.18.0
@rate_limit above a max_requests=/window_seconds= limiter used to be
collected by nothing and dropped in silence, so the route ran unthrottled.
If you have a decorated route on that form of the middleware, it is being
enforced from this version on — check the limit is the one you want.
For handlers you cannot decorate, the overrides map is the central alternative.
Its key is the route's full path template as matched at runtime (the value of
request.url_rule), so a route on a blueprint with url_prefix="/api" uses
"/api/login", not "/login":
from veloce import RateLimitMiddleware, TokenBucket
app.add_middleware(
RateLimitMiddleware(
strategy=TokenBucket(rate=1000, per=60),
overrides={"/api/login": TokenBucket(rate=5, per=60)},
)
)
An override key that matches no registered route stops startup with
AuditFailed, so a wrong prefix fails before any request is served rather than
silently doing nothing. An explicit
overrides entry wins over a @rate_limit tag on the same route.
Per-route limits resolve against the entry route
Like exclude_middleware, a per-route limit is resolved against the route
matched when the request arrives. A before_request hook that rewrites the
path to a different route does not change which limit applies — rate limiting
runs before those hooks.
Added in version 0.4.0
Selectable strategy/backend, the rate_limit per-route decorator, and the
per-route overrides map on RateLimitMiddleware. The bare
max_requests/window_seconds form is unchanged.
Function middleware¶
For one-off logic, register a function with @app.middleware("http").
It receives the request and a call_next callable that runs the rest of
the stack:
@app.middleware("http")
async def add_timing_header(request, call_next):
response = await call_next(request)
response.headers["X-Powered-By"] = "veloce"
return response
Class-based middleware¶
For reusable middleware, subclass BaseHTTPMiddleware and implement
dispatch:
from veloce import BaseHTTPMiddleware
class RequestIDMiddleware(BaseHTTPMiddleware):
async def dispatch(self, request, call_next):
response = await call_next(request)
response.headers["X-Request-ID"] = new_id()
return response
app.add_http_middleware(RequestIDMiddleware())
Ordering¶
Middleware runs in the order it is added on the way in, and in reverse on the way out — the first one added is the outermost layer.
That ordering covers the middleware registered with add_middleware().
Function middleware registered with @app.middleware("http") is a separate
layer that always wraps the whole add_middleware() chain, so its position in
your source file relative to an add_middleware() call does not change the
result:
from veloce import Middleware, Veloce
app = Veloce()
class Stamp(Middleware):
async def process_response(self, request, response):
response.headers["X-Stamp"] = "1"
return response
@app.middleware("http")
async def outer(request, call_next):
# Runs before Stamp on the way in, and after it on the way out —
# regardless of whether this decorator appears above or below
# `app.add_middleware(Stamp)`.
return await call_next(request)
app.add_middleware(Stamp)
Note
Because the two registration styles form separate layers, a function
middleware always sees the response after every add_middleware()
middleware has processed it. Register both halves of a cooperating pair in
the same style if their relative order matters.
Short-circuiting the request¶
process_request returns None to let the request carry on. Returning a
Response instead short-circuits the
pipeline: the handler never runs, and neither does the process_request of any
middleware added after this one. The returned response still travels back out
through the response phase.
from veloce import JSONResponse, Middleware, Request, Response, Veloce
app = Veloce()
class APIKeyMiddleware(Middleware):
async def process_request(self, request: Request) -> Response | None:
if request.headers.get("X-API-Key") != "expected-key":
return JSONResponse({"detail": "Unauthorized"}, status_code=401)
return None # carry on to the handler
app.add_middleware(APIKeyMiddleware)
This is how every built-in that rejects a request before it reaches your code
works — CSRFMiddleware, RateLimitMiddleware, TrustedHostMiddleware, and
CORSMiddleware's preflight reply all short-circuit from process_request.
Excluding middleware per route¶
A route can opt out of middleware with exclude_middleware. Each entry is
either the middleware class or its resolved name:
from veloce import CSRFMiddleware, Veloce
app = Veloce()
@app.post("/webhooks/stripe", exclude_middleware=[CSRFMiddleware])
async def stripe_webhook():
return {"ok": True}
Prefer the class. It matches by type, so it covers a subclass of that middleware, and it cannot be misspelled — a wrong class is an error where you write it, while a wrong name is an exclusion that silently matches nothing and leaves the middleware running.
A string matches a middleware's resolved name, which defaults to its class name;
pass name= to the middleware when two instances of the same class must be
addressed independently, and exclude them by those names. A string entry matches
that one name exactly and does not reach subclasses.
The opt-out applies to both the request and response phases, so a skipped middleware never runs for that route at all.
Changed in version 0.13
exclude_middleware accepts middleware classes. An entry that is neither a
class nor a string now raises TypeError at registration; it was previously
accepted and matched nothing.
The exclusion set is keyed on the route matched at dispatch entry. The same set
of middleware that runs process_request runs process_response, so setup and
teardown stay balanced. A before_request hook that rewrites the request path
to a different route does not change which middleware run for that request - the
entry route's exclude_middleware is authoritative.
RateLimitMiddleware counts per process by default
The default InMemoryRateLimitBackend keeps its state in one process, so
under uvicorn --workers N the effective limit is roughly N x the
configured one. For a shared cross-worker limit pass a
RedisRateLimitBackend from
veloce.contrib.redis (pip install veloceframework[redis]), which keeps
the state in Redis:
from redis.asyncio import Redis
from veloce import RateLimitMiddleware, TokenBucket
from veloce.contrib.redis import RedisRateLimitBackend
client = Redis.from_url("redis://localhost:6379/0")
app.add_middleware(
RateLimitMiddleware(
strategy=TokenBucket(rate=100, per=60),
backend=RedisRateLimitBackend(client),
)
)
app.add_middleware(CSRFMiddleware(secret="..."))
app.add_middleware(RateLimitMiddleware(max_requests=100, window_seconds=60))
# Inbound webhooks can't carry a CSRF token, and the health probe should
# never be rate limited.
@app.post("/webhooks/stripe", exclude_middleware=["CSRFMiddleware"])
async def stripe_webhook(request):
...
@app.get("/health", exclude_middleware=["RateLimitMiddleware"])
async def health():
return {"status": "ok"}
This works on @app.route/@app.get/@app.post/… and the imperative
add_api_route, and on Blueprint and Router routes. Routes that declare no
exclusions run every registered middleware and pay no extra per-request cost.
Inspecting the pipeline¶
app.middlewares is the registered Middleware instances as a tuple, in the
order they will run - registration order, or priority order when priorities are
set. It answers "is this installed?" without reaching into the app:
from veloce import CORSMiddleware, Veloce
app = Veloce()
if not any(isinstance(m, CORSMiddleware) for m in app.middlewares):
app.add_middleware(CORSMiddleware, allow_origins=["https://example.com"])
It is a snapshot, and a tuple: adding middleware must go through
add_middleware(), which is what maintains the ordering and the compiled
pipeline. Standard ASGI middleware classes do not appear here - they wrap the
application when the ASGI stack is assembled rather than running in this
pipeline.
See also¶
- CORS — the full
CORSMiddlewareparameter reference. - Sessions —
SessionMiddlewareandServerSessionMiddleware. - Configuration —
SECRET_KEYand the app-level defaults. - Deployment
- Routing
- Dependency injection
- The API reference lists every middleware class.