Performance checklist
In one sentence
Five things to check before you trust a FastAPI service under real concurrency. Each one maps to a demo with captured before-and-after numbers.
Tick items as you go. Progress is saved in this browser's local storage only. Nothing is sent anywhere, and the page works without it.
1. Event loop hygiene
Demo: AP1. Three concurrent requests: 3.05 s blocking against 1.04 s async.
2. Dependency lifecycle
Demo: AP2. 17.5 ms against 1.7 ms mean, about 10×, from rebuilding an engine per request.
3. ORM loading strategy
Demo: AP3. Plain attribute access crashes with MissingGreenlet. awaitable_attrs works but costs 6 queries against 2 for five orders.
4. Pydantic usage
Demo: AP4. About 1.9× slower and about 1.8× more function calls, from a redundant validation round-trip.
5. Connection pool sizing
Demo: AP5. Same load, same code: 2.1% failures at 47.8 req/s with the undersized pool, against 0% failures at 179.3 req/s with the tuned pool.
Profiling toolkit
| Tool | Use it for | Command |
|---|---|---|
py-spy | Zero-instrumentation sampling profiler, flame graphs | py-spy record --pid <PID> -o out.svg |
asyncio debug mode | Catching blocking calls inside async def | PYTHONASYNCIODEBUG=1 python app.py |
cProfile and pstats | Function-level CPU cost, for example Pydantic overhead | python -m cProfile -o out.prof script.py |
SQLAlchemy echo=True | Seeing every SQL statement and the query counts | create_async_engine(url, echo=True) |
locust | Realistic concurrent load, pool and timeout behaviour | locust --headless -u 50 -r 10 -t 30s |
More detail is on the profiling toolkit page.