Skip to content

AP4 — Pydantic validation overhead ​

What goes wrong ​

Pydantic v2 is fast. The problem is redundant work. Middleware and serializer layers often grow a pattern like this: build a dict by hand from the ORM object, validate it, dump it, validate the dump again, and dump again. Each validation re-runs every field validator, and each round-trip allocates new objects.

The overhead is pure CPU, and it runs on the event-loop thread. It is easy to miss in a single request, and it adds up across thousands of rows. Most people recognise the bad version from their own codebase.

The bad code ​

bad_transform is in orders-api-demo/app/schemas_ap4.py. It builds a dict, validates it, dumps it, and validates it again. The second validation is the redundant one.

python
def bad_transform(order) -> dict:
    raw = {
        "id": str(order.id),
        "customer_id": str(order.customer_id),
        "status": order.status,
        "items": [
            {
                "id": str(item.id),
                "sku": item.sku,
                "quantity": item.quantity,
                "unit_price": float(item.unit_price),
            }
            for item in order.items
        ],
    }
    validated = BadOrder.model_validate(raw)
    dumped = validated.model_dump()  # round-trip #1
    revalidated = BadOrder.model_validate(dumped)  # round-trip #2 — redundant
    return revalidated.model_dump()

BadOrder and BadOrderItem have chained field_validators, for example the SKU regex and a price normalisation. Those run on every validation, so the round-trip runs them twice.

The good code ​

good_transform validates the ORM object directly, once. No manual dict, no round-trip.

python
class GoodOrder(BaseModel):
    model_config = ConfigDict(from_attributes=True)

    id: uuid.UUID
    customer_id: uuid.UUID
    status: str
    items: list[GoodOrderItem]


def good_transform(order) -> dict:
    validated = GoodOrder.model_validate(order)
    return validated.model_dump(mode="json")

For data that never crosses a process boundary, the file also includes internal_summary, which returns a TypedDict and skips Pydantic entirely.

Try it ​

bash
curl "localhost:8000/demo/ap4/bad?limit=500"
curl "localhost:8000/demo/ap4/good?limit=500"
bash
PYTHONPATH=. uv run python scripts/pydantic_cprofile_demo.py

Compare elapsed_seconds. Only the transform is timed, not the database fetch. Expected, local SQLite (DEMO_GUIDE): bad is roughly 2 to 2.7× slower than good, for example 5.8 ms against 2.9 ms at 500 orders.

The cProfile script prints the function-call counts. Expected, local SQLite: about 290k calls for bad and 111k for good.

Live, the time difference is only milliseconds, so it will not read on a projector. Show the captured numbers instead, and use the endpoints if you want a live moment.

Measured evidence ​

0.056 s → 0.030 s
cProfile total time
Postgres capture, 200 orders × 20 repeats
204,481 → 116,463
function calls
Postgres capture
~1.9×
slower
time, bad ÷ good
~290k vs 111k
function calls, SQLite local
working-tree capture, not committed
Bad
204481 calls
Good
116463 calls
1.76× bad ÷ good · function calls, Postgres capture (committed)

Captured header lines (benchmarks/ap4-pydantic/cprofile-bad.txt, cprofile-good.txt):

text
=== BAD (nested validators + round-trip): 200 orders x 20 repeats ===
204481 function calls in 0.056 seconds

=== GOOD (from_attributes, single pass): 200 orders x 20 repeats ===
116463 function calls (116319 primitive calls) in 0.030 seconds

The automated test tests/test_ap4_pydantic.py asserts identical output counts and that bad is more than 1.3× slower than good.

Why call counts matter more than wall time

Wall-clock timings move with the machine. Function-call counts are stable across machines, so they make a better comparison.

How to detect it in your own service ​

ToolCommandWhat to look for
cProfilepython -m cProfile -o out.prof script.pymodel_validate and model_dump called twice per object
Code reviewSearch for model_dump() followed by model_validate(A round-trip with no change in between
BenchmarkCompare call counts over a realistic batch sizeA large call count for a simple mapping

Checklist ​

Talking points ​

  • Pydantic v2 is fast, but a redundant round-trip times thousands of rows is not free, and it is CPU on the event-loop thread.
  • Measure call counts, not just wall time. They are stable across machines.
  • The "skip Pydantic for internal data" point is the one people forget. Mention it explicitly.
  • Do not demo this live. Show the captured cProfile numbers.

Released under the MIT License. Speaker: Satyam Soni, PyCon Hong Kong 2026.