Back to Blog
fastapi
api
python
scalability
security

Async, Secure, and Scalable: Why FastAPI Is My Go-To for Modern APIs

A deep dive into FastAPI's strengths for building robust APIs, with practical examples and tips for scaling your next project.

Niklas L.
23 min read

⚡ Async, Secure, and Scalable: Why FastAPI Is My Go-To for Modern APIs

Let me paint you a picture. It's 3 AM, your API is getting hammered with 50,000 concurrent requests because someone posted about your app on Hacker News. Your PostgreSQL connection pool is maxed out, Redis is screaming, and you're watching your server metrics climb into the red zone. This exact scenario happened to me last month, and you know what? My FastAPI backend didn't even flinch. CPU usage peaked at 40%, memory stayed stable, and every single request got served in under 200ms.

This isn't some theoretical benchmark bullshit—this is real production traffic on a $20/month VPS. And it's why I've been building everything with FastAPI for the past three years.

The Async Revolution Nobody's Talking About Properly

Everyone mentions that FastAPI is "async by default," but most tutorials show you some trivial async def endpoint that returns "Hello World" and call it a day. Let's talk about what async actually means when you're building something real.

Traditional Flask or Django apps handle requests like a government office—one at a time, making everyone else wait in line. When your endpoint hits a database or calls an external API, that worker just sits there, twiddling its thumbs, while other requests pile up. It's like having a chef who stands and watches the oven instead of prepping the next dish.

FastAPI flips this model on its head. When an async endpoint hits an I/O operation, it immediately moves on to handle another request. Here's what this looks like in practice:

import asyncio
import httpx
from fastapi import FastAPI
from typing import List

app = FastAPI()

async def fetch_user_data(user_id: int):
    async with httpx.AsyncClient() as client:
        # This doesn't block - other requests get processed while waiting
        response = await client.get(f"https://api.example.com/users/{user_id}")
        return response.json()

@app.get("/dashboard/{user_id}")
async def get_dashboard(user_id: int):
    # Fire off multiple async operations simultaneously
    tasks = [
        fetch_user_data(user_id),
        get_recent_activity(user_id),
        calculate_analytics(user_id),
        fetch_notifications(user_id)
    ]

    # All four operations happen in parallel
    results = await asyncio.gather(*tasks)

    return {
        "user": results[0],
        "activity": results[1],
        "analytics": results[2],
        "notifications": results[3]
    }

In a synchronous framework, those four operations would happen sequentially. If each takes 200ms, you're looking at 800ms total. With async, they all happen simultaneously, so you're back down to 200ms. That's a 4x performance improvement with zero additional infrastructure.

The WebSocket Game Changer

But here's where it gets really interesting. I was building a collaborative document editor last year—think Google Docs but for technical specifications. Users needed to see each other's changes in real-time, cursor positions, who's typing where, the whole nine yards.

WebSockets in FastAPI are stupid simple but incredibly powerful:

from fastapi import WebSocket, WebSocketDisconnect
from typing import Dict, Set
import json

class ConnectionManager:
    def __init__(self):
        self.active_connections: Dict[str, Set[WebSocket]] = {}

    async def connect(self, websocket: WebSocket, document_id: str):
        await websocket.accept()
        if document_id not in self.active_connections:
            self.active_connections[document_id] = set()
        self.active_connections[document_id].add(websocket)

    async def broadcast(self, document_id: str, message: dict, sender: WebSocket):
        if document_id in self.active_connections:
            # Send to everyone except the sender
            connections = self.active_connections[document_id] - {sender}
            for connection in connections:
                try:
                    await connection.send_json(message)
                except:
                    # Connection is dead, clean it up
                    self.active_connections[document_id].discard(connection)

manager = ConnectionManager()

@app.websocket("/ws/{document_id}")
async def websocket_endpoint(websocket: WebSocket, document_id: str):
    await manager.connect(websocket, document_id)
    try:
        while True:
            data = await websocket.receive_json()

            # Validate and process the change
            if data["type"] == "cursor_move":
                await manager.broadcast(document_id, {
                    "type": "cursor_update",
                    "user_id": data["user_id"],
                    "position": data["position"]
                }, websocket)
            elif data["type"] == "text_change":
                # Here's where you'd update your database
                await update_document(document_id, data["changes"])
                await manager.broadcast(document_id, data, websocket)
    except WebSocketDisconnect:
        manager.disconnect(websocket, document_id)

This handles hundreds of concurrent users editing the same document. The beautiful part? It's running on a single FastAPI instance. No separate WebSocket server, no complex message queue setup, just clean async Python.

Security: Beyond the "Just Use JWT" Nonsense

Every FastAPI tutorial tells you to slap JWT on your endpoints and call it secure. That's like putting a deadbolt on your door while leaving the windows open. Real security is layered, and FastAPI gives you the tools to build it properly.

The Authentication Stack That Actually Works

Here's the authentication system I've been evolving over dozens of projects:

from fastapi import Depends, HTTPException, Security
from fastapi.security import HTTPBearer, HTTPAuthorizationCredentials
from passlib.context import CryptContext
from datetime import datetime, timedelta
import jwt
import redis.asyncio as redis
from typing import Optional

pwd_context = CryptContext(schemes=["argon2"], deprecated="auto")
security = HTTPBearer()

class AuthenticationService:
    def __init__(self):
        self.redis_client = redis.Redis(decode_responses=True)
        self.ACCESS_TOKEN_EXPIRE = timedelta(minutes=15)
        self.REFRESH_TOKEN_EXPIRE = timedelta(days=30)

    async def create_tokens(self, user_id: str):
        # Short-lived access token
        access_payload = {
            "sub": user_id,
            "exp": datetime.utcnow() + self.ACCESS_TOKEN_EXPIRE,
            "type": "access",
            "jti": generate_token_id()  # Unique token ID for revocation
        }

        # Long-lived refresh token
        refresh_payload = {
            "sub": user_id,
            "exp": datetime.utcnow() + self.REFRESH_TOKEN_EXPIRE,
            "type": "refresh",
            "jti": generate_token_id()
        }

        access_token = jwt.encode(access_payload, SECRET_KEY, algorithm="HS256")
        refresh_token = jwt.encode(refresh_payload, REFRESH_SECRET, algorithm="HS256")

        # Store refresh token in Redis with expiration
        await self.redis_client.setex(
            f"refresh_token:{user_id}:{refresh_payload['jti']}",
            self.REFRESH_TOKEN_EXPIRE.total_seconds(),
            refresh_token
        )

        return access_token, refresh_token

    async def verify_token(self, credentials: HTTPAuthorizationCredentials = Security(security)):
        token = credentials.credentials

        try:
            payload = jwt.decode(token, SECRET_KEY, algorithms=["HS256"])

            # Check if token is revoked
            if await self.redis_client.exists(f"revoked:{payload['jti']}"):
                raise HTTPException(status_code=401, detail="Token has been revoked")

            return payload["sub"]
        except jwt.ExpiredSignatureError:
            raise HTTPException(status_code=401, detail="Token expired")
        except jwt.JWTError:
            raise HTTPException(status_code=401, detail="Invalid token")

But here's the kicker—authentication is just the beginning. You need rate limiting, request validation, and proper CORS handling.

Rate Limiting That Scales

Most rate limiting tutorials show you some in-memory counter that resets when your server restarts. Here's production-grade rate limiting using Redis:

from fastapi import Request
import hashlib

class RateLimiter:
    def __init__(self, requests_per_minute: int = 60):
        self.requests_per_minute = requests_per_minute
        self.redis_client = redis.Redis(decode_responses=True)

    async def check_rate_limit(self, request: Request, user_id: Optional[str] = None):
        # Use IP for anonymous users, user_id for authenticated
        if user_id:
            key = f"rate_limit:user:{user_id}"
            limit = self.requests_per_minute * 2  # Authenticated users get more
        else:
            # Hash the IP for privacy
            ip_hash = hashlib.sha256(request.client.host.encode()).hexdigest()[:16]
            key = f"rate_limit:ip:{ip_hash}"
            limit = self.requests_per_minute

        pipe = self.redis_client.pipeline()
        pipe.incr(key)
        pipe.expire(key, 60)

        current_requests, _ = await pipe.execute()

        if current_requests > limit:
            # Calculate time until reset
            ttl = await self.redis_client.ttl(key)
            raise HTTPException(
                status_code=429,
                detail=f"Rate limit exceeded. Try again in {ttl} seconds",
                headers={"Retry-After": str(ttl)}
            )

        return current_requests

# Use it as a dependency
@app.get("/api/sensitive-endpoint")
async def sensitive_endpoint(
    rate_limit: int = Depends(rate_limiter.check_rate_limit),
    user_id: str = Depends(auth_service.verify_token)
):
    return {"requests_used": rate_limit}

Input Validation That Catches Everything

Pydantic isn't just for type hints—it's your first line of defense against malicious input. But you need to be paranoid about it:

from pydantic import BaseModel, validator, Field
from typing import Optional
import re
import bleach

class UserInput(BaseModel):
    username: str = Field(..., min_length=3, max_length=20)
    email: str
    bio: Optional[str] = Field(None, max_length=500)
    age: int = Field(..., ge=13, le=120)

    @validator('username')
    def username_alphanumeric(cls, v):
        if not re.match(r'^[a-zA-Z0-9_]+$', v):
            raise ValueError('Username must be alphanumeric with underscores only')

        # Check for sneaky Unicode lookalikes
        if v != v.encode('ascii', 'ignore').decode('ascii'):
            raise ValueError('Username contains invalid characters')

        # Prevent username squatting
        reserved = ['admin', 'root', 'api', 'www', 'mail']
        if v.lower() in reserved:
            raise ValueError('Username is reserved')

        return v

    @validator('email')
    def email_valid(cls, v):
        # Don't just check format, check for common typos
        if '@' not in v or '.' not in v.split('@')[1]:
            raise ValueError('Invalid email format')

        # Check for disposable email domains
        domain = v.split('@')[1].lower()
        disposable_domains = load_disposable_domains()  # Load from file/DB
        if domain in disposable_domains:
            raise ValueError('Disposable email addresses not allowed')

        return v.lower()

    @validator('bio')
    def sanitize_bio(cls, v):
        if v:
            # Strip HTML but keep basic formatting
            cleaned = bleach.clean(v, tags=['b', 'i', 'u', 'br'], strip=True)
            # Remove excessive whitespace
            cleaned = ' '.join(cleaned.split())
            return cleaned
        return v

Scaling: From Side Project to Series A

Here's something nobody tells you about scaling—it's not about the big moments. It's about the hundred tiny decisions you make along the way. FastAPI makes most of those decisions correctly by default, but you still need to know what you're doing.

Database Connections: The Silent Killer

The number one reason APIs fall over under load? Database connection exhaustion. Here's how to handle it properly:

from sqlalchemy.ext.asyncio import create_async_engine, AsyncSession
from sqlalchemy.orm import sessionmaker
from contextlib import asynccontextmanager
import asyncpg

# Create the async engine with proper pooling
engine = create_async_engine(
    "postgresql+asyncpg://user:pass@localhost/db",
    pool_size=20,  # Number of connections to maintain
    max_overflow=10,  # Maximum overflow connections
    pool_timeout=30,  # Timeout for getting connection from pool
    pool_recycle=1800,  # Recycle connections after 30 minutes
    pool_pre_ping=True,  # Test connections before using
)

AsyncSessionLocal = sessionmaker(
    engine, class_=AsyncSession, expire_on_commit=False
)

# Connection dependency with automatic cleanup
async def get_db():
    async with AsyncSessionLocal() as session:
        try:
            yield session
            await session.commit()
        except Exception:
            await session.rollback()
            raise
        finally:
            await session.close()

# But here's the advanced stuff - query optimization
from sqlalchemy import select, func
from sqlalchemy.orm import selectinload, joinedload

@app.get("/users/{user_id}/full-profile")
async def get_full_profile(user_id: int, db: AsyncSession = Depends(get_db)):
    # Bad: N+1 query problem
    # user = await db.get(User, user_id)
    # posts = await user.posts  # Another query
    # for post in posts:
    #     comments = await post.comments  # N more queries!

    # Good: Everything in one query
    result = await db.execute(
        select(User)
        .options(
            selectinload(User.posts).selectinload(Post.comments),
            joinedload(User.profile)
        )
        .where(User.id == user_id)
    )

    user = result.scalar_one_or_none()

    if not user:
        raise HTTPException(status_code=404)

    return user

Caching: The Performance Multiplier

Redis isn't just for rate limiting. Used properly, it can reduce your database load by 90%:

import pickle
from typing import Optional, Any
import hashlib

class CacheService:
    def __init__(self):
        self.redis = redis.Redis(decode_responses=False)  # Binary mode for pickle

    def make_cache_key(self, prefix: str, **kwargs) -> str:
        """Generate consistent cache keys from parameters"""
        # Sort kwargs for consistent keys
        sorted_params = sorted(kwargs.items())
        param_string = ":".join([f"{k}={v}" for k, v in sorted_params])

        # Hash long keys to avoid Redis key length limits
        if len(param_string) > 200:
            param_string = hashlib.md5(param_string.encode()).hexdigest()

        return f"{prefix}:{param_string}"

    async def get_or_set(
        self,
        key: str,
        getter_func,
        ttl: int = 300,
        skip_cache: bool = False
    ) -> Any:
        """Get from cache or execute function and cache result"""
        if not skip_cache:
            cached = await self.redis.get(key)
            if cached:
                return pickle.loads(cached)

        # Execute the function
        result = await getter_func()

        # Cache the result
        await self.redis.setex(key, ttl, pickle.dumps(result))

        return result

    async def invalidate_pattern(self, pattern: str):
        """Invalidate all keys matching a pattern"""
        cursor = 0
        while True:
            cursor, keys = await self.redis.scan(cursor, match=pattern)
            if keys:
                await self.redis.delete(*keys)
            if cursor == 0:
                break

cache = CacheService()

# Usage in endpoints
@app.get("/analytics/{user_id}")
async def get_analytics(
    user_id: int,
    force_refresh: bool = False,
    db: AsyncSession = Depends(get_db)
):
    cache_key = cache.make_cache_key("analytics", user_id=user_id)

    async def calculate_analytics():
        # Expensive calculation
        return await db.execute(
            select(
                func.count(Post.id).label("total_posts"),
                func.sum(Post.views).label("total_views"),
                func.avg(Post.rating).label("avg_rating")
            )
            .where(Post.user_id == user_id)
        ).first()

    return await cache.get_or_set(
        cache_key,
        calculate_analytics,
        ttl=3600,  # Cache for 1 hour
        skip_cache=force_refresh
    )

Background Tasks: Don't Make Users Wait

One of FastAPI's most underutilized features is background tasks. Stop making users wait for emails to send or logs to write:

from fastapi import BackgroundTasks
import asyncio
from typing import List

class EmailService:
    def __init__(self):
        self.queue = asyncio.Queue()
        self.batch_size = 10
        self.batch_timeout = 5.0

    async def send_batch(self, emails: List[dict]):
        """Send multiple emails efficiently"""
        # Use your email service's batch API
        async with httpx.AsyncClient() as client:
            await client.post(
                "https://api.sendgrid.com/v3/mail/send",
                json={"personalizations": emails},
                headers={"Authorization": f"Bearer {SENDGRID_KEY}"}
            )

@app.post("/signup")
async def signup(
    user_data: UserInput,
    background_tasks: BackgroundTasks,
    db: AsyncSession = Depends(get_db)
):
    # Create user immediately
    user = User(**user_data.dict())
    db.add(user)
    await db.commit()

    # Queue email for background processing
    background_tasks.add_task(
        email_service.send_welcome_email,
        user.email,
        user.username
    )

    # Queue analytics event
    background_tasks.add_task(
        analytics.track,
        "user_signup",
        {"user_id": user.id, "source": "api"}
    )

    # User gets response immediately
    return {"id": user.id, "status": "created"}

Production Deployment: The Real World

Everyone can get a FastAPI app running locally. Here's how to deploy it so it doesn't fall over when TechCrunch writes about you.

Docker Configuration That Actually Works

# Multi-stage build for smaller images
FROM python:3.11-slim as builder

WORKDIR /app

# Install build dependencies
RUN apt-get update && apt-get install -y \
    gcc \
    g++ \
    && rm -rf /var/lib/apt/lists/*

# Copy requirements first for better caching
COPY requirements.txt .
RUN pip install --user --no-cache-dir -r requirements.txt

# Production image
FROM python:3.11-slim

WORKDIR /app

# Create non-root user
RUN useradd -m -u 1000 appuser && chown -R appuser:appuser /app

# Copy installed packages from builder
COPY --from=builder --chown=appuser:appuser /root/.local /home/appuser/.local

# Copy application code
COPY --chown=appuser:appuser . .

USER appuser

# Update PATH
ENV PATH=/home/appuser/.local/bin:$PATH

# Health check
HEALTHCHECK --interval=30s --timeout=3s --start-period=5s --retries=3 \
    CMD python -c "import httpx; httpx.get('http://localhost:8000/health')"

# Run with optimal settings
CMD ["uvicorn", "main:app", \
     "--host", "0.0.0.0", \
     "--port", "8000", \
     "--workers", "4", \
     "--loop", "uvloop", \
     "--access-log", \
     "--log-config", "logging.yaml"]

Nginx Configuration for Production

upstream fastapi_backend {
    least_conn;  # Better than round-robin for varying request times
    server app1:8000 max_fails=3 fail_timeout=30s;
    server app2:8000 max_fails=3 fail_timeout=30s;
    keepalive 32;  # Keep connections alive
}

server {
    listen 80;
    server_name api.yourdomain.com;

    # Redirect to HTTPS
    return 301 https://$server_name$request_uri;
}

server {
    listen 443 ssl http2;
    server_name api.yourdomain.com;

    # SSL configuration
    ssl_certificate /etc/nginx/ssl/cert.pem;
    ssl_certificate_key /etc/nginx/ssl/key.pem;
    ssl_protocols TLSv1.2 TLSv1.3;
    ssl_ciphers HIGH:!aNULL:!MD5;

    # Security headers
    add_header X-Content-Type-Options nosniff;
    add_header X-Frame-Options DENY;
    add_header X-XSS-Protection "1; mode=block";
    add_header Strict-Transport-Security "max-age=31536000; includeSubDomains" always;

    # Rate limiting zones
    limit_req_zone $binary_remote_addr zone=general:10m rate=10r/s;
    limit_req_zone $binary_remote_addr zone=auth:10m rate=3r/s;

    # General API endpoints
    location /api {
        limit_req zone=general burst=20 nodelay;

        proxy_pass http://fastapi_backend;
        proxy_http_version 1.1;
        proxy_set_header Connection "";
        proxy_set_header Host $host;
        proxy_set_header X-Real-IP $remote_addr;
        proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
        proxy_set_header X-Forwarded-Proto $scheme;

        # Timeouts
        proxy_connect_timeout 5s;
        proxy_send_timeout 60s;
        proxy_read_timeout 60s;

        # Buffering
        proxy_buffering on;
        proxy_buffer_size 4k;
        proxy_buffers 8 4k;
    }

    # Auth endpoints with stricter limits
    location /api/auth {
        limit_req zone=auth burst=5 nodelay;
        proxy_pass http://fastapi_backend;
        # ... same proxy settings
    }

    # WebSocket support
    location /ws {
        proxy_pass http://fastapi_backend;
        proxy_http_version 1.1;
        proxy_set_header Upgrade $http_upgrade;
        proxy_set_header Connection "upgrade";
        proxy_read_timeout 86400;
    }
}

Monitoring: You Can't Fix What You Can't See

The difference between hobby projects and production systems? Monitoring. Here's my battle-tested setup:

from prometheus_client import Counter, Histogram, Gauge, generate_latest
from functools import wraps
import time
import traceback

# Metrics
request_count = Counter('api_requests_total', 'Total requests', ['method', 'endpoint', 'status'])
request_duration = Histogram('api_request_duration_seconds', 'Request duration', ['method', 'endpoint'])
active_connections = Gauge('api_active_connections', 'Active connections')
db_pool_size = Gauge('db_pool_size', 'Database pool size')

# Middleware for automatic metrics
@app.middleware("http")
async def track_metrics(request: Request, call_next):
    start_time = time.time()
    active_connections.inc()

    try:
        response = await call_next(request)
        duration = time.time() - start_time

        # Track metrics
        request_count.labels(
            method=request.method,
            endpoint=request.url.path,
            status=response.status_code
        ).inc()

        request_duration.labels(
            method=request.method,
            endpoint=request.url.path
        ).observe(duration)

        return response
    except Exception as e:
        # Track errors
        request_count.labels(
            method=request.method,
            endpoint=request.url.path,
            status=500
        ).inc()

        # Log to Sentry
        sentry_sdk.capture_exception(e)

        raise
    finally:
        active_connections.dec()

# Metrics endpoint for Prometheus
@app.get("/metrics")
async def metrics():
    # Update dynamic metrics
    db_pool_size.set(engine.pool.size())

    return Response(
        content=generate_latest(),
        media_type="text/plain"
    )

# Custom decorator for tracking specific operations
def track_operation(operation_name: str):
    def decorator(func):
        @wraps(func)
        async def wrapper(*args, **kwargs):
            with operation_duration.labels(operation=operation_name).time():
                try:
                    result = await func(*args, **kwargs)
                    operation_success.labels(operation=operation_name).inc()
                    return result
                except Exception as e:
                    operation_errors.labels(
                        operation=operation_name,
                        error_type=type(e).__name__
                    ).inc()
                    raise
        return wrapper
    return decorator

The Template That Changes Everything

Look, I've built this stack probably 50 times now. Same authentication, same database setup, same monitoring, same Docker configuration. That's why I finally put together FastLaunchAPI.dev—it's everything I just showed you, pre-configured and ready to deploy.

It's not some generic boilerplate. It's the exact setup I use in production, refined over three years and dozens of projects. JWT with refresh tokens, Redis caching, async SQLAlchemy, Prometheus metrics, the works. You clone it, change a few environment variables, and you're running the same stack that handles millions of requests on my production apps.

Performance Tricks Nobody Talks About

Here are some discoveries that took me way too long to figure out:

1. Connection Pool Warmup

Cold starts kill performance. Pre-warm your connection pools:

@app.on_event("startup")
async def warmup():
    # Pre-create database connections
    tasks = []
    for _ in range(engine.pool.size()):
        tasks.append(verify_database_connection())

    await asyncio.gather(*tasks, return_exceptions=True)

    # Pre-warm Redis connections
    for _ in range(10):
        await redis_client.ping()

    # Pre-compile regex patterns
    global EMAIL_PATTERN, URL_PATTERN
    EMAIL_PATTERN = re.compile(r'^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}$')
    URL_PATTERN = re.compile(r'https?://[^\s]+')

2. Smart Query Batching

Instead of N database queries, batch them intelligently:

from dataclasses import dataclass
from typing import List, Dict
import asyncio

@dataclass
class BatchRequest:
    key: str
    resolver: callable
    result: asyncio.Future

class BatchProcessor:
    def __init__(self, batch_size: int = 100, timeout: float = 0.01):
        self.batch_size = batch_size
        self.timeout = timeout
        self.pending: List[BatchRequest] = []
        self.processing = False

    async def get(self, key: str, resolver: callable):
        future = asyncio.Future()
        request = BatchRequest(key, resolver, future)
        self.pending.append(request)

        if len(self.pending) >= self.batch_size:
            await self._process_batch()
        elif not self.processing:
            asyncio.create_task(self._process_after_timeout())

        return await future

    async def _process_after_timeout(self):
        self.processing = True
        await asyncio.sleep(self.timeout)
        await self._process_batch()
        self.processing = False

    async def _process_batch(self):
        if not self.pending:
            return

        batch = self.pending[:self.batch_size]
        self.pending = self.pending[self.batch_size:]

        # Group by resolver function
        grouped: Dict[callable, List[BatchRequest]] = {}
        for request in batch:
            if request.resolver not in grouped:
                grouped[request.resolver] = []
            grouped[request.resolver].append(request)

        # Execute each resolver once with all its keys
        for resolver, requests in grouped.items():
            keys = [r.key for r in requests]
            try:
                results = await resolver(keys)
                for request in requests:
                    request.result.set_result(results.get(request.key))
            except Exception as e:
                for request in requests:
                    request.result.set_exception(e)

3. Response Streaming for Large Data

Don't load everything into memory:

from fastapi.responses import StreamingResponse
import csv
from io import StringIO

@app.get("/export/users")
async def export_users(db: AsyncSession = Depends(get_db)):
    async def generate():
        # Write CSV header
        output = StringIO()
        writer = csv.writer(output)
        writer.writerow(['id', 'username', 'email', 'created_at'])
        yield output.getvalue()

        # Stream results in chunks
        offset = 0
        chunk_size = 1000

        while True:
            result = await db.execute(
                select(User)
                .offset(offset)
                .limit(chunk_size)
            )
            users = result.scalars().all()

            if not users:
                break

            output = StringIO()
            writer = csv.writer(output)
            for user in users:
                writer.writerow([user.id, user.username, user.email, user.created_at])

            yield output.getvalue()
            offset += chunk_size

            # Let other requests process
            await asyncio.sleep(0)

    return StreamingResponse(
        generate(),
        media_type='text/csv',
        headers={'Content-Disposition': 'attachment; filename="users.csv"'}
    )

## The Mistakes That Nearly Killed My Startup

Let me tell you about the time our API crashed during a live demo with investors. 50,000 requests hit us in 30 seconds because someone accidentally left a while loop in the frontend. The server ran out of memory, PostgreSQL connections maxed out, and the whole thing just died. That's when I learned these lessons the hard way.

### Memory Leaks in Async Code

Python's garbage collector doesn't always play nice with async code. Here's a leak that took me three days to find:

```python
# BAD: This leaks memory like crazy
class WebSocketManager:
    def __init__(self):
        self.connections = {}  # This grows forever!

    async def connect(self, user_id: str, websocket: WebSocket):
        self.connections[user_id] = websocket
        # If user reconnects, old connection stays in memory

# GOOD: Proper cleanup
class WebSocketManager:
    def __init__(self):
        self.connections = {}
        self.connection_tasks = {}

    async def connect(self, user_id: str, websocket: WebSocket):
        # Clean up existing connection
        if user_id in self.connections:
            await self.disconnect(user_id)

        self.connections[user_id] = websocket

        # Track the connection task
        task = asyncio.current_task()
        self.connection_tasks[user_id] = task

        # Set up automatic cleanup
        try:
            await websocket.accept()
            while True:
                data = await websocket.receive_text()
                # Process data
        except WebSocketDisconnect:
            pass
        finally:
            await self.disconnect(user_id)

    async def disconnect(self, user_id: str):
        if user_id in self.connections:
            del self.connections[user_id]
        if user_id in self.connection_tasks:
            del self.connection_tasks[user_id]

The Thundering Herd Problem

When your cache expires, suddenly every request tries to rebuild it simultaneously. Here's the fix nobody teaches:

import asyncio
from typing import Optional, Dict, Any

class ThunderingHerdLock:
    def __init__(self):
        self.locks: Dict[str, asyncio.Lock] = {}
        self.results: Dict[str, Any] = {}

    async def get_or_compute(self, key: str, compute_func):
        # Create a lock for this specific key
        if key not in self.locks:
            self.locks[key] = asyncio.Lock()

        # Try to get result without lock first (fast path)
        if key in self.results:
            return self.results[key]

        # Acquire lock for this key
        async with self.locks[key]:
            # Double-check after acquiring lock
            if key in self.results:
                return self.results[key]

            # Only one coroutine computes the value
            result = await compute_func()
            self.results[key] = result

            # Clean up after some time to prevent memory growth
            asyncio.create_task(self._cleanup_after(key, ttl=300))

            return result

    async def _cleanup_after(self, key: str, ttl: int):
        await asyncio.sleep(ttl)
        self.results.pop(key, None)
        self.locks.pop(key, None)

# Usage
herd_lock = ThunderingHerdLock()

@app.get("/expensive-computation/{item_id}")
async def expensive_endpoint(item_id: int):
    async def compute():
        # Simulate expensive operation
        await asyncio.sleep(5)
        return {"result": item_id * 42}

    return await herd_lock.get_or_compute(f"compute:{item_id}", compute)

Database Transaction Gotchas

This one cost us $3,000 in AWS bills before we figured it out:

# BAD: Long-running transaction blocks other queries
@app.post("/process-batch")
async def process_batch_bad(items: List[Item], db: AsyncSession = Depends(get_db)):
    async with db.begin():  # Transaction starts here
        for item in items:  # Could be thousands
            await process_item(item)  # Each takes 100ms
            await external_api_call(item)  # Network call inside transaction!
        # Transaction held for potentially minutes

# GOOD: Minimal transaction scope
@app.post("/process-batch")
async def process_batch_good(items: List[Item], db: AsyncSession = Depends(get_db)):
    # Prepare everything outside transaction
    processed_items = []
    for item in items:
        result = await external_api_call(item)
        processed_items.append(prepare_for_db(item, result))

    # Quick transaction just for writes
    async with db.begin():
        for item in processed_items:
            db.add(item)
        # Transaction is milliseconds, not minutes

Advanced Patterns That Scale to Millions

After three years of FastAPI in production, these are the patterns that separate toys from real systems.

Circuit Breakers for External Services

When external APIs fail, don't let them take you down with them:

from datetime import datetime, timedelta
from enum import Enum
from typing import Optional
import asyncio

class CircuitState(Enum):
    CLOSED = "closed"  # Normal operation
    OPEN = "open"      # Failing, reject requests
    HALF_OPEN = "half_open"  # Testing if service recovered

class CircuitBreaker:
    def __init__(
        self,
        failure_threshold: int = 5,
        recovery_timeout: int = 60,
        expected_exception: type = Exception
    ):
        self.failure_threshold = failure_threshold
        self.recovery_timeout = recovery_timeout
        self.expected_exception = expected_exception
        self.failure_count = 0
        self.last_failure_time: Optional[datetime] = None
        self.state = CircuitState.CLOSED

    async def call(self, func, *args, **kwargs):
        if self.state == CircuitState.OPEN:
            if datetime.now() - self.last_failure_time > timedelta(seconds=self.recovery_timeout):
                self.state = CircuitState.HALF_OPEN
            else:
                raise Exception("Circuit breaker is OPEN")

        try:
            result = await func(*args, **kwargs)
            self._on_success()
            return result
        except self.expected_exception as e:
            self._on_failure()
            raise

    def _on_success(self):
        self.failure_count = 0
        self.state = CircuitState.CLOSED

    def _on_failure(self):
        self.failure_count += 1
        self.last_failure_time = datetime.now()

        if self.failure_count >= self.failure_threshold:
            self.state = CircuitState.OPEN

# Usage
payment_circuit = CircuitBreaker(failure_threshold=3, recovery_timeout=30)

@app.post("/checkout")
async def checkout(order: Order):
    try:
        result = await payment_circuit.call(
            process_payment,
            order.payment_info
        )
        return {"status": "success", "transaction_id": result.id}
    except Exception as e:
        # Fallback behavior when circuit is open
        await queue_for_retry(order)
        return {"status": "queued", "message": "Payment processing delayed"}

Event Sourcing for Critical Operations

Never lose data again. Every change becomes an immutable event:

from datetime import datetime
from typing import List, Dict, Any
from enum import Enum

class EventType(Enum):
    USER_CREATED = "user_created"
    USER_UPDATED = "user_updated"
    ORDER_PLACED = "order_placed"
    PAYMENT_PROCESSED = "payment_processed"

class Event(BaseModel):
    id: str
    aggregate_id: str  # The entity this event relates to
    event_type: EventType
    event_data: Dict[str, Any]
    event_time: datetime
    event_version: int
    metadata: Dict[str, Any]  # User ID, IP, etc.

class EventStore:
    def __init__(self, db: AsyncSession):
        self.db = db

    async def append(self, event: Event):
        # Events are immutable - only INSERT, never UPDATE
        db_event = EventModel(
            id=event.id,
            aggregate_id=event.aggregate_id,
            event_type=event.event_type.value,
            event_data=json.dumps(event.event_data),
            event_time=event.event_time,
            event_version=event.event_version,
            metadata=json.dumps(event.metadata)
        )
        self.db.add(db_event)
        await self.db.commit()

        # Publish to event bus for real-time processing
        await event_bus.publish(event)

    async def get_events(self, aggregate_id: str) -> List[Event]:
        result = await self.db.execute(
            select(EventModel)
            .where(EventModel.aggregate_id == aggregate_id)
            .order_by(EventModel.event_version)
        )
        return [self._to_event(e) for e in result.scalars()]

    async def replay_to_state(self, aggregate_id: str) -> Dict[str, Any]:
        """Rebuild current state from event history"""
        events = await self.get_events(aggregate_id)
        state = {}

        for event in events:
            state = self.apply_event(state, event)

        return state

# Usage
@app.post("/users")
async def create_user(
    user_data: UserInput,
    db: AsyncSession = Depends(get_db),
    current_user: User = Depends(get_current_user)
):
    event = Event(
        id=generate_uuid(),
        aggregate_id=generate_uuid(),  # New user ID
        event_type=EventType.USER_CREATED,
        event_data=user_data.dict(),
        event_time=datetime.utcnow(),
        event_version=1,
        metadata={
            "created_by": current_user.id,
            "ip_address": request.client.host
        }
    )

    await event_store.append(event)

    # Build materialized view asynchronously
    background_tasks.add_task(rebuild_user_view, event.aggregate_id)

    return {"user_id": event.aggregate_id}

Distributed Tracing That Actually Helps

When your API calls five other services, you need to track requests across all of them:

from opentelemetry import trace
from opentelemetry.exporter.jaeger import JaegerExporter
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.instrumentation.fastapi import FastAPIInstrumentor
from opentelemetry.instrumentation.httpx import HTTPXClientInstrumentor
from opentelemetry.instrumentation.sqlalchemy import SQLAlchemyInstrumentor

# Set up tracing
trace.set_tracer_provider(TracerProvider())
tracer = trace.get_tracer(__name__)

# Export to Jaeger
jaeger_exporter = JaegerExporter(
    agent_host_name="localhost",
    agent_port=6831,
)

span_processor = BatchSpanProcessor(jaeger_exporter)
trace.get_tracer_provider().add_span_processor(span_processor)

# Auto-instrument libraries
FastAPIInstrumentor.instrument_app(app)
HTTPXClientInstrumentor().instrument()
SQLAlchemyInstrumentor().instrument(engine=engine.sync_engine)

# Manual instrumentation for business logic
@app.get("/complex-operation")
async def complex_operation(user_id: int):
    with tracer.start_as_current_span("complex_operation") as span:
        span.set_attribute("user.id", user_id)

        # Each step gets its own span
        with tracer.start_as_current_span("fetch_user"):
            user = await get_user(user_id)

        with tracer.start_as_current_span("calculate_recommendations"):
            recommendations = await calculate_recommendations(user)
            span.set_attribute("recommendation.count", len(recommendations))

        with tracer.start_as_current_span("fetch_metadata"):
            tasks = [fetch_metadata(r.id) for r in recommendations]
            metadata = await asyncio.gather(*tasks)

        return {"recommendations": recommendations, "metadata": metadata}

Testing Strategies That Actually Work

Unit tests are great, but they won't catch the bugs that take down production. Here's what will:

Chaos Testing Your API

import random
import asyncio
from fastapi import Request

@app.middleware("http")
async def chaos_middleware(request: Request, call_next):
    # Only in staging environment
    if os.getenv("ENABLE_CHAOS") != "true":
        return await call_next(request)

    chaos_type = random.choice(["none"] * 90 + ["delay", "error", "timeout"])

    if chaos_type == "delay":
        # Random delay between 1-5 seconds
        await asyncio.sleep(random.uniform(1, 5))
    elif chaos_type == "error" and random.random() < 0.1:
        # 10% chance of 500 error
        return JSONResponse(status_code=500, content={"error": "Chaos monkey!"})
    elif chaos_type == "timeout":
        # Simulate timeout
        await asyncio.sleep(31)

    return await call_next(request)

Load Testing That Mirrors Reality

# locustfile.py for realistic load testing
from locust import HttpUser, task, between
import random

class APIUser(HttpUser):
    wait_time = between(1, 3)

    def on_start(self):
        # Login once
        response = self.client.post("/auth/login", json={
            "username": f"user{random.randint(1, 1000)}",
            "password": "testpass"
        })
        self.token = response.json()["access_token"]
        self.headers = {"Authorization": f"Bearer {self.token}"}

    @task(10)
    def view_dashboard(self):
        self.client.get("/dashboard", headers=self.headers)

    @task(5)
    def create_post(self):
        self.client.post("/posts",
            headers=self.headers,
            json={"title": "Test", "content": "Lorem ipsum"}
        )

    @task(2)
    def heavy_operation(self):
        # Simulate expensive operations
        self.client.get(f"/analytics/{random.randint(1, 100)}",
                       headers=self.headers)

    @task(1)
    def websocket_test(self):
        # Test WebSocket connections
        with self.client.websocket_connect("/ws") as ws:
            ws.send_json({"type": "ping"})
            response = ws.receive_json()

The Reality Check

Here's the thing nobody tells you about building APIs—perfection is the enemy of shipping. I've seen too many developers spend months optimizing their database queries while their startup runs out of runway. Start with FastAPI's defaults, they're good enough for your first 100,000 users. Add caching when you actually need it. Implement circuit breakers when external services actually start failing.

But also, don't be naive. That authentication system I showed you? Implement it from day one. The monitoring setup? Get it running before you have users. The Docker configuration? Use it even for your MVP. These aren't premature optimizations—they're insurance policies.

The beauty of FastAPI is that it scales with you. You can start with a simple CRUD API on Friday and have it handling WebSockets, background tasks, and complex async operations by Monday. Try doing that with Django or Flask and you'll be refactoring for weeks.

One Last Thing

Every API I've built has taught me something new. The payment processor that had to handle exactly-once semantics. The social platform that needed real-time updates for 50,000 concurrent users. The analytics engine that processed 10GB of data per hour on a budget VPS. Each one pushed FastAPI in different ways, and it never let me down.

If you're starting a new project, grab FastLaunchAPI.dev and save yourself the setup time. It's got all the patterns from this post already implemented, tested, and documented. But more importantly, understand why these patterns exist. Know when to use connection pooling versus when to increase your database's max connections. Understand why event sourcing might save your ass when the auditors come knocking. Learn when WebSockets make sense and when they're overkill.

Building APIs isn't about following a checklist—it's about understanding trade-offs. FastAPI gives you the tools to make those trade-offs intelligently. Use them wisely, ship fast, and remember: the best API is the one that's actually running in production, serving real users, making real money.

Now stop reading blog posts and go build something.

Related Articles