Back to Blog
gpt-5
fastapi
celery
python
ai
web-development
api

How to Build a Smart Website Analyzer with GPT-5, FastAPI, and Celery

Step-by-step guide to creating a scalable API that uses GPT-5 to analyze websites, featuring FastAPI and Celery for background processing.

Niklas L.
8 min read

How to Build a Smart Website Analyzer with GPT-5, FastAPI, and Celery

Imagine this — you want to quickly understand what a website is about: who they serve, how they make money, what makes them tick in their market. Doing this by hand takes hours of research. But with a smart system powered by GPT-5, you can get detailed insights in seconds.

This post walks you through building an API that accepts any website URL, analyzes it with GPT-5 behind the scenes, and gives you a full business analysis. The trick is combining FastAPI for the web part, Celery for background processing, and GPT-5 for the AI magic.

Let’s break down the key pieces, code snippets included, and explain what’s really going on.


Step 1: Accept the Website URL and Kick Off the Analysis

@router.post("/analyze")
async def analyze_website(request: WebsiteAnalysisRequest):
    url = str(request.url)
    task = analyze_any_website.delay(url)
    return {
        "task_id": task.id,
        "status": "started",
        "message": f"AI analysis started for {url}",
        "check_results": f"/results/{task.id}"
    }

What’s happening here?

This is your entry point — the API endpoint clients hit when they want a website analyzed.

Why is this great? The API stays lightning fast. The client doesn’t have to wait for GPT-5 to finish; they can check back when ready. This pattern is key to scaling any system with slow, heavy operations.


Step 2: Fetch Website Content with Proper Headers and Limits

headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36'
}
response = requests.get(url, headers=headers, timeout=10)
website_content = response.text[:8000]  # Limit content size

Why this matters

Web scraping can be tricky. Some sites block requests that don’t look like they come from real browsers. That’s why we add a realistic User-Agent header — it’s a little trick to convince servers we’re a normal browser and avoid being blocked.

Next, we fetch the content with a timeout of 10 seconds. This means if the site is slow or offline, our worker won’t hang forever. This keeps your system responsive.

We also cap the fetched content to 8,000 characters. GPT-5 has input limits, and the cost of processing huge inputs is higher. Plus, the most important info — like the homepage’s key messaging — usually appears near the top. This is a practical way to balance detail and performance.


Step 3: Prepare a Thoughtful GPT-5 Prompt

template = """You are a business analyst and web expert. Analyze this website and provide a comprehensive business report.

Create a detailed analysis including:
1. BUSINESS MODEL: What does this company do and how do they make money?
2. TARGET AUDIENCE: Who are their customers?
3. COMPETITIVE POSITION: How do they compare in their market?
4. WEBSITE QUALITY: User experience and design analysis
5. GROWTH OPPORTUNITIES: What could they improve or expand?
6. RISK ASSESSMENT: Potential challenges or threats
7. MARKETING STRATEGY: How do they attract customers?
8. REVENUE ESTIMATE: Educated guess on their revenue range
9. INVESTMENT POTENTIAL: Would you invest in this company?
10. ACTIONABLE RECOMMENDATIONS: 3-5 specific improvements

Be specific and insightful. Return as JSON with detailed analysis."""

Why this prompt is the backbone of your AI’s understanding

GPT models respond best to clear instructions. Here, we act like a consulting firm, telling GPT-5 to wear the hat of a business analyst and deliver a deep dive report.

We explicitly list 10 categories we want analyzed — from the business model and target audience, all the way to investment potential and actionable advice. This guides GPT-5 to cover all important angles and avoid vague, generic answers.

We also instruct GPT-5 to return the response in JSON. This is crucial because it makes downstream parsing and consumption much easier — instead of digging through messy paragraphs, your system can programmatically handle the data.


Step 4: Inject Real Website Data into the Prompt

human_template = """
Website URL: {url}
Domain: {domain}

Website Content Analysis:
{content}

Provide a comprehensive business analysis of this website as if you're a consulting firm.
"""

Why separate system and human parts?

This structure comes from Langchain’s approach to conversational prompts:

By keeping these separate, we create a cleaner, more maintainable prompt. Plus, it allows for dynamic injection of the URL, domain, and website content, making the prompt flexible for any site.


Step 5: Format and Send Prompt to GPT-5

chat_prompt = ChatPromptTemplate.from_messages([
    ("system", template),
    ("human", human_template)
])

formatted_prompt = chat_prompt.format_messages(
    url=url,
    domain=domain,
    content=website_content
)

result = chat_model.invoke(formatted_prompt)

Breaking down this key interaction

This abstraction through Langchain makes prompt management much easier, especially as your app grows or you add more complex conversations.


Step 6: Handle GPT-5’s Response Safely

try:
    ai_analysis = json.loads(result.content)
except json.JSONDecodeError:
    ai_analysis = result.content

Why error handling here is non-negotiable

GPT-5 tries to return JSON, but sometimes it slips up — maybe a missing comma or stray character. If your code blindly tries to parse without catching exceptions, your whole system can crash.

By wrapping it in a try-except, you gracefully handle bad JSON by falling back to the raw text. This means your service stays up and you can later debug or fix the prompt.

This approach acknowledges the imperfection of AI while keeping your system robust.


Step 7: Compose the Result with Metadata

return {
    "status": "success",
    "data": {
        "analyzed_url": url,
        "domain": domain,
        "generated_at": datetime.now().isoformat(),
        "analysis": ai_analysis,
        "analysis_type": "website_business_intelligence",
        "content_preview": website_content[:500] + "...",
        "report_id": f"website_analysis_{domain.replace('.', '_')}_{datetime.now().strftime('%H%M%S')}"
    }
}

What makes this response valuable?

This dictionary packs more than just the raw analysis. It includes:

This level of detail is crucial when you want to build real apps that store, search, or display AI-generated data.


Step 8: Checking Task Progress with a Results Endpoint

@router.get("/results/{task_id}")
async def get_results(task_id: str):
    task_result = AsyncResult(task_id)

    if task_result.status == "SUCCESS":
        return {
            "status": "completed",
            "website": task_result.result["data"]["analyzed_url"],
            "analysis": task_result.result["data"]["analysis"],
            "generated_at": task_result.result["data"]["generated_at"],
        }
    elif task_result.status in ["PENDING", "STARTED"]:
        return {
            "status": "processing",
            "message": "Analysis still running... check back soon."
        }
    else:
        return {"status": "failed", "error": str(task_result.info)}

Why this endpoint is

a user-friendly must-have

Since analysis runs asynchronously, users need a way to check if their report is ready.

This clear feedback loop improves user experience and lets clients build nice UIs with loading indicators or notifications.


Final Thoughts and Best Practices


Watch the Full Video Walkthrough

If you want to see all this put together in real time, check out my full video tutorial here: Watch on YouTube

It’s a great way to follow along and deepen your understanding.


FAQ

Q: Why use Celery instead of analyzing directly in FastAPI? Because GPT-5 can take several seconds or longer to respond. Running it inside FastAPI would block other requests and degrade performance. Celery runs analysis in the background, keeping the API responsive.

Q: Can this work with other AI models? Absolutely. The prompt design and task queue structure are model-agnostic. Just swap GPT-5 for any other supported model.

Q: How do I secure this API? Add authentication layers (OAuth, API keys), validate inputs strictly, and use HTTPS.

Related Articles