How to Build a Smart Website Analyzer with GPT-5, FastAPI, and Celery
Step-by-step guide to creating a scalable API that uses GPT-5 to analyze websites, featuring FastAPI and Celery for background processing.
How to Build a Smart Website Analyzer with GPT-5, FastAPI, and Celery
Imagine this — you want to quickly understand what a website is about: who they serve, how they make money, what makes them tick in their market. Doing this by hand takes hours of research. But with a smart system powered by GPT-5, you can get detailed insights in seconds.
This post walks you through building an API that accepts any website URL, analyzes it with GPT-5 behind the scenes, and gives you a full business analysis. The trick is combining FastAPI for the web part, Celery for background processing, and GPT-5 for the AI magic.
Let’s break down the key pieces, code snippets included, and explain what’s really going on.
Step 1: Accept the Website URL and Kick Off the Analysis
@router.post("/analyze")
async def analyze_website(request: WebsiteAnalysisRequest):
url = str(request.url)
task = analyze_any_website.delay(url)
return {
"task_id": task.id,
"status": "started",
"message": f"AI analysis started for {url}",
"check_results": f"/results/{task.id}"
}
What’s happening here?
This is your entry point — the API endpoint clients hit when they want a website analyzed.
- The client sends a JSON payload with the
url. - We immediately convert the URL to a string for safety.
- Instead of analyzing the website right here (which could take several seconds or more), we hand off the job to a Celery task using
.delay(url). This queues the job asynchronously. - Celery returns a
task.id, a unique identifier for this analysis job. - We reply instantly with a message that the job started, the task ID, and a URL to check the results later.
Why is this great? The API stays lightning fast. The client doesn’t have to wait for GPT-5 to finish; they can check back when ready. This pattern is key to scaling any system with slow, heavy operations.
Step 2: Fetch Website Content with Proper Headers and Limits
headers = {
'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36'
}
response = requests.get(url, headers=headers, timeout=10)
website_content = response.text[:8000] # Limit content size
Why this matters
Web scraping can be tricky. Some sites block requests that don’t look like they come from real browsers. That’s why we add a realistic User-Agent header — it’s a little trick to convince servers we’re a normal browser and avoid being blocked.
Next, we fetch the content with a timeout of 10 seconds. This means if the site is slow or offline, our worker won’t hang forever. This keeps your system responsive.
We also cap the fetched content to 8,000 characters. GPT-5 has input limits, and the cost of processing huge inputs is higher. Plus, the most important info — like the homepage’s key messaging — usually appears near the top. This is a practical way to balance detail and performance.
Step 3: Prepare a Thoughtful GPT-5 Prompt
template = """You are a business analyst and web expert. Analyze this website and provide a comprehensive business report.
Create a detailed analysis including:
1. BUSINESS MODEL: What does this company do and how do they make money?
2. TARGET AUDIENCE: Who are their customers?
3. COMPETITIVE POSITION: How do they compare in their market?
4. WEBSITE QUALITY: User experience and design analysis
5. GROWTH OPPORTUNITIES: What could they improve or expand?
6. RISK ASSESSMENT: Potential challenges or threats
7. MARKETING STRATEGY: How do they attract customers?
8. REVENUE ESTIMATE: Educated guess on their revenue range
9. INVESTMENT POTENTIAL: Would you invest in this company?
10. ACTIONABLE RECOMMENDATIONS: 3-5 specific improvements
Be specific and insightful. Return as JSON with detailed analysis."""
Why this prompt is the backbone of your AI’s understanding
GPT models respond best to clear instructions. Here, we act like a consulting firm, telling GPT-5 to wear the hat of a business analyst and deliver a deep dive report.
We explicitly list 10 categories we want analyzed — from the business model and target audience, all the way to investment potential and actionable advice. This guides GPT-5 to cover all important angles and avoid vague, generic answers.
We also instruct GPT-5 to return the response in JSON. This is crucial because it makes downstream parsing and consumption much easier — instead of digging through messy paragraphs, your system can programmatically handle the data.
Step 4: Inject Real Website Data into the Prompt
human_template = """
Website URL: {url}
Domain: {domain}
Website Content Analysis:
{content}
Provide a comprehensive business analysis of this website as if you're a consulting firm.
"""
Why separate system and human parts?
This structure comes from Langchain’s approach to conversational prompts:
- The system message (
template) sets the overall context and instructions. - The human message (
human_template) includes the actual data we want analyzed.
By keeping these separate, we create a cleaner, more maintainable prompt. Plus, it allows for dynamic injection of the URL, domain, and website content, making the prompt flexible for any site.
Step 5: Format and Send Prompt to GPT-5
chat_prompt = ChatPromptTemplate.from_messages([
("system", template),
("human", human_template)
])
formatted_prompt = chat_prompt.format_messages(
url=url,
domain=domain,
content=website_content
)
result = chat_model.invoke(formatted_prompt)
Breaking down this key interaction
ChatPromptTemplate.from_messagescombines our system and human templates into a single prompt object.- We call
format_messagesto insert the actual website URL, domain, and content into the prompt placeholders. chat_model.invoke(formatted_prompt)sends the final prompt to GPT-5 and waits for the AI’s response.
This abstraction through Langchain makes prompt management much easier, especially as your app grows or you add more complex conversations.
Step 6: Handle GPT-5’s Response Safely
try:
ai_analysis = json.loads(result.content)
except json.JSONDecodeError:
ai_analysis = result.content
Why error handling here is non-negotiable
GPT-5 tries to return JSON, but sometimes it slips up — maybe a missing comma or stray character. If your code blindly tries to parse without catching exceptions, your whole system can crash.
By wrapping it in a try-except, you gracefully handle bad JSON by falling back to the raw text. This means your service stays up and you can later debug or fix the prompt.
This approach acknowledges the imperfection of AI while keeping your system robust.
Step 7: Compose the Result with Metadata
return {
"status": "success",
"data": {
"analyzed_url": url,
"domain": domain,
"generated_at": datetime.now().isoformat(),
"analysis": ai_analysis,
"analysis_type": "website_business_intelligence",
"content_preview": website_content[:500] + "...",
"report_id": f"website_analysis_{domain.replace('.', '_')}_{datetime.now().strftime('%H%M%S')}"
}
}
What makes this response valuable?
This dictionary packs more than just the raw analysis. It includes:
- The original URL and domain for easy reference.
- A timestamp of when the report was generated — handy for logs and freshness checks.
- The full AI analysis structured as JSON or raw text.
- A preview snippet of the website content, useful if you want to display it alongside the report.
- A unique
report_idconstructed from domain and timestamp, which can be used as a key if storing in databases or file systems.
This level of detail is crucial when you want to build real apps that store, search, or display AI-generated data.
Step 8: Checking Task Progress with a Results Endpoint
@router.get("/results/{task_id}")
async def get_results(task_id: str):
task_result = AsyncResult(task_id)
if task_result.status == "SUCCESS":
return {
"status": "completed",
"website": task_result.result["data"]["analyzed_url"],
"analysis": task_result.result["data"]["analysis"],
"generated_at": task_result.result["data"]["generated_at"],
}
elif task_result.status in ["PENDING", "STARTED"]:
return {
"status": "processing",
"message": "Analysis still running... check back soon."
}
else:
return {"status": "failed", "error": str(task_result.info)}
Why this endpoint is
a user-friendly must-have
Since analysis runs asynchronously, users need a way to check if their report is ready.
- If the task succeeded, we return the detailed analysis.
- If it’s still running, we let users know to hang tight.
- If it failed, we send back error details to aid troubleshooting.
This clear feedback loop improves user experience and lets clients build nice UIs with loading indicators or notifications.
Final Thoughts and Best Practices
- Cache your results! If multiple users analyze the same website, don’t pay GPT-5 every time. Cache and reuse.
- Validate URLs to block malicious requests and protect your backend.
- Set timeouts and retry policies on requests to avoid stalled workers.
- Store results in a database for querying and dashboards.
- Secure your API keys with environment variables, never hardcode secrets.
- Monitor your Celery workers and logs for errors and performance bottlenecks.
Watch the Full Video Walkthrough
If you want to see all this put together in real time, check out my full video tutorial here: Watch on YouTube
It’s a great way to follow along and deepen your understanding.
FAQ
Q: Why use Celery instead of analyzing directly in FastAPI? Because GPT-5 can take several seconds or longer to respond. Running it inside FastAPI would block other requests and degrade performance. Celery runs analysis in the background, keeping the API responsive.
Q: Can this work with other AI models? Absolutely. The prompt design and task queue structure are model-agnostic. Just swap GPT-5 for any other supported model.
Q: How do I secure this API? Add authentication layers (OAuth, API keys), validate inputs strictly, and use HTTPS.
Related Articles
Modern API Design: Balancing Speed, Maintainability, and Developer Experience
A deep dive into best practices for designing APIs today, including async patterns, structured code, deployment strategies, and performance tradeoffs—plus a subtle nod to FastAPI for rapid development.
From Zero to Production: Launching Your First SaaS with FastAPI in 30 Days
A practical roadmap for solo founders to go from idea to live SaaS in a month using FastAPI — simple, fast, and realistic.
How I Automated My Side Hustle with FastAPI (and Made My First $500)
A personal story of how I used FastAPI to turn repetitive freelance work into a simple automation tool that started earning money on its own.