Complete REST API reference for OpenWebAnswer.
Base URL: http://localhost:8000
OpenWebAnswer provides a REST API for question-answering with integrated caching, source tracking, and analytics.
Currently, OpenWebAnswer does not require authentication for public deployment. For production deployments, enable API key authentication via environment variables.
| Property | Value |
|---|---|
| Protocol | HTTP/HTTPS |
| Content-Type | application/json |
| Response Format | JSON |
| Authentication | Optional (disabled by default) |
Submit a question and get an answer with sources, followups, and analytics.
Request URL:
POST /api/queryRequest Parameters:
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
question |
string | β Yes | - | User's question (min 1 char, max 2000 chars) |
mode |
string | β No | hybrid |
Query mode: realtime (web search only), cached (vector store only), hybrid (intelligent selection) |
top_k |
integer | β No | 5 | Number of source documents to retrieve (1-20) |
stream |
boolean | β No | false | Stream response (line-by-line LLM output) |
Request Example:
curl -X POST http://localhost:8000/api/query \
-H "Content-Type: application/json" \
-d '{
"question": "What is machine learning?",
"mode": "hybrid",
"top_k": 5,
"stream": false
}'Response (200 OK):
{
"id": "550e8400-e29b-41d4-a716-446655440000",
"question": "What is machine learning?",
"answer": "Machine learning is a subset of artificial intelligence that enables systems to learn and improve from experience without being explicitly programmed...",
"sources": [
{
"url": "https://en.wikipedia.org/wiki/Machine_learning",
"title": "Machine learning - Wikipedia",
"snippet": "Machine learning (ML) is a subset of artificial intelligence...",
"relevance_score": 0.95,
"quality_score": 0.94,
"quality_reason": "High authority domain with comprehensive content",
"metadata": {
"source_type": "encyclopedia",
"language": "en"
}
},
{
"url": "https://www.ibm.com/topics/machine-learning",
"title": "What is Machine Learning?",
"snippet": "Machine learning algorithms build mathematical models based on sample data...",
"relevance_score": 0.92,
"quality_score": 0.87,
"quality_reason": "Trusted industry source with clear explanations",
"metadata": {
"source_type": "article",
"language": "en"
}
}
],
"followups": [
"What are the main types of machine learning algorithms?",
"How is machine learning used in industry?",
"What programming languages are best for machine learning?"
],
"mode": "hybrid",
"response_time_ms": 2534.5,
"sources_count": 2,
"average_quality_score": 0.905,
"timestamp": "2024-12-13T12:00:00",
"cache_hit": false
}Response Schema:
| Field | Type | Description |
|---|---|---|
id |
string (UUID) | Unique query identifier for tracking and analytics |
question |
string | Original user question |
answer |
string | Generated answer combining source information |
sources |
array[Document] | Retrieved source documents ranked by relevance |
sources[].url |
string | Source URL |
sources[].title |
string | Page title |
sources[].snippet |
string | Relevant text excerpt |
sources[].relevance_score |
float (0-1) | Relevance to query (0 = not relevant, 1 = highly relevant) |
sources[].quality_score |
float (0-1) | Source quality score based on authority and freshness |
sources[].quality_reason |
string | Explanation for quality score |
sources[].metadata |
object | Additional metadata (source type, language, etc.) |
followups |
array[string] | 3 suggested follow-up questions |
mode |
string | Query mode used (realtime, cached, hybrid) |
response_time_ms |
float | Total response time in milliseconds |
sources_count |
integer | Number of sources retrieved |
average_quality_score |
float (0-1) | Average quality score of all sources |
timestamp |
datetime | Response generation timestamp (ISO 8601) |
cache_hit |
boolean | Whether response was retrieved from cache |
Error Responses:
// 400 Bad Request - Invalid query
{
"status": "error",
"message": "Invalid question",
"detail": "Question must be between 1 and 2000 characters",
"timestamp": "2024-12-13T12:00:00"
}
// 500 Internal Server Error
{
"status": "error",
"message": "Failed to process query",
"detail": "LLM service unavailable. Please try again later.",
"timestamp": "2024-12-13T12:00:00"
}Query Modes:
-
realtime: Always performs web search, fetches new documents, generates fresh answer- Use for: Current events, breaking news, time-sensitive information
- Response time: 4-15 seconds
-
cached: Uses only cached documents and vector index- Use for: Faster responses, offline operation, reduced API calls
- Response time: 100-500ms
- May return stale information
-
hybrid(default): Checks cache first, falls back to web search if needed- Use for: Balanced approach, good for most queries
- Response time: 50-500ms if cached, 4-15s if web search needed
- Best overall performance/freshness trade-off
Example Python:
import requests
import json
url = "http://localhost:8000/api/query"
payload = {
"query": "What is Python used for?",
"mode": "hybrid"
}
response = requests.post(url, json=payload)
result = response.json()
print("Answer:", result["answer"])
print("Sources:", len(result["sources"]))
print("Follow-ups:", result["followups"])
print("Response time:", result["response_time_ms"], "ms")
print("Cache hit:", result["cache_hit"])All analytics endpoints are read-only and provide insights into system performance and query patterns.
Get overall system insights and analytics summary.
Request URL:
GET /api/analytics/insightsResponse (200 OK):
{
"total_queries": 1547,
"cache_hit_rate": 0.623,
"avg_response_time_ms": 2534.5,
"total_documents_indexed": 4283,
"vector_index_size": 34215,
"system_uptime_hours": 168.5
}Get the most frequently executed queries.
Request URL:
GET /api/analytics/top-queries?limit=10&offset=0Query Parameters:
| Parameter | Type | Default | Description |
|---|---|---|---|
limit |
integer | 10 | Number of queries to return (max 100) |
offset |
integer | 0 | Pagination offset |
Response (200 OK):
{
"queries": [
{
"query_text": "python tutorial",
"count": 156,
"avg_response_time_ms": 2400.5,
"last_executed": "2024-12-13T12:00:00"
},
{
"query_text": "machine learning",
"count": 143,
"avg_response_time_ms": 2650.3,
"last_executed": "2024-12-13T11:45:00"
}
],
"total": 2,
"limit": 10,
"offset": 0
}Get queries with the longest response times.
Request URL:
GET /api/analytics/slowest-queries?limit=10&offset=0Query Parameters:
| Parameter | Type | Default | Description |
|---|---|---|---|
limit |
integer | 10 | Number of queries to return (max 100) |
offset |
integer | 0 | Pagination offset |
Response (200 OK):
{
"queries": [
{
"query_text": "latest AI breakthroughs",
"response_time_ms": 12500.5,
"executed_at": "2024-12-13T10:30:00"
},
{
"query_text": "quantum computing explained",
"response_time_ms": 11200.3,
"executed_at": "2024-12-13T09:15:00"
}
],
"total": 2,
"limit": 10,
"offset": 0
}Get aggregated performance metrics across all queries.
Request URL:
GET /api/analytics/performanceResponse (200 OK):
{
"total_queries": 1547,
"avg_response_time_ms": 2534.5,
"median_response_time_ms": 2100.0,
"p95_response_time_ms": 8500.0,
"p99_response_time_ms": 12000.0,
"min_response_time_ms": 50.0,
"max_response_time_ms": 18234.0,
"cache_hit_rate": 0.623,
"web_search_rate": 0.377,
"avg_sources_per_query": 4.2,
"timestamp": "2024-12-13T12:00:00"
}Get statistics about indexed documents and caching.
Request URL:
GET /api/analytics/documentsResponse (200 OK):
{
"total_documents_indexed": 4283,
"total_chunks": 45892,
"documents_cached_7d": 1245,
"documents_cached_30d": 3156,
"avg_quality_score": 0.847,
"faiss_index_vectors": 34215,
"faiss_index_size_mb": 127.8,
"storage_used_mb": 245.3,
"last_index_rebuild": "2024-12-13T00:00:00",
"timestamp": "2024-12-13T12:00:00"
}Get comprehensive system health and component status.
Request URL:
GET /api/analytics/healthResponse (200 OK):
{
"status": "healthy",
"timestamp": "2024-12-13T12:00:00",
"version": "1.0.0",
"database": "ok",
"vector_store": "ok",
"llm": "ok",
"uptime_hours": 168.5,
"memory_usage_mb": 512.3,
"cpu_usage_percent": 23.5
}Possible Status Values:
healthy: All systems operationaldegraded: Some components have issues but system is functionalunhealthy: One or more critical components are failing
Get detailed cache hit rate statistics.
Request URL:
GET /api/analytics/cache-hit-rateResponse (200 OK):
{
"overall_cache_hit_rate": 0.623,
"cache_hits": 1547,
"cache_misses": 934,
"total_queries": 2481,
"by_mode": {
"hybrid": 0.65,
"cached": 0.95,
"realtime": 0.0
},
"by_time_range": {
"last_hour": 0.58,
"last_24h": 0.62,
"last_7d": 0.623
},
"cache_efficiency": {
"response_time_saved_ms": 8500000,
"api_calls_saved": 1547,
"processing_time_saved_hours": 12.4
},
"timestamp": "2024-12-13T12:00:00"
}Submit user rating/feedback for a query result.
Request URL:
POST /api/analytics/rate-queryRequest Body:
{
"query_id": "550e8400-e29b-41d4-a716-446655440000",
"rating": 5,
"feedback": "Very helpful answer",
"tags": ["accurate", "comprehensive"]
}Request Parameters:
| Field | Type | Required | Description |
|---|---|---|---|
query_id |
string (UUID) | β Yes | ID of the query to rate |
rating |
integer | β Yes | Rating (1-5 stars) |
feedback |
string | β No | Optional text feedback |
tags |
array[string] | β No | Optional tags (accurate, helpful, outdated, etc.) |
Response (200 OK):
{
"status": "success",
"message": "Rating recorded",
"query_id": "550e8400-e29b-41d4-a716-446655440000"
}Get a comprehensive analytics report.
Request URL:
GET /api/analytics/report?time_range=7dQuery Parameters:
| Parameter | Type | Default | Description |
|---|---|---|---|
time_range |
string | 7d |
1d, 7d, 30d, 90d |
Response (200 OK):
{
"time_period": "last 7 days",
"total_queries": 2481,
"unique_queries": 1547,
"cache_statistics": {
"hit_rate": 0.623,
"hits": 1547,
"misses": 934
},
"performance": {
"avg_response_time_ms": 2534.5,
"median_response_time_ms": 2100.0,
"p95_response_time_ms": 8500.0
},
"documents": {
"total_indexed": 4283,
"total_cached": 3156,
"avg_quality_score": 0.847
},
"top_queries": [
{"text": "python", "count": 156},
{"text": "machine learning", "count": 143}
],
"timestamp": "2024-12-13T12:00:00"
}Get quick summary statistics.
Request URL:
GET /api/analytics/summaryResponse (200 OK):
{
"total_queries_processed": 2481,
"cache_hit_rate": 0.623,
"avg_response_time_ms": 2534.5,
"system_status": "healthy",
"documents_indexed": 4283,
"vector_store_status": "ok"
}Check system health and component status.
Request:
GET /api/health HTTP/1.1
Host: localhost:8000Response (200 OK):
{
"status": "healthy",
"timestamp": "2025-12-10T15:30:45Z",
"version": "1.0.0",
"components": {
"database": {
"status": "connected",
"response_time_ms": 2
},
"vector_store": {
"status": "ready",
"total_vectors": 34215,
"index_size_mb": 127.8
},
"llm": {
"status": "connected",
"provider": "openai",
"model": "gpt-3.5-turbo"
},
"search_engine": {
"status": "operational",
"engine": "duckduckgo",
"last_check": "2025-12-10T15:29:45Z"
},
"cache": {
"status": "operational",
"items": 1547,
"size_mb": 145.3
}
},
"background_jobs": {
"cleanup_job": "running",
"index_maintenance": "running"
}
}Status Values:
healthy- All systems operationaldegraded- Some components have issues but system is functionalunhealthy- Critical components are down
All errors follow this standard format:
{
"detail": "Human-readable error message",
"error_code": "ERROR_CODE",
"request_id": "uuid-for-tracking",
"timestamp": "2025-12-10T15:30:45Z"
}| Code | Meaning | Example |
|---|---|---|
| 200 | OK | Query successful |
| 400 | Bad Request | Invalid query format |
| 404 | Not Found | Query ID doesn't exist |
| 429 | Too Many Requests | Rate limit exceeded |
| 500 | Internal Server Error | Unexpected error |
| 503 | Service Unavailable | System maintenance |
// 400: Query too short
{
"detail": "Query must be at least 3 characters",
"error_code": "INVALID_QUERY_LENGTH"
}
// 400: Invalid mode
{
"detail": "Mode must be one of: realtime, cached, hybrid",
"error_code": "INVALID_MODE"
}
// 429: Rate limited
{
"detail": "Rate limit exceeded. Max 60 requests per minute.",
"error_code": "RATE_LIMIT_EXCEEDED"
}
// 503: Vector store not ready
{
"detail": "Vector store is initializing. Try again in a moment.",
"error_code": "VECTOR_STORE_INITIALIZING"
}Rate limiting is configurable in .env:
RATE_LIMIT_ENABLED=false # Set to true for production
RATE_LIMIT_REQUESTS_PER_MINUTE=60
RATE_LIMIT_REQUESTS_PER_HOUR=1000When rate limited, response includes retry information:
{
"detail": "Rate limit exceeded",
"error_code": "RATE_LIMIT_EXCEEDED",
"retry_after_seconds": 45
}import requests
from datetime import datetime
class OpenWebAnswerClient:
def __init__(self, base_url="http://localhost:8000"):
self.base_url = base_url
self.session = requests.Session()
def ask(self, query: str, mode: str = "hybrid"):
"""Ask a question"""
response = self.session.post(
f"{self.base_url}/api/query",
json={"query": query, "mode": mode}
)
response.raise_for_status()
return response.json()
def rate_response(self, query_id: str, rating: int, feedback: str = ""):
"""Rate a query response"""
response = self.session.post(
f"{self.base_url}/api/analytics/rate-query",
json={
"query_id": query_id,
"rating": rating,
"feedback": feedback
}
)
response.raise_for_status()
return response.json()
def get_cache_status(self):
"""Get cache statistics"""
response = self.session.get(
f"{self.base_url}/api/analytics/cache-status"
)
response.raise_for_status()
return response.json()
# Usage
client = OpenWebAnswerClient()
# Ask a question
result = client.ask("What is machine learning?")
print("Answer:", result["answer"][:100] + "...")
print("Sources:", len(result["sources"]))
print("Cache hit:", result["cache_hit"])
print("Response time:", result["response_time_ms"], "ms")
# Rate the response
client.rate_response(
query_id=result["query_id"],
rating=5,
feedback="Excellent answer!"
)
# Get cache stats
stats = client.get_cache_status()
print("Cache hit rate:", f"{stats['cache_hit_rate']*100:.1f}%")async function askQuestion(query, mode = 'hybrid') {
const response = await fetch('http://localhost:8000/api/query', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ query, mode })
});
if (!response.ok) {
throw new Error(`API error: ${response.statusText}`);
}
return response.json();
}
async function rateResponse(queryId, rating, feedback = '') {
const response = await fetch('http://localhost:8000/api/analytics/rate-query', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ query_id: queryId, rating, feedback })
});
return response.json();
}
// Usage
const result = await askQuestion('What is Python?');
console.log('Answer:', result.answer.substring(0, 100) + '...');
console.log('Sources:', result.sources.length);
console.log('Response time:', result.response_time_ms, 'ms');
// Rate it
await rateResponse(result.query_id, 5, 'Great answer!');# Ask a question
curl -X POST http://localhost:8000/api/query \
-H "Content-Type: application/json" \
-d '{
"query": "What is quantum computing?",
"mode": "hybrid"
}' | jq .
# Rate a response
curl -X POST http://localhost:8000/api/analytics/rate-query \
-H "Content-Type: application/json" \
-d '{
"query_id": "550e8400-e29b-41d4-a716-446655440000",
"rating": 5
}'
# Get cache status
curl http://localhost:8000/api/analytics/cache-status | jq .
# Get top queries
curl "http://localhost:8000/api/analytics/top-queries?limit=5" | jq .
# Check health
curl http://localhost:8000/api/health | jq .Currently on API version 1.0.0. Future versions will be prefixed:
/api/v1/query # Current (implied)
/api/v2/query # Future versionsBackward compatibility is maintained within major versions.
Last Updated: December 13, 2025 (Phase 7: Final Verification) API Version: 1.0.0
#!/bin/bash
QUESTION="What is the capital of France?"
curl -s -X POST "http://127.0.0.1:8000/api/query" \
-H "Content-Type: application/json" \
-d "{\"question\": \"$QUESTION\"}" | jq '.answer'
# Output: The capital of France is Paris...import requests
import json
from concurrent.futures import ThreadPoolExecutor
questions = [
"What is machine learning?",
"How does AI work?",
"What is deep learning?"
]
def query(question):
response = requests.post(
"http://127.0.0.1:8000/api/query",
json={"question": question}
)
return response.json()
# Run queries in parallel
with ThreadPoolExecutor(max_workers=3) as executor:
results = list(executor.map(query, questions))
# Process results
for i, result in enumerate(results):
print(f"Q: {questions[i]}")
print(f"A: {result['answer'][:100]}...")
print(f"Cache Hit: {result['cache_hit']}\n")import requests
import matplotlib.pyplot as plt
from datetime import datetime, timedelta
# Get performance metrics
response = requests.get("http://127.0.0.1:8000/api/analytics/performance")
metrics = response.json()
print(f"P50: {metrics['response_time_metrics']['median_ms']}ms")
print(f"P95: {metrics['response_time_metrics']['p95_ms']}ms")
print(f"P99: {metrics['response_time_metrics']['p99_ms']}ms")
print(f"Cache Hit Rate: {metrics['cache_metrics']['cache_hit_rate']:.1%}")
# Visualize
times = [50, 100, 200, 500, 1000, 2000, 5000, 10000]
plt.hist(times, bins=20)
plt.xlabel("Response Time (ms)")
plt.ylabel("Frequency")
plt.title("Response Time Distribution")
plt.show()// Collect user feedback and improve system
const queryAndGetFeedback = async (question) => {
// Get answer
const response = await fetch('http://127.0.0.1:8000/api/query', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ question })
});
const result = await response.json();
// Display answer
document.getElementById('answer').textContent = result.answer;
// Show feedback buttons
document.getElementById('feedback').innerHTML = `
<button onclick="rateFeedback('${result.id}', 5)">π Helpful</button>
<button onclick="rateFeedback('${result.id}', 1)">π Not Helpful</button>
`;
return result.id;
};
const rateFeedback = async (queryId, rating) => {
await fetch('http://127.0.0.1:8000/api/analytics/rate-query', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({
query_id: queryId,
rating: rating,
feedback: document.getElementById('feedback-text').value
})
});
alert('Thank you for your feedback!');
};Full interactive API documentation available at:
-
Issues: Report bugs and request features on GitHub
-
Documentation: See README.md and ARCHITECTURE.md
-
Development: See DEVELOPMENT.md
Last Updated: December 13, 2025 (Phase 7: Final Verification)
Version: 1.0.0