Skip to content

Latest commit

Β 

History

History
984 lines (767 loc) Β· 22.1 KB

File metadata and controls

984 lines (767 loc) Β· 22.1 KB

API Documentation

Complete REST API reference for OpenWebAnswer.

Base URL: http://localhost:8000


Table of Contents

  1. Overview
  2. Query Endpoint
  3. Analytics Endpoints
  4. Health & Status
  5. Error Handling
  6. Examples

Overview

OpenWebAnswer provides a REST API for question-answering with integrated caching, source tracking, and analytics.

Authentication

Currently, OpenWebAnswer does not require authentication for public deployment. For production deployments, enable API key authentication via environment variables.

Base Information

Property Value
Protocol HTTP/HTTPS
Content-Type application/json
Response Format JSON
Authentication Optional (disabled by default)

Query Endpoint

POST /api/query

Submit a question and get an answer with sources, followups, and analytics.

Request URL:

POST /api/query

Request Parameters:

Parameter Type Required Default Description
question string βœ… Yes - User's question (min 1 char, max 2000 chars)
mode string ❌ No hybrid Query mode: realtime (web search only), cached (vector store only), hybrid (intelligent selection)
top_k integer ❌ No 5 Number of source documents to retrieve (1-20)
stream boolean ❌ No false Stream response (line-by-line LLM output)

Request Example:

curl -X POST http://localhost:8000/api/query \
  -H "Content-Type: application/json" \
  -d '{
    "question": "What is machine learning?",
    "mode": "hybrid",
    "top_k": 5,
    "stream": false
  }'

Response (200 OK):

{
  "id": "550e8400-e29b-41d4-a716-446655440000",
  "question": "What is machine learning?",
  "answer": "Machine learning is a subset of artificial intelligence that enables systems to learn and improve from experience without being explicitly programmed...",
  "sources": [
    {
      "url": "https://en.wikipedia.org/wiki/Machine_learning",
      "title": "Machine learning - Wikipedia",
      "snippet": "Machine learning (ML) is a subset of artificial intelligence...",
      "relevance_score": 0.95,
      "quality_score": 0.94,
      "quality_reason": "High authority domain with comprehensive content",
      "metadata": {
        "source_type": "encyclopedia",
        "language": "en"
      }
    },
    {
      "url": "https://www.ibm.com/topics/machine-learning",
      "title": "What is Machine Learning?",
      "snippet": "Machine learning algorithms build mathematical models based on sample data...",
      "relevance_score": 0.92,
      "quality_score": 0.87,
      "quality_reason": "Trusted industry source with clear explanations",
      "metadata": {
        "source_type": "article",
        "language": "en"
      }
    }
  ],
  "followups": [
    "What are the main types of machine learning algorithms?",
    "How is machine learning used in industry?",
    "What programming languages are best for machine learning?"
  ],
  "mode": "hybrid",
  "response_time_ms": 2534.5,
  "sources_count": 2,
  "average_quality_score": 0.905,
  "timestamp": "2024-12-13T12:00:00",
  "cache_hit": false
}

Response Schema:

Field Type Description
id string (UUID) Unique query identifier for tracking and analytics
question string Original user question
answer string Generated answer combining source information
sources array[Document] Retrieved source documents ranked by relevance
sources[].url string Source URL
sources[].title string Page title
sources[].snippet string Relevant text excerpt
sources[].relevance_score float (0-1) Relevance to query (0 = not relevant, 1 = highly relevant)
sources[].quality_score float (0-1) Source quality score based on authority and freshness
sources[].quality_reason string Explanation for quality score
sources[].metadata object Additional metadata (source type, language, etc.)
followups array[string] 3 suggested follow-up questions
mode string Query mode used (realtime, cached, hybrid)
response_time_ms float Total response time in milliseconds
sources_count integer Number of sources retrieved
average_quality_score float (0-1) Average quality score of all sources
timestamp datetime Response generation timestamp (ISO 8601)
cache_hit boolean Whether response was retrieved from cache

Error Responses:

// 400 Bad Request - Invalid query
{
  "status": "error",
  "message": "Invalid question",
  "detail": "Question must be between 1 and 2000 characters",
  "timestamp": "2024-12-13T12:00:00"
}

// 500 Internal Server Error
{
  "status": "error",
  "message": "Failed to process query",
  "detail": "LLM service unavailable. Please try again later.",
  "timestamp": "2024-12-13T12:00:00"
}

Query Modes:

  • realtime: Always performs web search, fetches new documents, generates fresh answer

    • Use for: Current events, breaking news, time-sensitive information
    • Response time: 4-15 seconds
  • cached: Uses only cached documents and vector index

    • Use for: Faster responses, offline operation, reduced API calls
    • Response time: 100-500ms
    • May return stale information
  • hybrid (default): Checks cache first, falls back to web search if needed

    • Use for: Balanced approach, good for most queries
    • Response time: 50-500ms if cached, 4-15s if web search needed
    • Best overall performance/freshness trade-off

Example Python:

import requests
import json

url = "http://localhost:8000/api/query"
payload = {
    "query": "What is Python used for?",
    "mode": "hybrid"
}

response = requests.post(url, json=payload)
result = response.json()

print("Answer:", result["answer"])
print("Sources:", len(result["sources"]))
print("Follow-ups:", result["followups"])
print("Response time:", result["response_time_ms"], "ms")
print("Cache hit:", result["cache_hit"])

Analytics Endpoints

All analytics endpoints are read-only and provide insights into system performance and query patterns.

GET /api/analytics/insights

Get overall system insights and analytics summary.

Request URL:

GET /api/analytics/insights

Response (200 OK):

{
  "total_queries": 1547,
  "cache_hit_rate": 0.623,
  "avg_response_time_ms": 2534.5,
  "total_documents_indexed": 4283,
  "vector_index_size": 34215,
  "system_uptime_hours": 168.5
}

GET /api/analytics/top-queries

Get the most frequently executed queries.

Request URL:

GET /api/analytics/top-queries?limit=10&offset=0

Query Parameters:

Parameter Type Default Description
limit integer 10 Number of queries to return (max 100)
offset integer 0 Pagination offset

Response (200 OK):

{
  "queries": [
    {
      "query_text": "python tutorial",
      "count": 156,
      "avg_response_time_ms": 2400.5,
      "last_executed": "2024-12-13T12:00:00"
    },
    {
      "query_text": "machine learning",
      "count": 143,
      "avg_response_time_ms": 2650.3,
      "last_executed": "2024-12-13T11:45:00"
    }
  ],
  "total": 2,
  "limit": 10,
  "offset": 0
}

GET /api/analytics/slowest-queries

Get queries with the longest response times.

Request URL:

GET /api/analytics/slowest-queries?limit=10&offset=0

Query Parameters:

Parameter Type Default Description
limit integer 10 Number of queries to return (max 100)
offset integer 0 Pagination offset

Response (200 OK):

{
  "queries": [
    {
      "query_text": "latest AI breakthroughs",
      "response_time_ms": 12500.5,
      "executed_at": "2024-12-13T10:30:00"
    },
    {
      "query_text": "quantum computing explained",
      "response_time_ms": 11200.3,
      "executed_at": "2024-12-13T09:15:00"
    }
  ],
  "total": 2,
  "limit": 10,
  "offset": 0
}

GET /api/analytics/performance

Get aggregated performance metrics across all queries.

Request URL:

GET /api/analytics/performance

Response (200 OK):

{
  "total_queries": 1547,
  "avg_response_time_ms": 2534.5,
  "median_response_time_ms": 2100.0,
  "p95_response_time_ms": 8500.0,
  "p99_response_time_ms": 12000.0,
  "min_response_time_ms": 50.0,
  "max_response_time_ms": 18234.0,
  "cache_hit_rate": 0.623,
  "web_search_rate": 0.377,
  "avg_sources_per_query": 4.2,
  "timestamp": "2024-12-13T12:00:00"
}

GET /api/analytics/documents

Get statistics about indexed documents and caching.

Request URL:

GET /api/analytics/documents

Response (200 OK):

{
  "total_documents_indexed": 4283,
  "total_chunks": 45892,
  "documents_cached_7d": 1245,
  "documents_cached_30d": 3156,
  "avg_quality_score": 0.847,
  "faiss_index_vectors": 34215,
  "faiss_index_size_mb": 127.8,
  "storage_used_mb": 245.3,
  "last_index_rebuild": "2024-12-13T00:00:00",
  "timestamp": "2024-12-13T12:00:00"
}

GET /api/analytics/health

Get comprehensive system health and component status.

Request URL:

GET /api/analytics/health

Response (200 OK):

{
  "status": "healthy",
  "timestamp": "2024-12-13T12:00:00",
  "version": "1.0.0",
  "database": "ok",
  "vector_store": "ok",
  "llm": "ok",
  "uptime_hours": 168.5,
  "memory_usage_mb": 512.3,
  "cpu_usage_percent": 23.5
}

Possible Status Values:

  • healthy: All systems operational
  • degraded: Some components have issues but system is functional
  • unhealthy: One or more critical components are failing

GET /api/analytics/cache-hit-rate

Get detailed cache hit rate statistics.

Request URL:

GET /api/analytics/cache-hit-rate

Response (200 OK):

{
  "overall_cache_hit_rate": 0.623,
  "cache_hits": 1547,
  "cache_misses": 934,
  "total_queries": 2481,
  "by_mode": {
    "hybrid": 0.65,
    "cached": 0.95,
    "realtime": 0.0
  },
  "by_time_range": {
    "last_hour": 0.58,
    "last_24h": 0.62,
    "last_7d": 0.623
  },
  "cache_efficiency": {
    "response_time_saved_ms": 8500000,
    "api_calls_saved": 1547,
    "processing_time_saved_hours": 12.4
  },
  "timestamp": "2024-12-13T12:00:00"
}

POST /api/analytics/rate-query

Submit user rating/feedback for a query result.

Request URL:

POST /api/analytics/rate-query

Request Body:

{
  "query_id": "550e8400-e29b-41d4-a716-446655440000",
  "rating": 5,
  "feedback": "Very helpful answer",
  "tags": ["accurate", "comprehensive"]
}

Request Parameters:

Field Type Required Description
query_id string (UUID) βœ… Yes ID of the query to rate
rating integer βœ… Yes Rating (1-5 stars)
feedback string ❌ No Optional text feedback
tags array[string] ❌ No Optional tags (accurate, helpful, outdated, etc.)

Response (200 OK):

{
  "status": "success",
  "message": "Rating recorded",
  "query_id": "550e8400-e29b-41d4-a716-446655440000"
}

GET /api/analytics/report

Get a comprehensive analytics report.

Request URL:

GET /api/analytics/report?time_range=7d

Query Parameters:

Parameter Type Default Description
time_range string 7d 1d, 7d, 30d, 90d

Response (200 OK):

{
  "time_period": "last 7 days",
  "total_queries": 2481,
  "unique_queries": 1547,
  "cache_statistics": {
    "hit_rate": 0.623,
    "hits": 1547,
    "misses": 934
  },
  "performance": {
    "avg_response_time_ms": 2534.5,
    "median_response_time_ms": 2100.0,
    "p95_response_time_ms": 8500.0
  },
  "documents": {
    "total_indexed": 4283,
    "total_cached": 3156,
    "avg_quality_score": 0.847
  },
  "top_queries": [
    {"text": "python", "count": 156},
    {"text": "machine learning", "count": 143}
  ],
  "timestamp": "2024-12-13T12:00:00"
}

GET /api/analytics/summary

Get quick summary statistics.

Request URL:

GET /api/analytics/summary

Response (200 OK):

{
  "total_queries_processed": 2481,
  "cache_hit_rate": 0.623,
  "avg_response_time_ms": 2534.5,
  "system_status": "healthy",
  "documents_indexed": 4283,
  "vector_store_status": "ok"
}

Health & Status

GET /api/health

Check system health and component status.

Request:

GET /api/health HTTP/1.1
Host: localhost:8000

Response (200 OK):

{
  "status": "healthy",
  "timestamp": "2025-12-10T15:30:45Z",
  "version": "1.0.0",
  "components": {
    "database": {
      "status": "connected",
      "response_time_ms": 2
    },
    "vector_store": {
      "status": "ready",
      "total_vectors": 34215,
      "index_size_mb": 127.8
    },
    "llm": {
      "status": "connected",
      "provider": "openai",
      "model": "gpt-3.5-turbo"
    },
    "search_engine": {
      "status": "operational",
      "engine": "duckduckgo",
      "last_check": "2025-12-10T15:29:45Z"
    },
    "cache": {
      "status": "operational",
      "items": 1547,
      "size_mb": 145.3
    }
  },
  "background_jobs": {
    "cleanup_job": "running",
    "index_maintenance": "running"
  }
}

Status Values:

  • healthy - All systems operational
  • degraded - Some components have issues but system is functional
  • unhealthy - Critical components are down

Error Handling

Error Response Format

All errors follow this standard format:

{
  "detail": "Human-readable error message",
  "error_code": "ERROR_CODE",
  "request_id": "uuid-for-tracking",
  "timestamp": "2025-12-10T15:30:45Z"
}

HTTP Status Codes

Code Meaning Example
200 OK Query successful
400 Bad Request Invalid query format
404 Not Found Query ID doesn't exist
429 Too Many Requests Rate limit exceeded
500 Internal Server Error Unexpected error
503 Service Unavailable System maintenance

Common Errors

// 400: Query too short
{
  "detail": "Query must be at least 3 characters",
  "error_code": "INVALID_QUERY_LENGTH"
}

// 400: Invalid mode
{
  "detail": "Mode must be one of: realtime, cached, hybrid",
  "error_code": "INVALID_MODE"
}

// 429: Rate limited
{
  "detail": "Rate limit exceeded. Max 60 requests per minute.",
  "error_code": "RATE_LIMIT_EXCEEDED"
}

// 503: Vector store not ready
{
  "detail": "Vector store is initializing. Try again in a moment.",
  "error_code": "VECTOR_STORE_INITIALIZING"
}

Rate Limiting

Rate limiting is configurable in .env:

RATE_LIMIT_ENABLED=false          # Set to true for production
RATE_LIMIT_REQUESTS_PER_MINUTE=60
RATE_LIMIT_REQUESTS_PER_HOUR=1000

When rate limited, response includes retry information:

{
  "detail": "Rate limit exceeded",
  "error_code": "RATE_LIMIT_EXCEEDED",
  "retry_after_seconds": 45
}

Examples

Python Client Example

import requests
from datetime import datetime

class OpenWebAnswerClient:
    def __init__(self, base_url="http://localhost:8000"):
        self.base_url = base_url
        self.session = requests.Session()

    def ask(self, query: str, mode: str = "hybrid"):
        """Ask a question"""
        response = self.session.post(
            f"{self.base_url}/api/query",
            json={"query": query, "mode": mode}
        )
        response.raise_for_status()
        return response.json()

    def rate_response(self, query_id: str, rating: int, feedback: str = ""):
        """Rate a query response"""
        response = self.session.post(
            f"{self.base_url}/api/analytics/rate-query",
            json={
                "query_id": query_id,
                "rating": rating,
                "feedback": feedback
            }
        )
        response.raise_for_status()
        return response.json()

    def get_cache_status(self):
        """Get cache statistics"""
        response = self.session.get(
            f"{self.base_url}/api/analytics/cache-status"
        )
        response.raise_for_status()
        return response.json()

# Usage
client = OpenWebAnswerClient()

# Ask a question
result = client.ask("What is machine learning?")
print("Answer:", result["answer"][:100] + "...")
print("Sources:", len(result["sources"]))
print("Cache hit:", result["cache_hit"])
print("Response time:", result["response_time_ms"], "ms")

# Rate the response
client.rate_response(
    query_id=result["query_id"],
    rating=5,
    feedback="Excellent answer!"
)

# Get cache stats
stats = client.get_cache_status()
print("Cache hit rate:", f"{stats['cache_hit_rate']*100:.1f}%")

JavaScript/Node.js Example

async function askQuestion(query, mode = 'hybrid') {
  const response = await fetch('http://localhost:8000/api/query', {
    method: 'POST',
    headers: { 'Content-Type': 'application/json' },
    body: JSON.stringify({ query, mode })
  });

  if (!response.ok) {
    throw new Error(`API error: ${response.statusText}`);
  }

  return response.json();
}

async function rateResponse(queryId, rating, feedback = '') {
  const response = await fetch('http://localhost:8000/api/analytics/rate-query', {
    method: 'POST',
    headers: { 'Content-Type': 'application/json' },
    body: JSON.stringify({ query_id: queryId, rating, feedback })
  });

  return response.json();
}

// Usage
const result = await askQuestion('What is Python?');
console.log('Answer:', result.answer.substring(0, 100) + '...');
console.log('Sources:', result.sources.length);
console.log('Response time:', result.response_time_ms, 'ms');

// Rate it
await rateResponse(result.query_id, 5, 'Great answer!');

cURL Examples

# Ask a question
curl -X POST http://localhost:8000/api/query \
  -H "Content-Type: application/json" \
  -d '{
    "query": "What is quantum computing?",
    "mode": "hybrid"
  }' | jq .

# Rate a response
curl -X POST http://localhost:8000/api/analytics/rate-query \
  -H "Content-Type: application/json" \
  -d '{
    "query_id": "550e8400-e29b-41d4-a716-446655440000",
    "rating": 5
  }'

# Get cache status
curl http://localhost:8000/api/analytics/cache-status | jq .

# Get top queries
curl "http://localhost:8000/api/analytics/top-queries?limit=5" | jq .

# Check health
curl http://localhost:8000/api/health | jq .

API Versioning

Currently on API version 1.0.0. Future versions will be prefixed:

/api/v1/query     # Current (implied)
/api/v2/query     # Future versions

Backward compatibility is maintained within major versions.


Last Updated: December 13, 2025 (Phase 7: Final Verification) API Version: 1.0.0


Examples & Recipes

Recipe 1: Basic Question-Answer

#!/bin/bash

QUESTION="What is the capital of France?"

curl -s -X POST "http://127.0.0.1:8000/api/query" \
  -H "Content-Type: application/json" \
  -d "{\"question\": \"$QUESTION\"}" | jq '.answer'

# Output: The capital of France is Paris...

Recipe 2: Batch Queries

import requests
import json
from concurrent.futures import ThreadPoolExecutor

questions = [
    "What is machine learning?",
    "How does AI work?",
    "What is deep learning?"
]

def query(question):
    response = requests.post(
        "http://127.0.0.1:8000/api/query",
        json={"question": question}
    )
    return response.json()

# Run queries in parallel
with ThreadPoolExecutor(max_workers=3) as executor:
    results = list(executor.map(query, questions))

# Process results
for i, result in enumerate(results):
    print(f"Q: {questions[i]}")
    print(f"A: {result['answer'][:100]}...")
    print(f"Cache Hit: {result['cache_hit']}\n")

Recipe 3: Monitor Performance

import requests
import matplotlib.pyplot as plt
from datetime import datetime, timedelta

# Get performance metrics
response = requests.get("http://127.0.0.1:8000/api/analytics/performance")
metrics = response.json()

print(f"P50: {metrics['response_time_metrics']['median_ms']}ms")
print(f"P95: {metrics['response_time_metrics']['p95_ms']}ms")
print(f"P99: {metrics['response_time_metrics']['p99_ms']}ms")
print(f"Cache Hit Rate: {metrics['cache_metrics']['cache_hit_rate']:.1%}")

# Visualize
times = [50, 100, 200, 500, 1000, 2000, 5000, 10000]
plt.hist(times, bins=20)
plt.xlabel("Response Time (ms)")
plt.ylabel("Frequency")
plt.title("Response Time Distribution")
plt.show()

Recipe 4: Feedback Loop

// Collect user feedback and improve system

const queryAndGetFeedback = async (question) => {
  // Get answer
  const response = await fetch('http://127.0.0.1:8000/api/query', {
    method: 'POST',
    headers: { 'Content-Type': 'application/json' },
    body: JSON.stringify({ question })
  });
  
  const result = await response.json();
  
  // Display answer
  document.getElementById('answer').textContent = result.answer;
  
  // Show feedback buttons
  document.getElementById('feedback').innerHTML = `
    <button onclick="rateFeedback('${result.id}', 5)">πŸ‘ Helpful</button>
    <button onclick="rateFeedback('${result.id}', 1)">πŸ‘Ž Not Helpful</button>
  `;
  
  return result.id;
};

const rateFeedback = async (queryId, rating) => {
  await fetch('http://127.0.0.1:8000/api/analytics/rate-query', {
    method: 'POST',
    headers: { 'Content-Type': 'application/json' },
    body: JSON.stringify({
      query_id: queryId,
      rating: rating,
      feedback: document.getElementById('feedback-text').value
    })
  });
  
  alert('Thank you for your feedback!');
};

OpenAPI / Swagger

Full interactive API documentation available at:


Last Updated: December 13, 2025 (Phase 7: Final Verification)
Version: 1.0.0