Skip to content

Warning

🚧 This repository is under active development. Watch the repo, monitor branches and issues, and check the Changelog for the latest updates.

🧭 Navigation:
πŸ”΅ Home | Vision LLM Theory | UI | Deployment | CDK Stacks | Runtime | S3 Files | Lambda Specialists | Prompting System


🦑 BADGERS v4.0 as of August 2026

Broad Agentic Document Generative Extraction & Recognition System

BADGERS transforms document processing through vision-enabled AI and deep layout analysis. Unlike traditional text extraction tools, BADGERS understands document structure and meaning by recognizing visual hierarchies, reading patterns, and contextual relationships between elements.

πŸ€” Why BADGERS?

Traditional document processing tools extract text but lose context. They can't distinguish a header from body text, understand table relationships, or recognize that a diagram explains the adjacent paragraph. BADGERS solves this by:

  • πŸ—οΈ Preserving semantic structure - Maintains document hierarchy and element relationships
  • πŸ‘οΈ Understanding visual context - Recognizes how layout conveys meaning
  • πŸ“š Processing diverse content - Handles 21+ element types from handwriting to equations
  • πŸ€– Automating complex workflows - Orchestrates multiple specialized specialists via an AI agent

Use cases: research acceleration, compliance automation, content management, accessibility remediation.

πŸ“Έ Screenshots

A single React + Express app is both the testing workbench and the deployment/ops console β€” the same code runs locally via npm run dev or on ECS behind Cognito OIDC. Tabs are role-gated: the Testing row below is visible to all users, while an admin-only Deploy row (Stacks, Specialists, S3 Configs, Deploy Tags) is not pictured here. See the UI README for the full tab and role breakdown.

Testing tabs

Home Chat
Home Chat
Landing view with per-page navigation and the resolved environment (region, runtime ARN, gateway ID, config bucket). Streams messages to the AgentCore Runtime over WebSocket, with extended-thinking blocks, the live gateway tool list, and audit / dynamic-token toggles.
Create Specialist Evaluations
Create Specialist Evaluations
Four-step wizard β€” basic info, generated prompt review, examples, deploy β€” including primary model and two fallbacks. Pages through a session's specialist output and scores accuracy, element identification, and contextual understanding 1–5.
Pricing Observability
Pricing Observability
Basic and advanced Bedrock cost calculators with industry presets, per-model token pricing, and ingestion assumptions. Pulls traces and spans for a session from the CloudWatch aws/spans log group, with token usage and an event timeline.

Deployment CLI

Deployment CLI

The deployment menu tracks the eight ordered steps β€” Lambda layers, foundational infrastructure, prompts/manifests/schemas, specialist Lambdas, gateway, runtime, UI image, UI ECS service β€” per deployment ID and stack suffix. Steps can be run individually, all at once, or resumed from wherever the last run stopped.

βš™οΈ How It Works

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                           AgentCore Runtime                                 β”‚
β”‚   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”‚
β”‚   β”‚  PDF Analysis Agent (Strands)                                       β”‚   β”‚
β”‚   β”‚  - Claude Sonnet 4.5 with Extended Thinking                         β”‚   β”‚
β”‚   β”‚  - Session state management                                         β”‚   β”‚
β”‚   β”‚  - MCP tool orchestration                                           β”‚   β”‚
β”‚   β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                      β”‚
                                      β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                           AgentCore Gateway                                 β”‚
β”‚   - MCP Protocol (2025-03-26)                                               β”‚
β”‚   - Cognito JWT Authentication                                              β”‚
β”‚   - Semantic tool search                                                    β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                      β”‚
                   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                   β”‚                  β”‚                  β”‚
                   β–Ό                  β–Ό                  β–Ό
            β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
            β”‚   Lambda    β”‚    β”‚   Lambda    β”‚    β”‚   Lambda    β”‚
            β”‚ Specialist  β”‚    β”‚ Specialist  β”‚    β”‚ Specialist  β”‚
            β”‚ (26 tools)  β”‚    β”‚             β”‚    β”‚             β”‚
            β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                   β”‚                  β”‚                  β”‚
                   β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                      β–Ό
                               β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                               β”‚   Bedrock   β”‚
                               β”‚   Claude    β”‚
                               β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
  1. πŸ“„ User submits a document with analysis instructions
  2. 🧠 Strands Agent (running in AgentCore Runtime) interprets the request
  3. πŸ”§ Agent selects tools from a library of specialists via MCP Gateway
  4. πŸ“‹ Agent opens a job on the first specialist call, tagging every invocation in the run with the job and document it belongs to
  5. ⚑ Lambda specialists (standardized and domain-specific functions, including container-based) process document elements using Claude vision models, each recording its own outcome
  6. πŸ“Š Results aggregate with preserved structure and semantic relationships

πŸ› οΈ Tech Stack

Component Technology
πŸ€– Agent Framework Strands Agents
🏠 Agent Hosting Amazon Bedrock AgentCore Runtime
πŸšͺ Tool Gateway Amazon Bedrock AgentCore Gateway (MCP Protocol)
🧠 Foundation Model Claude Sonnet 4.5 (via Amazon Bedrock)
⚑ Compute AWS Lambda (modular specialist functions, including container-based)
πŸ“¦ Storage Amazon S3 (configs, prompts, outputs)
πŸ“‹ Job Tracking Amazon DynamoDB (document β†’ job β†’ subtask state)
πŸ–₯️ UI Hosting Amazon ECS Express Gateway service (in a VPC)
πŸ” Auth Amazon Cognito (OIDC + PKCE for the UI, OAuth 2.0 M2M for the Gateway)
πŸ—οΈ IaC AWS CDK (Python)
πŸ“ˆ Observability CloudWatch Logs, X-Ray Transaction Search
πŸ“Š Cost Tracking Bedrock Application Inference Profiles

πŸ”¬ Specialists

Specialist Purpose
πŸ“Έ pdf_to_images_converter Convert PDF pages to images
🏷️ classify_pdf_content Classify document content type
πŸ“ full_text_specialist Extract all text content
πŸ“Š table_specialist Extract and structure tables
πŸ“ˆ charts_specialist Analyze charts and graphs
πŸ”€ diagram_specialist Process diagrams and flowcharts
πŸ“ layout_specialist Document structure analysis
πŸ₯ decision_tree_specialist Medical/clinical document analysis
πŸ”¬ scientific_specialist Scientific paper analysis
✍️ handwriting_specialist Handwritten text recognition
πŸ”’ handwriting_math_specialist Handwritten mathematical notation recognition
πŸ’» code_block_specialist Extract code snippets
πŸ—‚οΈ metadata_generic_specialist Generic metadata extraction
πŸ—‚οΈ metadata_mads_specialist MADS metadata format extraction
πŸ—‚οΈ metadata_mods_specialist MODS metadata format extraction
πŸ”‘ keyword_topic_specialist Extract keywords and topics
πŸ”§ remediation_specialist PDF accessibility remediation (container, content stream tagging + structure tree builder)
πŸ“„ page_specialist Single page content analysis
🧱 elements_specialist Document element detection
🧱 robust_elements_specialist Enhanced element detection with fallbacks
πŸ‘οΈ general_visual_analysis_specialist General-purpose visual content analysis
✏️ editorial_specialist Editorial content and markup analysis
πŸ—ΊοΈ war_map_specialist Historical war map analysis
πŸŽ“ edu_transcript_specialist Educational transcript analysis
πŸ”— correlation_specialist Correlate multi-specialist results per page
πŸ–ΌοΈ image_enhancer Image enhancement and preprocessing

πŸš€ Deployment

Prerequisites

Quick Start

./deploy.sh

That is the whole command. deploy.sh asks which deployment to work on β€” listing anything it finds in .deploy-state/, or offering to start a new one β€” and then presents a menu. It is resumable and every step is idempotent, so re-run it after a failure and it picks up where it stopped.

DEPLOYMENT_ID is not read from the environment. It is always chosen interactively, because a value left exported in your shell silently targets another deployment's stacks.

Pick option 9 for a full deployment, or 12 to run only what is still outstanding. You can jump straight to one option β€” ./deploy.sh 6 β€” and the deployment is still chosen interactively first.

The eight steps:

# Step What it does
1 Lambda Layers foundation, PDF processing, Poppler/qpdf
2 Foundational Infra S3, Cognito, DynamoDB, IAM, ECR, Inference Profiles, X-Ray, Memory, VPC
3 Upload Config prompts, manifests and schemas to the config bucket
4 Specialist Lambdas container images, then the Lambda stack (26 specialists)
5 Gateway AgentCore MCP Gateway, records the Gateway URL
6 Runtime builds and pushes the agent image, then deploys the Runtime
7 UI β€” Build generates ui/.env from Cognito, builds the bundle and image
8 UI β€” Deploy ECS Express Gateway service, forces the rollout, waits for it

Plus 9 full deployment, 12 resume, 10 status, 11 reset state (deletes nothing in AWS), 0 exit.

Step 8 asks once whether the UI should be publicly reachable. That answer is fixed for the life of the VPC β€” see Network exposure.

For the full procedure, prerequisites in depth, and every environment variable, see the Deployment Guide.

Deployment Identity

DEPLOYMENT_ID is a short label you choose β€” lowercase, starting with a letter, 16 characters or fewer. deploy.sh generates a three-character random STACK_SUFFIX once and persists both in .deploy-state/{DEPLOYMENT_ID}.json:

  • Stack names are BADGERS-{Name}-{DEPLOYMENT_ID}-{suffix} β€” for example BADGERS-S3-dev-a1b
  • Resource names carry both parts β€” for example badgers-config-dev-a1b, and SSM parameters under /badgers-dev-a1b/

Because both are unique per deployment, several deployments can coexist in one account and region. Stack names include the deployment id as well as the suffix so each stack is self-describing β€” tooling reads a deployment's identity off the stack name, and a mistyped id matches no stacks instead of resolving someone else's. The state file also tracks which steps completed, which is what makes the script resumable.

Cleanup

./destroy.sh

Like deploy.sh it asks what to tear down, but it discovers deployments from CloudFormation rather than from .deploy-state/ β€” a state file can be deleted while the stacks are still live. You are then required to type the DEPLOYMENT_ID to confirm.

It empties the S3 buckets, deletes the ECS Express service and the AgentCore runtime before the VPC (CloudFormation cannot delete a VPC while any ENI is still attached), sweeps leftover ENIs, destroys every stack in reverse dependency order, verifies they are gone, and only then schedules the KMS key for deletion so its alias is freed for redeployment. A teardown that leaves stacks standing exits non-zero and says so rather than reporting success.

If a VPC stack still gets stuck on a lingering ENI:

DEPLOYMENT_ID=dev STACK_SUFFIX=a1b ./destroy.sh --vpc-cleanup-only

To tear down by hand when the script cannot run, follow Manual Teardown in the Console β€” the stack deletion order matters, and two resources have to be removed before any stack.

πŸ“ Project Structure

β”œβ”€β”€ deployment/
β”‚   β”œβ”€β”€ app.py                 # CDK app entry point
β”‚   β”œβ”€β”€ stacks/                # CDK stack definitions
β”‚   β”œβ”€β”€ lambdas/code/          # Specialist Lambda functions
β”‚   β”œβ”€β”€ runtime/               # AgentCore Runtime container
β”‚   β”œβ”€β”€ s3_files/              # Prompts, schemas, manifests
β”‚   └── badgers-foundation/    # Shared specialist framework
β”œβ”€β”€ ui/                        # BADGERS UI (React + Express, runs locally or deployed via Docker)
β”‚   β”œβ”€β”€ src/                   # React components (testing + admin tabs, role-gated)
β”‚   β”œβ”€β”€ server/                # Express API server (testing + admin routes, OIDC auth)
β”‚   └── Dockerfile             # Container image for AWS deployment
└── pyproject.toml

πŸ” Technical Deep Dive

πŸ“¦ Lambda Layers

BADGERS uses Lambda layers shared across specialist functions:

πŸ—οΈ Foundation Layer (layer.zip)

  • Built via deployment/lambdas/build_foundation_layer.sh
  • Contains the specialist framework (7 Python modules)
  • Includes dependencies: boto3, botocore
  • Includes core system prompts used by all specialists
layer/python/
β”œβ”€β”€ foundation/
β”‚   β”œβ”€β”€ specialist_foundation.py    # 🎯 Main orchestration class
β”‚   β”œβ”€β”€ bedrock_client.py         # πŸ”„ Bedrock API with retry/fallback
β”‚   β”œβ”€β”€ configuration_manager.py  # βš™οΈ Config loading/validation
β”‚   β”œβ”€β”€ image_processor.py        # πŸ–ΌοΈ Image optimization
β”‚   β”œβ”€β”€ message_chain_builder.py  # πŸ’¬ Claude message formatting
β”‚   β”œβ”€β”€ prompt_loader.py          # πŸ“œ Prompt file loading (local/S3)
β”‚   └── response_processor.py     # πŸ“€ Response extraction
β”œβ”€β”€ config/
β”‚   └── config.py
└── prompts/core_system_prompts/
    └── *.xml

πŸ“„ Poppler Layer (poppler-qpdf-layer.zip)

  • PDF rendering library for pdf_to_images_converter
  • Built via deployment/lambdas/build_poppler_qdf_layer.sh

πŸ”¬ How an Specialist Works

Each specialist follows the same pattern using SpecialistFoundation:

# Lambda handler (simplified)
def lambda_handler(event, context):
    # 1️⃣ Load config from S3 manifest
    config = load_manifest_from_s3(bucket, "full_text_specialist")

    # 2️⃣ Initialize foundation with S3-aware prompt loader
    specialist = SpecialistFoundation(...)

    # 3️⃣ Run analysis pipeline
    result = specialist.analyze(image_data)

    # 4️⃣ Save result to S3 and return
    save_result_to_s3(result, session_id)
    return {"result": result}

The analyze() method orchestrates:

  1. πŸ–ΌοΈ Image processing - Resize/optimize for Claude's vision API
  2. πŸ“œ Prompt loading - Combine wrapper + specialist prompts from S3
  3. πŸ’¬ Message building - Format for Bedrock Converse API
  4. ⚑ Dynamic token estimation - Score image complexity and set token budget (when enabled)
  5. πŸ€– Model invocation - Call Claude with retry/fallback logic
  6. βœ… Response processing - Extract and validate result

πŸ“œ Prompting System

Prompts are modular XML files composed at runtime:

s3://config-bucket/
β”œβ”€β”€ core_system_prompts/
β”‚   β”œβ”€β”€ prompt_system_wrapper.xml   # 🎁 Main template with placeholders
β”‚   β”œβ”€β”€ core_rules/rules.xml        # πŸ“ Shared rules for all specialists
β”‚   └── error_handling/*.xml        # ⚠️ Error response templates
β”œβ”€β”€ prompts/{specialist_name}/
β”‚   β”œβ”€β”€ {specialist}_job_role.xml     # πŸ‘€ Role definition
β”‚   β”œβ”€β”€ {specialist}_context.xml      # 🌍 Domain context
β”‚   β”œβ”€β”€ {specialist}_rules.xml        # πŸ“ Specialist-specific rules
β”‚   β”œβ”€β”€ {specialist}_tasks.xml        # βœ… Task instructions
β”‚   └── {specialist}_format.xml       # πŸ“‹ Output format spec
└── wrappers/
    └── prompt_system_wrapper.xml

The PromptLoader composes the final system prompt:

<!-- prompt_system_wrapper.xml -->
<system_prompt>
    {core_rules}           <!-- πŸ“ Injected from core_rules/rules.xml -->
    {composed_prompt}      <!-- 🧩 Injected from specialist prompt files -->
    {error_handler_general}
    {error_handler_not_found}
</system_prompt>

Placeholders like [[PIXEL_WIDTH]] and [[PIXEL_HEIGHT]] are replaced with actual image dimensions at runtime.

βš™οΈ Configuration System

Each specialist has a manifest file in S3:

// s3://config-bucket/manifests/full_text_specialist.json
{
    "tool": {
        "name": "analyze_full_text_tool",
        "description": "Extracts text content maintaining reading order...",
        "inputSchema": {
            "type": "object",
            "properties": {
                "image_path": { "type": "string" },
                "session_id": { "type": "string" },
                "audit_mode": { "type": "boolean" }
            },
            "required": ["image_path", "session_id"]
        }
    },
    "specialist": {
        "name": "full_text_specialist",
        "enhancement_eligible": true,
        "model_selections": {
            "primary": "us.anthropic.claude-sonnet-4-5-20250929-v1:0",
            "fallback_list": [
                "us.anthropic.claude-haiku-4-5-20251001-v1:0",
                "us.amazon.nova-premier-v1:0"
            ]
        },
        "max_retries": 3,
        "prompt_files": [
            "full_text_job_role.xml",
            "full_text_context.xml",
            "full_text_rules.xml",
            "full_text_tasks_extraction.xml",
            "full_text_format.xml"
        ],
        "max_examples": 0,
        "analysis_text": "full text content",
        "expected_output_tokens": 6000,
        "output_extension": "xml"
    }
}

Key configuration features:

  • πŸ”„ Model fallback chain - Primary model with ordered fallbacks
  • πŸ” Retry logic - Configurable retry count per specialist
  • 🧩 Prompt composition - List of XML files to combine
  • πŸ“‹ Tool schema - MCP-compatible input schema for Gateway
  • πŸ–ΌοΈ Enhancement eligible - Flag indicating specialist benefits from image preprocessing (used by image_enhancer tool)

Global settings (from environment or defaults):

{
    "max_tokens": 8000,
    "temperature": 0.1,
    "max_image_size": 20971520,  # 20MB
    "max_dimension": 2048,
    "jpeg_quality": 85,
    "throttle_delay": 1.0,
    "aws_region": "us-west-2"
}

⚑ Dynamic Token Estimation

When enabled, BADGERS estimates the optimal max_tokens per image based on visual complexity, reducing cost on simple documents and avoiding truncation on dense ones. The scorer runs on the already-processed image bytes β€” no extra I/O.

Four metrics are combined into a complexity score: text pixel ratio, grayscale entropy, edge density, and color standard deviation. The score maps to a token budget (8K / 12K / 16K / 24K).

Enabling: Toggle "Dynamic Token Estimation" in the chat UI, or set the Lambda environment variable DYNAMIC_TOKENS_ENABLED=true.

Tuning: Add a dynamic_tokens block to an specialist manifest to customize weights and thresholds:

"dynamic_tokens": {
    "weights": {
        "text_ratio": 0.2,
        "entropy": 0.3,
        "edge_density": 0.3,
        "color_std": 0.2
    },
    "thresholds": [
        {"max_score": 0.20, "max_tokens": 8000},
        {"max_score": 0.30, "max_tokens": 12000},
        {"max_score": 0.45, "max_tokens": 16000},
        {"max_score": 1.00, "max_tokens": 24000}
    ]
}

Observability: When active, logs report the estimated budget, actual token usage, and utilization percentage for calibration.

πŸ“Š Inference Profiles for Cost Tracking

BADGERS uses Application Inference Profiles to enable cost allocation and usage monitoring. The system maps model IDs to profile ARNs at runtime:

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                        Inference Profile Flow                               β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚                                                                             β”‚
β”‚  1. CDK deploys InferenceProfilesStack                                      β”‚
β”‚     └─> Creates ApplicationInferenceProfile for each model                  β”‚
β”‚         β€’ badgers-claude-sonnet-{id}  (US)                               β”‚
β”‚         β€’ badgers-claude-haiku-{id}   (US)                               β”‚
β”‚         β€’ badgers-claude-opus-{id}    (US)                               β”‚
β”‚         β€’ badgers-nova-premier-{id}   (US)                               β”‚
β”‚                                                                             β”‚
β”‚  2. Runtime receives profile ARNs as environment variables                  β”‚
β”‚     └─> CLAUDE_SONNET_PROFILE_ARN, CLAUDE_HAIKU_PROFILE_ARN, etc.           β”‚
β”‚                                                                             β”‚
β”‚  3. At invocation, bedrock_client.py maps model_id β†’ profile ARN            β”‚
β”‚     └─> "us.anthropic.claude-sonnet-4-5-*" β†’ $CLAUDE_SONNET_PROFILE_ARN    β”‚
β”‚                                                                             β”‚
β”‚  4. Bedrock invoked with profile ARN (enables cost tracking)                β”‚
β”‚     └─> Falls back to model ID if no profile configured                     β”‚
β”‚                                                                             β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Model ID to environment variable mapping:

Model Pattern Environment Variable
*claude-sonnet-4-5* CLAUDE_SONNET_PROFILE_ARN
*claude-haiku-4-5* CLAUDE_HAIKU_PROFILE_ARN
*claude-opus-4-6* CLAUDE_OPUS_PROFILE_ARN
*nova-premier* NOVA_PREMIER_PROFILE_ARN

βž• Adding a New Specialist

Option 1: Use the Wizard (Recommended)

cd local_testing
npm run dev

The Specialist Creation Wizard is available as the πŸ§™ Create Specialist tab in the UI.

Option 2: Manual Creation

  1. πŸ“œ Create prompt files in deployment/s3_files/prompts/{specialist_name}/
  2. πŸ“‹ Create manifest in deployment/s3_files/manifests/{specialist_name}.json
  3. πŸ“ Create schema in deployment/s3_files/schemas/{specialist_name}.json
  4. ⚑ Create Lambda code in deployment/lambdas/code/{specialist_name}/lambda_handler.py
  5. πŸ“ Register in deployment/stacks/lambda_stack.py
  6. πŸš€ Redeploy: cdk deploy BADGERS-Lambda-{id}-{suffix} BADGERS-Gateway-{id}-{suffix}

πŸ”§ Troubleshooting

Service Control Policy (SCP) Blocks Cross-Region Inference

If your AWS organization uses strict SCPs that deny cross-region Bedrock operations, you may see:

AccessDeniedException: ... is not authorized to perform: bedrock:InvokeModelWithResponseStream
on resource: arn:aws:bedrock:::foundation-model/anthropic.claude-* with an explicit deny
in a service control policy

BADGERS defaults to regional (us.anthropic.*) inference profiles which avoid cross-region routing. If you previously deployed with global.anthropic.* profiles, redeploy after pulling the latest code.

Marketplace Subscription Error on First Invocation

After a fresh deployment, the first model invocation may fail with:

AccessDeniedException: Model access is denied due to IAM user or service role is not authorized
to perform the required AWS Marketplace actions (aws-marketplace:ViewSubscriptions,
aws-marketplace:Subscribe)

The IAM stack now includes aws-marketplace:ViewSubscriptions and aws-marketplace:Subscribe permissions. If you see this error on an older deployment, redeploy the IAM stack. As a workaround, manually invoke the model once in the Bedrock console playground to trigger the Marketplace subscription.


Notices

Customers are responsible for making their own independent assessment of the information in this Guidance. This Guidance: (a) is for informational purposes only, (b) represents AWS current product offerings and practices, which are subject to change without notice, and (c) does not create any commitments or assurances from AWS and its affiliates, suppliers or licensors. AWS products or services are provided "as is" without warranties, representations, or conditions of any kind, whether express or implied. AWS responsibilities and liabilities to its customers are controlled by AWS agreements, and this Guidance is not part of, nor does it modify, any agreement between AWS and its customers.


Authors

  • Randall Potter

πŸ“– Further Reading

πŸ€– Amazon Bedrock & Foundation Models

πŸš€ Amazon Bedrock AgentCore

⚑ AWS Lambda

πŸ” Amazon Cognito

πŸ“ˆ Observability

πŸ“¦ Amazon S3

πŸ’» Amazon Kiro IDE

About

Guidance on deploying a generative AI document analysis with Amazon Bedrock AgentCore. Auto-classifies, enhances, and aggregates multi-type documents using Gestalt-informed vision prompts. Custom analyzer creation wizard. Scripted CDK deployment. Gradio frontend included.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

15 stars

Watchers

0 watching

Forks

Packages

Used by

Contributors

Languages