Skip to content

About

Image-to-world skills for GPT-powered Codex: 3D environments, object meshes, and SFX.

Resources

Stars

0 stars

Watchers

0 watching

Forks

 
 

Repository files navigation

image-blaster for Codex / GPT

Turn an image into an explorable 3D environment, standalone object meshes, and sound effects with GPT-powered Codex skills, World Labs, and FAL.

This is the Realsee Developer fork of neilsonnn/image-blaster, licensed under MIT. This fork targets Codex CLI and the Codex desktop app.

Quickstart

  1. Install Node.js 20+ and Bun. For terminal use, install Codex CLI with npm install -g @openai/codex and sign in. Alternatively, open this repository as a project in the Codex desktop app.

  2. Clone and install:

    git clone https://github.com/realsee-developer/image-blaster.git
    cd image-blaster
    bun install --frozen-lockfile
    cp .env.example .env
  3. Add WORLD_LABS_API_KEY from World Labs and FAL_KEY from FAL to the local .env, or set them in your environment. Keep keys out of chat and Git. Image analysis and the viewer do not need provider keys; generation uses the providers' paid APIs.

  4. Run bun run check:setup. Install ffmpeg and ffprobe if you want object sound trimming/normalization.

  5. Put an image in input/, then launch codex from the repository root (or use the desktop project). Ask:

    Use $image-blast-project to create a project named room-demo from input/.
    Analyze the image, show the object candidates, and wait for my selection
    before paid generation.
    

    For an authorized full workflow:

    Blast input/room.png as room-demo. Generate all suitable standalone objects,
    a clean plate, the static world, ambience, and object impact sounds.
    

Codex discovers repository skills in .agents/skills/ and project instructions in AGENTS.md. Select a vision-capable GPT model in your Codex host. Ordinary ChatGPT chat without repository filesystem and shell access cannot run this workflow. There is no separate OpenAI API key requirement when using signed-in Codex; FAL and World Labs keys are still required for their respective generation steps.

Workflow and outputs

Skill Purpose
image-blast-project Stage images, inspect project state, manage local files
image-blast-uncover Analyze visible content and identify standalone objects
image-blast-plate Remove selected objects from the background image
image-blast-world Generate/resume a World Labs static environment
image-blast-3d Generate one object mesh and its reference image
image-blast-sfx Generate ambient loops or object sound effects
image-blast-image-edit Edit an image through FAL
image-blast-wildcard Discover and execute another explicitly selected FAL model

Projects live in worlds/<slug>/. The normal outputs include .glb / .obj object meshes, .spz Gaussian splats, a world collider mesh, images, and .mp3 sounds, depending on provider results. Indexed files and hidden request metadata preserve request history and support recovery. Generation is synchronous at the script level; Codex waits on the same terminal session rather than submitting duplicates.

Viewer

Run bun run dev and open the URL printed by Vite. The viewer uses React and Three.js with Spark for splats and Rapier for physics. Generated assets load from local files. Its terminal button opens codex at the repository root on macOS; on other platforms or without the CLI, launch Codex yourself or use the desktop project. The terminal button is available in the development server only.

Generation options

GPT performs source-image analysis and orchestrates the skills. Asset generation continues to use the upstream providers:

  • World Labs marble-1.1 for environments.
  • nano-banana for image edits by default; --provider gpt-image-2 selects the existing FAL-backed GPT image editor for edit/plate workflows.
  • Hunyuan 3D for objects by default; Meshy when explicitly requested.
  • ElevenLabs SFX through FAL for sounds.

Hunyuan options: --face-count 40000-1500000 (default 50000), --enable-pbr true|false (default true), --generate-type Normal|LowPoly|Geometry (default Normal), and --polygon-type triangle|quadrilateral for LowPoly (default triangle).

Assets can be imported into Unity, Unreal, Godot, Blender, or other compatible 3D applications. Quality, generation time, and fees depend on the input and providers.

Development and migration

bun run test
bun run typecheck
bun run build

Tests run offline without provider keys. They do not verify real paid generation. The original React viewer and asset contracts are retained. Claude-specific skill frontmatter, argument substitution, subagent definitions, and hooks were replaced with Codex skills, AGENTS.md, and an explicit setup check. Executable helpers now live in scripts/; update any old .claude/scripts/ commands to that path.

See OpenAI's skill documentation for repository discovery and $skill-name invocation. For the original Claude version, use the upstream repository.

About

Image-to-world skills for GPT-powered Codex: 3D environments, object meshes, and SFX.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages