---
title: "Image Generation — AVCodex Docs"
description: "Image Generation — AVCodex documentation for AV integrators, programmers, and ops teams."
lang: en
json-ld:
---

[](/)

Solutions

[Pricing](/pricing)[The Signal](/blog)[Resources](/resources)

Learn

[Free AI Assessment](/scorecard)[Get Started →](/pricing)

[Documentation Home](/docs)

Guides 

Getting Started

-   [The Alchemist Copilot](/docs/guides/the-alchemist-copilot)
-   [Choosing a Model](/docs/guides/choosing-a-model)
-   [Skills & Templates](/docs/guides/skills-and-templates)
-   [Pricing & Usage](/docs/guides/pricing-and-usage)
-   [Understanding Tokens](/docs/guides/understanding-tokens)
-   [Maximize AVCodex Capabilities](/docs/guides/maximize-avcodex-capabilities)

Knowledge & Memory

-   [How Knowledge Sources Work](/docs/guides/how-knowledge-sources-work)
-   [Knowledge Retrieval Settings](/docs/guides/knowledge-retrieval-settings)
-   [User Memory](/docs/guides/user-memory)
-   [Consumer Brain](/docs/guides/consumer-brain)

Agent Capabilities

-   [Image Recognition](/docs/guides/image-recognition)
-   [Image Generation](/docs/guides/image-generation)
-   [Video Generation](/docs/guides/video-generation)
-   [Deep Research and Deep Thinking](/docs/guides/deep-research-and-deep-thinking)
-   [Heartbeat (Proactive AI Outreach)](/docs/guides/heartbeat-proactive-ai-outreach)
-   [Database Connections](/docs/guides/database-connections)
-   [Agent-to-Agent Links](/docs/guides/agent-to-agent-links)
-   [Message Tagging](/docs/guides/message-tagging)
-   [Lead Generation Forms](/docs/guides/lead-generation-forms)
-   [Multilingual Apps](/docs/guides/multilingual-apps)
-   [Understanding Evaluations](/docs/guides/understanding-evaluations)

Design & Experience

-   [Style Studio](/docs/guides/style-studio)
-   [Component Studio](/docs/guides/component-studio)
-   [HQ Profile](/docs/guides/hq-profile)
-   [Multiplayer Chat](/docs/guides/multiplayer-chat)
-   [Circles](/docs/guides/circles)
-   [Desktop Agent](/docs/guides/desktop-agent)

Voice & Phone

-   [Phone Numbers](/docs/guides/phone-numbers)
-   [Outbound Calling](/docs/guides/outbound-calling)
-   [Voice Cloning](/docs/guides/voice-cloning)

Publish & Share

-   [Embed Chat Widget](/docs/guides/embed-chat-widget)
-   [Custom Domains](/docs/guides/custom-domains)
-   [PWA Installation](/docs/guides/pwa-installation)
-   [AVCodex Sites](/docs/guides/avcodex-sites)
-   [Embed on Kajabi](/docs/guides/embed-on-kajabi)
-   [How to Use AVCodex with Claude Code](/docs/guides/how-to-use-avcodex-with-claude-code)

Monetization & Access

-   [Selling Access](/docs/guides/selling-access)
-   [Consumer Monetization](/docs/guides/consumer-monetization)
-   [Access Control](/docs/guides/access-control)
-   [Bring Your Own Auth](/docs/guides/bring-your-own-auth)
-   [Clever SSO for Schools](/docs/guides/clever-sso-for-schools)

Analytics & Operations

-   [Analytics & Chat History](/docs/guides/analytics-and-chat-history)
-   [Performance Dashboard](/docs/guides/performance-dashboard)
-   [Programmatic Usage Stats](/docs/guides/programmatic-usage-stats)
-   [Session Lifecycle Webhooks](/docs/guides/session-lifecycle-webhooks)
-   [Audit Logs](/docs/guides/audit-logs)

Teams & White-Label

-   [Team Management](/docs/guides/team-management)
-   [Enterprise Whitelabel](/docs/guides/enterprise-whitelabel)

Alchemist Platform

-   [Alchemist Tickets](/docs/guides/alchemist-tickets)
-   [Alchemist Getting Started](/docs/guides/alchemist-getting-started)
-   [Alchemist Working with Tickets](/docs/guides/alchemist-working-with-tickets)
-   [Alchemist Local Development](/docs/guides/alchemist-local-development)

Alchemist Operations

-   [Alchemist Environment Variables](/docs/guides/alchemist-environment-variables)
-   [Alchemist Deploys and Domains](/docs/guides/alchemist-deploys-and-domains)
-   [Alchemist Self-Healing](/docs/guides/alchemist-self-healing)

Alchemist API & Automation

-   [Alchemist API Keys](/docs/guides/alchemist-api-keys)
-   [Alchemist MCP Server](/docs/guides/alchemist-mcp-server)
-   [Alchemist Pipeline Configuration](/docs/guides/alchemist-pipeline-configuration)
-   [Alchemist Pipeline Permutations](/docs/guides/alchemist-pipeline-permutations)

Developer Platform

-   [Building Custom MCP Servers](/docs/guides/building-custom-mcp-servers)
-   [Consumer OAuth for Custom MCP Servers](/docs/guides/consumer-oauth-for-custom-mcp-servers)

AVCodex MCP Server

-   [Overview](/docs/guides/overview)
-   [MCP Reference](/docs/guides/mcp-reference)
-   [Setup & Installation](/docs/guides/setup-and-installation)
-   [Authentication](/docs/guides/authentication)
-   [Tools Reference](/docs/guides/tools-reference)
-   [Common Workflows](/docs/guides/common-workflows)
-   [Rate Limits](/docs/guides/rate-limits)

Custom Actions 

Pro Actions 

API 

Builder API 

Agentic Commerce (ACP) 

Integrations 

[Docs](/docs)/ Guides / Agent Capabilities 

# Image Generation

Last updated · MAR 2026 · [Read as Markdown](/docs/guides/image-generation.md)

AVCodex agents can generate and edit images directly in the chat conversation. Users describe what they want and the agent creates it. No separate tools or plugins required.

## [Available Models# ](#available-models)

AVCodex supports five image generation models across four providers. Each has different strengths, costs, and capabilities.

Model

Provider

Best For

Cost per Image

Max Resolution

**Gemini 3 Pro**

Google

Highest quality, text rendering, multi-reference editing

~$0.17

4K

**GPT Image 1.5**

OpenAI

Precise instruction-following edits

~$0.04

4096x4096

**FLUX.1 Kontext Pro**

Black Forest Labs

Conversational editing, preserving the original image

~$0.05

1024x1024

**Stability AI SD 3.5**

Stability AI

Inpainting, search-and-replace, background removal

~$0.05

1024x1024

**Gemini 2.5 Flash**

Google

Fast, affordable, high-volume generation

~$0.003

Dynamic

> **Note:** GPT Image 1.5 is the default for new agents. Gemini 3 Pro is the recommended premium option when you need the highest quality output (rendering equipment labels in a rack diagram, for example, where text fidelity matters).

### [Model Capabilities# ](#model-capabilities)

Not every model supports every operation:

-   **Generation**: Create images from text descriptions. All models support this.
-   **Editing**: Modify an existing image based on instructions. Supported by all models, though Gemini 2.5 Flash has limited editing.
-   **Inpainting**: Fill in or replace specific regions of an image. Supported by Gemini 3 Pro and Stability AI.
-   **Blending**: Combine multiple reference images into a new composition. Supported by Gemini 3 Pro and Gemini 2.5 Flash.

## [Enabling Image Generation# ](#enabling-image-generation)

Image generation is enabled by default for all agents on all tiers. To choose a model or adjust settings:

**Open Build Settings.**

Open your agent in the Builder and click the **Build** tab.

**Find the Image Model Section.**

Scroll to the **Image Model** card. You'll see a grid of available models.

**Select a Model.**

Click the model you want to use. The selected model is highlighted with your brand color.

**Adjust Model Settings.**

Each model exposes different configuration options. Adjust them below the model grid after selecting a model.

**Save.**

Click **Save** at the top of the Build tab. The new model is used for all future image-generation requests.

## [Model Settings# ](#model-settings)

Each model offers different configuration parameters. Users don't see these. They're builder-level defaults applied to every generation.

### [GPT Image 1.5 (OpenAI)# ](#gpt-image-1-5-openai)

-   **Quality**: Low (fastest, cheapest), Medium (balanced, default), or High (best quality). Higher quality costs more tokens.
-   **Size**: 1024x1024 (square), 1536x1024 (landscape), or 1024x1536 (portrait).
-   **Background**: Auto (model decides), Transparent (PNG only, useful for equipment icons or signal-flow markers), or Opaque (solid background).

### [Gemini 3 Pro# ](#gemini-3-pro)

-   **Aspect Ratio**: 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, or 2:3. The model adapts composition to fit.

### [FLUX.1 Kontext Pro# ](#flux-1-kontext-pro)

-   **Aspect Ratio**: Same options as Gemini 3 Pro.
-   **Prompt Upsampling**: When enabled, the model automatically enhances the user's prompt for better results.

### [Stability AI SD 3.5# ](#stability-ai-sd-3-5)

-   **Aspect Ratio**: Similar options. Only applies to new generations. Edits preserve original dimensions.
-   **Negative Prompt**: Describe elements to exclude (e.g. "blurry, low quality, watermark"). Steers the model away from unwanted content.
-   **Edit Strength**: Slider from 0.1 to 1.0. Low values make subtle tweaks. High values allow dramatic transformations. Default 0.7.

## [How It Works for Users# ](#how-it-works-for-users)

Users interact with image generation through natural language. They don't need to know which model is configured. They just ask.

### [Generating New Images# ](#generating-new-images)

Users describe what they want and the agent calls the `generateImage` tool automatically:

-   "Create a high-level signal-flow diagram for a typical 12-person huddle room with a Logitech Rally Bar and a Q-SYS Core."
-   "Generate a mockup of a corporate boardroom with dual 98-inch displays and a center table mic array."
-   "Draw an icon set for a control system status dashboard: green check, yellow warning, red error."

The generated image appears inline in the chat conversation.

### [Editing Existing Images# ](#editing-existing-images)

Users can upload an image and ask for changes. The agent uses the uploaded image as a reference:

-   "Mark up this rack elevation. Highlight the DSP and label it." (after uploading a CAD rack drawing).
-   "Remove the background from this product shot of the QSC Core 110f." (after uploading a manufacturer photo).
-   "Add a callout pointing to the input gain knob in this photo of the Yamaha mixer."

There's also a dedicated `editImage` tool that automatically resolves the most recently uploaded image, so users can say "change the color to red" without re-uploading.

### [Multi-Image Blending# ](#multi-image-blending)

With Gemini models, users can upload multiple reference images and ask the agent to combine them:

-   "Blend these two photos into a single composition."
-   "Take the rack from image 1 and put it in the room from image 2."

> **Tip:** Gemini 3 Pro supports up to 14 reference images in a single request, which makes it the best choice for complex multi-image workflows (combining a room photo, a CAD drawing, and a brand style sheet, for example).

## [Disabling Image Generation# ](#disabling-image-generation)

If your agent doesn't need image generation, disable it in the Build tab. When disabled, the `generateImage` and `editImage` tools are removed from the agent's toolset entirely. Users won't be able to request images.

## [Billing# ](#billing)

Image generation is billed per image through your organization's Stripe Token Billing balance. The cost per image depends on the model selected (see the table above). Costs shown in the Builder include a 30% platform markup over the raw provider cost.

\*AVCodex · Your AV expertise. Amplified by AI.\*

Was this helpful? 

[Edit this page →](#)

[

Previous

Image Recognition

](/docs/guides/image-recognition)[

Next

Video Generation

](/docs/guides/video-generation)

On this page

-   [Available Models](#available-models)
-   [Model Capabilities](#model-capabilities)
-   [Enabling Image Generation](#enabling-image-generation)
-   [Model Settings](#model-settings)
-   [GPT Image 1.5 (OpenAI)](#gpt-image-1-5-openai)
-   [Gemini 3 Pro](#gemini-3-pro)
-   [FLUX.1 Kontext Pro](#flux-1-kontext-pro)
-   [Stability AI SD 3.5](#stability-ai-sd-3-5)
-   [How It Works for Users](#how-it-works-for-users)
-   [Generating New Images](#generating-new-images)
-   [Editing Existing Images](#editing-existing-images)
-   [Multi-Image Blending](#multi-image-blending)
-   [Disabling Image Generation](#disabling-image-generation)
-   [Billing](#billing)

[](/)

The AI platform built exclusively for professional AV. Build, deploy, and sell AI tools that understand your industry.

### Platform

-   What You Can Build
-   Templates
-   [Pricing](/pricing)

### Services

-   [Done-For-You](/pricing)
-   [Academy](/academy)
-   [Contact](/contact)

### Company

-   About
-   [The Signal](/blog)
-   [Docs](/docs)
-   [LinkedIn](#)

© 2026 AVCodex. A Future Ready Holdings Inc. product. SOC 2 Type II Certified · HIPAA Compliant