Image Generation Models
Image generation models let Xagent create and modify images from text prompts. They power the generate_image and edit_image tools.
Generation vs. Understanding
Image generation models create images. Analysing or reading existing images (OCR, chart analysis) is handled by Vision LLMs, not image generation models.
Supported Providers
| Provider | Models | Abilities |
|---|---|---|
| OpenAI & compatible | gpt-image-1 and compatible | generate + edit |
| DashScope (Alibaba Cloud) | qwen-image | generate (can support edit) |
| Gemini (Google) | gemini-3-pro-preview-image | generate only |
| Xinference | Stable Diffusion variants | configurable (default generate, supports edit) |
Editing support
Not every model supports editing. Gemini image models support generation only — use generate_image. DashScope editing accepts JPG/PNG/BMP/TIFF/WEBP up to 10MB.
Model availability follows declared abilities
Whether generate_image or edit_image shows up as usable depends on whether at least one configured model declares that ability (generate and/or edit, set in step 3 below). If no configured model declares edit, the tools list explains why edit_image is unavailable instead of silently omitting it, and the agent gets an actionable error naming the configured models and their declared abilities rather than a generic failure.
Transparent backgrounds (gpt-image only)
The gpt-image-1 model family can return a PNG with a real alpha channel instead of an opaque background — pass transparent_background to generate_image or edit_image. Asking for a transparent background in the prompt text alone does not work; the parameter is required, and the call fails on any model that doesn't support it. Gemini, DashScope, and Xinference models cannot produce transparency at all.
Configuration
- Add an image generation provider under Models and enter credentials.
- Configure parameters: image size (e.g. 1024×1024), response format (
urlorb64_json), and number of images. - Set the model's abilities — generate, edit, or both.
Usage Examples
User: "Create a promotional poster for a coffee shop"
Xagent: [uses generate_image] -> saves image to workspace
User: "Change the color scheme to warm tones"
Xagent: [uses edit_image] -> modifies existing imageSecurity & Privacy
Generated images are saved to the task workspace. Review your provider's content policy and data-retention terms, and confirm you have rights to the generated content.