Skip to main content
POST
Image Editing / Multi-image Fusion / Batch Sequence
One endpoint, multiple modes: Seedream has no separate /v1/images/edits endpoint. Editing, multi-image fusion, and batch sequence all run through POST /v1/images/generations. This page’s Playground hits the same endpoint as Text-to-Image — the only difference is the image and sequential_image_generation parameters in the body.
Modes:
  • Single-image editing — image: ["url"] + sequential_image_generation: "disabled"
  • Multi-image fusion — image: ["url1", "url2", ...] + disabled
  • Batch sequence — sequential_image_generation: "auto" + sequential_image_generation_options.max_images: N
  • Image-to-sequence — combine the two: image array + auto + max_images
🖥️ Browser Playground limitation (b64_json mode only)In the default response_format: "url" mode, the Playground works fine (the response is just a temporary BytePlus TOS link). If you switch to response_format: "b64_json", the response contains a multi-MB base64 string and the browser Playground may show 请求时发生错误: unable to complete request — the request actually succeeded; the browser just can’t render such a long base64 string.Recommended workflow:
  • Just want to view the image? Keep the default url mode — the Playground returns the link directly (remember to download to your own storage within 24 hours).
  • Need b64_json? Copy the code sample below and run it locally — the code will decode and save the image to a file automatically.
⚠️ Key differences vs. OpenAI gpt-image-2 editing
  • No multipart/form-data uploads — upload your images to OSS or a public image host first, then pass URLs in the image array
  • image is a URL array, not a repeated image[] field (unlike OpenAI’s multipart/form-data format)
  • No mask field — Seedream does not support alpha-channel mask inpainting; the whole image is rewritten by the prompt
  • Hard limit on total count: input references + output ≤ 15 images
📎 Multi-image order mattersThe order of URLs in the image array becomes the “image 1 / image 2 / image 3” referenced in the prompt. Make ordering explicit:
Replace the clothing in image 1 with the outfit from image 2, keeping the lighting from image 3.
English prompts work best (the model is trained primarily on English), but Chinese is also supported as long as the wording is unambiguous.

Code Examples

About extra_body (important — don’t be misled into thinking it’s an extra nesting layer)image, sequential_image_generation, and watermark are not standard parameters of the OpenAI SDK’s images.generate(), so in the Python SDK you must put them inside extra_body to send them.But extra_body is just the SDK’s parameter container — its fields are flattened and merged into the top level of the request body, at the same level as model and prompt. The JSON that actually goes out is identical to the cURL example below (image sits at the top level); there is no real "extra_body": {...} nesting in the request.If you’re not using the OpenAI SDK and instead build the JSON directly (requests / fetch / etc.), do not write extra_body — just place image and the other fields at the same level as model.

Python (OpenAI SDK · single-image editing)

Python (OpenAI SDK · multi-image fusion)

Python (OpenAI SDK · batch sequence)

cURL (multi-image fusion)

Node.js (fetch · batch sequence)

Pass reference images as public URLs, not base64 (emphasized since 2026-09-11)When the entries in the image array are URLs, the request body is only a few KB and BytePlus downloads your images directly from its Singapore region (verified: the fetching IP belongs to BytePlus, and the APIYI gateway is not involved). When the entries are data:image/...;base64,... strings, a few high-resolution reference images easily push the body to 20 to 30 MB. That body must first be uploaded in full to the APIYI gateway and then forwarded to Singapore. When the cross-border upload is slow, it exceeds the provider ingress limit of 600 seconds for reading the request body and fails with 400 Error when parsing request. The whole request fails and does not fall back to URL mode, because the gateway does not convert base64 into a downloadable link for you.
  • ✅ Host the images on your own object storage or image host and pass public, unauthenticated URLs that can be fetched with a plain GET; keep each image under about 10 MB
  • ⚠️ If base64 is unavoidable, compress first: long edge within 2048 px, re-encode at quality 0.9, and keep the total under 6 MB for multiple images
  • ❌ Do not send base64 bodies above 20 MB, and do not pass private-network addresses or links that require login (BytePlus returns InvalidParameter: Error while downloading when the fetch fails)

Parameter Reference

Count Constraints in Multi-image and Sequence Modes

Iterative refinement: feed a previous output’s URL as the next input, with a fresh edit instruction, to refine progressively. Each round is billed per image — watch cumulative cost.

5.0 Flash Layer Separation (Transparent Layers)

seedream-5-0-flash-260915 can split a finished image into a background base plus separate elements on transparent layers, useful for re-layout, background swaps and asset reuse. Asking for a transparent background in the prompt does not produce an alpha channel, so use this for transparent assets.
Test results (2026-09-24, UTC+8):
  • With 1 reference image, the data array returned 12 images: the first is the background base with the elements removed and filled in (file name ends in _base.png, RGB); the other 11 are RGBA transparent layers (from _layer_1.png), each holding one element cropped to its own size
  • About 78 seconds at 1K; the same request at 2K returned nothing after 15 minutes but was still billed $0.216 for 12 images. Use 1K and set the client timeout to 300 seconds or more
  • Layer separation can fail: the same image once returned a 400 after about 60 seconds (the image content could not be processed for layer decomposition). Failures are not billed, so just retry
  • The areas removed from the base are painted in by the model, so review them before commercial use
Layer separation is billed per output image: each returned image costs $0.018. One request returned 12 images and cost $0.216 (usage.generated_images was 12). The number of layers depends on the image content and cannot be set in advance, and the same image can split differently each time (one image gave 12 and then 14 images, billed $0.216 and $0.252), so budget for a dozen or more.

Response Format

⚠️ The data array length reflects actual output count
  • sequential_image_generation: "disabled" → single-element data
  • sequential_image_generation: "auto" + max_images: N → typically N elements (occasionally fewer if the prompt produces less)
  • Billing is by usage.generated_images, not by max_images
Editing requests are billed identically to text-to-image — per output image. Reference image inputs are not separately billed.

Authorizations

Authorization
string
header
required

API Key obtained from APIYI Console

Body

application/json
model
enum<string>
default:seedream-5-0-260128
required

Model ID

Available options:
seedream-5-0-260128,
seedream-5-0-lite-260128,
seedream-4-5-251128,
seedream-4-0-250828,
seedream-5-0-pro-260628,
seedream-5-0-flash-260915
prompt
string
required

Editing / fusion / sequence instruction. For multi-image scenarios, refer to images explicitly as 'image 1 / image 2'

Example:

"Replace the clothing in image 1 with the outfit from image 2."

image
string<uri>[]

Reference image URL array. Up to 10 images (per official 4.5 / 5.0-pro / 5.0-flash docs; an 11th returns 400 on 5.0-flash in our tests). Note: input + output count ≤ 15

Maximum array length: 10
Example:
sequential_image_generation
enum<string>
default:disabled

Generation mode switch. disabled = single output (default); auto = batch sequence, paired with max_images. Not accepted by 5.0-pro / 5.0-flash — any value returns 400

Available options:
disabled,
auto
sequential_image_generation_options
object

Batch sequence options. Effective only when sequential_image_generation=auto

size
string
default:2K

Output size. Preset tiers (vary by version):

  • 1K (4.0 / 5.0-pro / 5.0-flash) / 1.5K (5.0-flash only) / 2K (all) / 3K (5.0 only) / 4K (4.5, 4.0)

Or exact pixel size WxH, total pixels ∈ [1280×720, 4096×4096], aspect ratio ∈ [1/16, 16]

Example:

"2K"

response_format
enum<string>
default:url
Available options:
url,
b64_json
output_format
enum<string>
default:jpeg

Output format. The 5.0 series supports png/jpeg; 4.5/4.0 only jpeg

Available options:
png,
jpeg
watermark
boolean
default:false
layer_decomposition
boolean
default:false

5.0-flash only: pass 1 reference image to get a background base plus several RGBA transparent layers. Billed per output image (12 in one test); use size 1K and a timeout of 300 seconds or more

stream
boolean
default:false

Streaming output. Recommended for long prompts and multi-image sequence scenarios. Not accepted by 5.0-pro / 5.0-flash (returns 400)

Response

Edited image generated successfully

model
string
Example:

"seedream-5-0-260128"

created
integer
Example:

1768518000

data
object[]

Result array. disabled mode returns 1 element; auto mode typically returns max_images elements (may be fewer)

usage
object

Billed by generated_images actual count, NOT by max_images