Skip to main content

Model Gateway - Image Generation

The ModelGatewayImageService generates images from text prompts using any image model available through the IBM watsonx.ai Model Gateway (DALL-E 3, gpt-image-1, and others). Only providers that expose image models can be used here. To see what your gateway actually offers, ask ModelGatewayCatalogService, see Catalog.

Setup required: The Model Gateway must be installed and configured by an administrator before use. See Model Gateway Prerequisites.

Quick Start​

ModelGatewayImageService service = ModelGatewayImageService.builder()
.baseUrl(CloudRegion.DALLAS)
.apiKey(WATSONX_API_KEY)
.modelId("gpt-image-1")
.build();

ModelGatewayImageResponse response = service.generate("A futuristic city at sunset");
String b64 = response.data().get(0).b64Json();

Service Configuration​

Basic Setup​

ModelGatewayImageService service = ModelGatewayImageService.builder()
.baseUrl(CloudRegion.DALLAS)
.apiKey(WATSONX_API_KEY)
.modelId("gpt-image-1")
.build();

Builder Parameters​

ParameterTypeRequiredDescription
apiKeyStringConditionalAPI key for IBM Cloud authentication
authenticatorAuthenticatorConditionalCustom authentication (alternative to apiKey)
baseUrlString / CloudRegionYeswatsonx.ai ML endpoint
modelIdStringYesImage model identifier (e.g., "gpt-image-1")
timeoutDurationNoRequest timeout (default: 60 seconds)
logRequestsBooleanNoEnable request logging (default: false)
logResponsesBooleanNoEnable response logging (default: false)
httpClientHttpClientNoCustom HTTP client
verifySslBooleanNoSSL certificate verification (default: true)
versionStringNoAPI version override

Either apiKey or authenticator must be provided.

On-premises deployments​

apiKey configures an IBM Cloud authenticator. On IBM watsonx.ai software (on-premises, CP4D) pass a CP4DAuthenticator through authenticator and use your instance URL as the baseUrl. The CloudRegion enum does not apply there. See Authentication.

ModelGatewayImageService service = ModelGatewayImageService.builder()
.baseUrl("https://cpd.example.com")
.authenticator(
CP4DAuthenticator.builder()
.url("https://cpd.example.com")
.username(CP4D_USERNAME)
.apiKey(CP4D_API_KEY)
.build()
)
.modelId("gpt-image-1")
.build();

Generating Images​

From a Prompt String​

ModelGatewayImageResponse response = service.generate("A serene mountain landscape");

With Parameters​

Use ModelGatewayImageParameters to configure the optional request options:

ModelGatewayImageParameters parameters = ModelGatewayImageParameters.builder()
.n(1)
.size(Size.SIZE_1024X1024)
.quality(Quality.HIGH)
.responseFormat(ResponseFormat.B64_JSON)
.style(Style.VIVID)
.outputFormat(OutputFormat.PNG)
.background(Background.TRANSPARENT)
.moderation(Moderation.LOW)
.user("user-123")
.build();

ModelGatewayImageResponse response = service.generate("A serene mountain landscape", parameters);

With a Request Object​

ModelGatewayImageRequest bundles the prompt and the parameters into a single value you can build once and reuse:

ModelGatewayImageRequest request = ModelGatewayImageRequest.builder()
.prompt("A serene mountain landscape")
.parameters(parameters)
.build();

ModelGatewayImageResponse response = service.generate(request);

Request Fields​

FieldTypeRequiredDescription
promptStringYesText description of the desired image.
parametersModelGatewayImageParametersNoOptional request options

Image Parameters​

ParameterTypeDefaultDescription
backgroundBackground / StringautoBackground transparency: transparent, opaque, or auto. Transparency requires an outputFormat that supports it, so png or webp
moderationModeration / StringautoContent moderation level: low for less restrictive filtering, or auto
nInteger1Number of images to generate, from 1 to 10
outputCompressionInteger100Compression level from 0 to 100, for the webp and jpeg formats only
outputFormatOutputFormat / StringjpegFile format: png, jpeg, webp, or auto
partialImagesInteger0Number of partial images streamed before the final result, from 0 to 3. With 0 the image arrives in a single event
qualityQuality / StringautoImage quality: auto, high, medium, low, hd, or standard
responseFormatResponseFormat / StringurlReturn format: url or b64_json
sizeSize / String1024x1024Dimensions of the generated image
styleStyle / StringvividVisual style: vivid for hyper-real and dramatic images, natural for more natural ones
userStringUnique identifier for the end-user, passed through to the upstream provider to help it detect abuse

Response Fields​

FieldTypeDescription
created()longUNIX timestamp in seconds of when the model response was created
data()List<ImageData>Generated image objects, between 1 and 10 of them
background()StringBackground setting used, never auto
outputFormat()StringOutput format used, never auto
quality()StringQuality level used, never auto
size()StringSize used, never auto
usage()UsageToken usage, or null if not returned. On OpenAI only gpt-image-1 reports it

ImageData Fields​

FieldTypeDescription
url()StringImage URL, or null when b64_json format was requested. Unsupported by gpt-image-1
b64Json()StringBase64-encoded image data, or null when url format was used
revisedPrompt()StringRevised prompt, if the model modified it. On OpenAI only dall-e-3 returns it

The returned data() list is unmodifiable.

Response formats​

responseFormat decides which of the two ImageData fields is populated. With b64_json the image bytes travel inline and you decode them yourself:

byte[] image = Base64.getDecoder().decode(response.data().get(0).b64Json());
Files.write(Path.of("image.png"), image);

With url the provider stores the image and returns a link to it. On OpenAI that link stays valid for 60 minutes after generation, so download the image before you need it again:

String url = response.data().get(0).url();

The field you did not request comes back null, so it also tells you which format the response came back in. Not every model honours the setting: OpenAI supports responseFormat only on dall-e-2 and dall-e-3, while gpt-image-1 always returns Base64.

Usage Fields​

FieldTypeDescription
inputTokens()longTokens in the input prompt, images and text together
outputTokens()longOutput tokens generated by the model
totalTokens()longTotal tokens used
inputTokensDetails()InputTokensDetailsBreakdown by token type

InputTokensDetails Fields​

FieldTypeDescription
textTokens()longText tokens in the prompt
imageTokens()longImage tokens in the prompt