Skip to main content

Model Gateway - Embedding

The ModelGatewayEmbeddingService generates vector embeddings from text using any third-party embedding model available through the IBM watsonx.ai Model Gateway (OpenAI, Azure OpenAI, Mistral, and others). Only providers that expose embedding models can be used here. To see what your gateway actually offers, ask ModelGatewayCatalogService, see Catalog.

Setup required: The Model Gateway must be installed and configured by an administrator before use. See Model Gateway Prerequisites.

Quick Start​

ModelGatewayEmbeddingService service = ModelGatewayEmbeddingService.builder()
.baseUrl(CloudRegion.DALLAS)
.apiKey(WATSONX_API_KEY)
.modelId("text-embedding-3-small")
.build();

ModelGatewayEmbeddingResponse response = service.embed("Hello, world!");

List<Float> vector = response.data().get(0).embedding();
// → [0.0023064255, -0.009327292, -0.0028842222, ...]

Service Configuration​

Basic Setup​

ModelGatewayEmbeddingService service = ModelGatewayEmbeddingService.builder()
.baseUrl(CloudRegion.DALLAS)
.apiKey(WATSONX_API_KEY)
.modelId("text-embedding-3-small")
.build();

Builder Parameters​

ParameterTypeRequiredDescription
apiKeyStringConditionalAPI key for IBM Cloud authentication
authenticatorAuthenticatorConditionalCustom authentication (alternative to apiKey)
baseUrlString / CloudRegionYeswatsonx.ai ML endpoint
modelIdStringYesEmbedding model identifier (e.g., "text-embedding-3-small")
timeoutDurationNoRequest timeout (default: 60 seconds)
logRequestsBooleanNoEnable request logging (default: false)
logResponsesBooleanNoEnable response logging (default: false)
httpClientHttpClientNoCustom HTTP client
verifySslBooleanNoSSL certificate verification (default: true)
versionStringNoAPI version override

Either apiKey or authenticator must be provided.

On-premises deployments​

apiKey configures an IBM Cloud authenticator. On IBM watsonx.ai software (on-premises, CP4D) pass a CP4DAuthenticator through authenticator and use your instance URL as the baseUrl. The CloudRegion enum does not apply. See Authentication.

ModelGatewayEmbeddingService service = ModelGatewayEmbeddingService.builder()
.baseUrl("https://cpd.example.com")
.authenticator(
CP4DAuthenticator.builder()
.url("https://cpd.example.com")
.username(CP4D_USERNAME)
.apiKey(CP4D_API_KEY)
.build()
)
.modelId("text-embedding-3-small")
.build();

Generating Embeddings​

Single Input​

ModelGatewayEmbeddingResponse response = service.embed("Hello, world!");

Multiple Inputs​

// varargs
ModelGatewayEmbeddingResponse response = service.embed("Hello", "World", "Goodbye");

// List<String>
ModelGatewayEmbeddingResponse response = service.embed(List.of("Hello", "World"));

With Parameters​

Use ModelGatewayEmbeddingParameters to configure the optional request options:

ModelGatewayEmbeddingParameters parameters = ModelGatewayEmbeddingParameters.builder()
.dimensions(512)
.encodingFormat(EncodingFormat.FLOAT)
.user("user-123")
.build();

ModelGatewayEmbeddingResponse response = service.embed(List.of("Hello, world!"), parameters);

With a Request Object​

ModelGatewayEmbeddingRequest bundles the inputs and the parameters into a single value you can build once and reuse:

ModelGatewayEmbeddingRequest request = ModelGatewayEmbeddingRequest.builder()
.input("Hello, world!")
.parameters(parameters)
.build();

ModelGatewayEmbeddingResponse response = service.embed(request);

Request Fields​

FieldTypeRequiredDescription
inputString... / List<String>YesText to embed (at least one input is required)
parametersModelGatewayEmbeddingParametersNoOptional request options

Embedding Parameters​

ParameterTypeDescription
dimensionsIntegerNumber of dimensions in the output embedding vector
encodingFormatEncodingFormat / StringFormat of the returned embeddings (default: FLOAT)
userStringUnique identifier for the end-user (passed through to the upstream provider)

The encodingFormat parameter accepts either the ModelGatewayEmbeddingParameters.EncodingFormat enum (preferred) or a raw string:

// preferred — using the enum
.encodingFormat(EncodingFormat.FLOAT) // → "float"
.encodingFormat(EncodingFormat.BASE64) // → "base64"

// also accepted — raw string
.encodingFormat("float")
.encodingFormat("base64")
Enum constantString valueWire representation
EncodingFormat.FLOAT"float"A JSON array of numbers
EncodingFormat.BASE64"base64"A Base64 string encoding the same vector as binary float32 values

Either way embedding() returns a List<Float>, so the format you request never changes the code that reads the vector. See Encoding formats.


Response Fields​

FieldTypeDescription
object()StringAlways "list"
model()StringThe model used to generate the embeddings
data()List<Embedding>The list of generated embedding objects
usage()UsageToken usage information

Embedding Fields​

FieldTypeDescription
object()StringAlways "embedding"
index()intPosition of this embedding in the input list
embedding()List<Float>The embedding vector, one Float per dimension
base64()StringThe raw Base64 payload, or null when the "float" format was used

Encoding formats​

embedding() always returns a List<Float>, whichever encodingFormat you requested. With "base64" the SDK decodes the payload for you, so the same code reads the vector in both cases:

List<Float> vector = response.data().get(0).embedding();

The two formats differ only in what travels over the wire, with "base64" sending the vector as binary float32 values instead of JSON numbers, which makes the response smaller. When you need the untouched payload, for example to forward it to another service, base64() gives you the original string:

String base64 = response.data().get(0).base64();

base64() is null for the "float" format, so it also tells you which format the response came back in.

The returned vector is unmodifiable.

Usage Fields​

FieldTypeDescription
promptTokens()intNumber of tokens in the input
totalTokens()intTotal tokens used