Evidence at a glance

Google 2026 10 6 EmbeddingGemma 2Evidence
use Apache 2.0Evidence
768Evidence
parameters 740MEvidence
270M parametersEvidence
170M parametersEvidence

An Embedding Service Is Not a One-Off Dependency

Google released EmbeddingGemma 2 on October 6, 2026. It is an embedding model, not a chat model that directly generates answers. It converts text, images, video, and audio into representations that can be compared for tasks such as search, retrieval, classification, and similarity matching. Google says the previous EmbeddingGemma had been downloaded more than 20 million times, while the new version extends support to multiple modalities.

Beyond the release itself, Simon Willison's Hacker News comment raises a question closer to day-to-day systems operations: what happens to vectors that have already been generated and stored when a model service is no longer available? Embeddings are rarely disposable results produced for a single interaction. They are often calculated in batches, added to an index, and queried repeatedly over time. That makes the model a long-term dependency in the index-production pipeline, and its availability can affect far more than the next request.

A Service Shutdown Can Propagate Through the Index

Vectors record a model's representation of content. If a provider stops offering a particular model, a team may need to use another model to calculate new embeddings and update the associated index. Even if the replacement model is more capable, that does not mean its vectors can simply be substituted for the old model's outputs. For a system with a large content collection, the process can require compute, time, and engineering work, not merely a different API address.

Willison notes that in April 2024, OpenAI said it would cover the financial cost of users re-embedding content with new models. His concern is that teams should not treat this as a commitment every provider will make. This is not a cost estimate for different service contracts, and it does not specify how many resources recalculating vectors would take. It is an architectural reminder: when an index is an important part of a production system, responsibility and budget for a model shutdown should not sit solely in a provider's product plans.

Multimodal Support Expands the Scope of the Dependency

EmbeddingGemma 2 maps text, images, video, and audio into the same 768-dimensional vector space, allowing content from different modalities to enter similarity-based retrieval workflows. Its encoders are modular: the text component has 270 million parameters, the vision encoder 170 million, and the audio encoder 300 million, for a total of 740 million. The text component can run on its own, while the vision and audio encoders can be loaded as needed. By default, video is sampled at one frame per second and passed to the vision encoder.

These design choices let teams combine components according to their tasks and give a single retrieval system the potential to cover different kinds of content. But once an index expands from text to images, video, and audio, a change in model availability may affect more than a text corpus. A shared vector space describes how the model represents inputs. It does not mean every application must adopt every modality at once. Deployments still need to load components according to the data they use and account for how indexes for different content are maintained.

Shorter Vectors Save Storage, with a Quality Trade-Off

The model supports Matryoshka Representation Learning, which allows generated vectors to be shortened to fewer dimensions in order to reduce storage requirements. The model card recommends re-normalizing vectors after truncation and warns that using fewer dimensions may reduce quality. This gives teams room to tune the balance between storage and retrieval performance, but it does not remove the decision. How much to truncate, and whether the resulting quality loss is acceptable, depends on the retrieval task.

This capability addresses a different cost from the risk of a service shutdown. Shortening vectors targets storage use, while Apache 2.0 open weights address whether the model can continue to be used outside a hosted service. The former does not protect a team from model migration if a provider changes its offering. The latter does not automatically reduce the storage footprint of an existing index. Technical leaders who group both under the broad label of cost reduction risk overlooking the limits of each.

Open Weights Offer a Takeover Path, Not a Free Migration

EmbeddingGemma 2 is available under the Apache 2.0 license, leaving teams with two different ways to run it. In normal operation, they can use a hosted service and leave deployment and operations to the provider. If the provider stops hosting the model, they can run the open-weight version themselves or look for another provider that can. This is the combination Willison values: he is willing to pay for hosting, but does not want the end of a hosted service to eliminate the possibility of continuing to use the original model.

Taking over the original model is not the same as migrating to a new one. Self-hosting still requires a team to manage deployment and resources, while switching models may mean recalculating existing vectors. A practical response is to track model versions, paths for obtaining weights, and embedding pipelines alongside the index itself. Teams should distinguish two plans: how to keep the original model running if hosting ends, and how to assess index-rebuilding work if they choose a replacement. Open weights provide a continuity buffer, not a guarantee that operations or migration will be free.