A comment on Hacker News highlights that EmbeddingGemma 2 being released under the Apache 2.0 license offers practical benefits for applications that rely on embeddings. The commenter argues that closed, hosted‑only models create long‑term risk because a vendor may discontinue a model, forcing expensive re‑embedding of stored vectors.
Why the license matters
The commenter points out that typical embedding use cases compute and retain thousands or even millions of embedding vectors for later comparisons. If the model is proprietary and only available as a hosted service, a vendor could eventually stop offering it or replace it with a different model. That scenario would usually require re‑computing existing stored vectors to match the new model, which can incur substantial time and financial costs.
Example and caveat
The comment notes that in April 2024, OpenAI offered to cover the financial cost for users re‑embedding content when moving to certain new models. However, the author cautions that such compensation is not a guarantee from every provider and therefore cannot be relied upon universally.
Preference for hosted service with an escape hatch
Crucially, the commenter does not advocate for self‑hosting as the preferred workflow. Instead, they would rather pay for a hosted model while retaining the option to run the open weights themselves if the provider stops hosting. The Apache 2.0 license enables that fallback: users can run the model locally or engage another vendor to do so if needed.
Implications for applications
- Reliability: an open license reduces single‑vendor lock‑in because model weights remain accessible if hosting ends.
- Cost control: the potential need to re‑embed large volumes of stored vectors can be mitigated if organizations can continue to run the original model.
- Flexibility: developers can continue using hosted services for convenience but keep the legal and technical ability to migrate or self‑host.
Conclusion
According to the Hacker News comment, EmbeddingGemma 2’s Apache 2.0 licensing is a meaningful advantage for systems that depend on large‑scale embeddings. It lowers the operational and financial risks associated with vendor‑hosted, proprietary models while preserving the option to use hosted services for day‑to‑day operations.



