1 July 2026 (inference)
Wednesday, July 1, 2026View original
Prompt caching for open-source models in serverless inference chat completions and responses API is now in public preview. Open-source models cache context automatically, so you do not need to set the cache_control or prompt_cache_retention …
Read the full article on the original site.
Read Full Article