Install
Build and serve custom generative AI model with Friendli Endpoints, saving GPU costs and accelerating AI inference. FriendliAI offers best inference solutions to optimize LLM.
- 18articles · 90d
- 5+ day agolatest article
- Jun 24, 2026earliest in window
- 78%with images
- 286avg words
- Computers & Electronics 18
- Science & Technology 18
- Software Dev. 18
Please confirm you are human
This browser or connection looks automated. Press and continuously hold the control for 3 seconds to enable Google-hosted web results and, when separately allowed, AI-assisted answers.
A successful check enables 100 search requests. Interactive access does not authorize scraping, systematic collection, or reuse of search output.
News
Qwen3.8-27B-AWQ-INT4 API & Inference Endpoint
1+ week, 15+ hour ago (503+ words) GLM-5.3 is live. Run Z.ai's latest model on Friendli Model APIs. Try it today ➜ Run this model inference on single tenant GPU with unmatched speed and reliability at scale. SOC 2® Type II For streamlined integration, we recommend using Qwen3.8 via APIs....
MaralGPT-Mythos-9B-2606 API & Inference Endpoint
2+ mon, 1+ week ago (175+ words) GLM-5.2 is live. #1 throughput on OpenRouter, pay-per-token on FriendliAI. Try it today ➜ Run this model inference on single tenant GPU with unmatched speed and reliability at scale. Talk with our engineer to get a quote for reserved GPU instances with…...
Kimi-K2.7-Code API & Inference Endpoint
2+ mon, 1+ week ago (352+ words) GLM-5.2 is live. #1 throughput on OpenRouter, pay-per-token on FriendliAI. Try it today ➜ Run this model inference on single tenant GPU with unmatched speed and reliability at scale. Run this model inference with full control and performance in your environment. Talk…...
zai-org/GLM-5.2 API & Inference Endpoint
2+ mon, 1+ week ago (153+ words) GLM-5.2 is live. #1 throughput on OpenRouter, pay-per-token on FriendliAI. Try it today ➜ Run this model inference with a simple API call. Run this model inference on single tenant GPU with unmatched speed and reliability at scale. Run this model inference…...
moonshotai/Kimi-K2.7-Code API & Inference Endpoint
2+ mon, 2+ week ago (350+ words) ⚡ Hit your SLA, cut costs. Download the Friendli Guide to Inference Performance Optimization ➜ Run this model inference on single tenant GPU with unmatched speed and reliability at scale. Run this model inference with full control and performance in your environment....