Inference ExchangeEarly access
Serve your models on GPUs in Latin America.
Inference Exchange serves large language models on local GPUs in Latin America, close to your users, on hardware EdgeUno runs today. We are onboarding early customers now.
What you get
Local GPUs
Inference runs on GPUs in the region instead of overseas.
Close to your users
Shorter round trips for every request your application makes.
Early access
Talk to sales to join the first customers and shape the offer with us.
Join the early access.
Tell us which models you serve and we will get back to you.