Run your own model behind the Fluxion API

Deploy a model you trained or licensed on dedicated GPUs, and call it with the same key, routes and request shape as the catalogue. You choose the hardware and how far it scales; we run it, keep it healthy and show you what it is doing.

  1. 01

    Bring your model

    From a Hugging Face repo, a container, or a fine-tune made here.

  2. 02

    Choose hardware

    Pick the GPU, then how far it scales up and down.

  3. 03

    One API

    Same key, routes and request shape as the catalogue.

  4. 04

    Monitor

    Requests, failures, replicas, GPU time and logs.

01/Bring your model

Bring a model you trained or licensed

Point Fluxion at a Hugging Face repository, a container image of your own, or a model you fine-tuned here. We pull the weights, build the serving image and run a health check before any request reaches it.

  • A Hugging Face repository or your own container
  • Your Fluxion fine-tunes, without exporting anything
  • Health-checked before it takes traffic

02/Choose hardware

Choose the hardware and how it scales

Pick a GPU type and how many each replica gets, then set the fewest and the most replicas to run. The deployment adds replicas when requests start to queue and removes them when traffic falls, down to zero when nothing is running, if you allow it.

  • GPU type and count per replica
  • Minimum and maximum replicas
  • Scale to zero after an idle timeout

03/One API

Call it through the API you already use

A deployment is a model id on your account, served from the same endpoint as the catalogue. Your application keeps its key, its routes and its polling code; moving a request between a catalogue model and your own is a change to one field. Only your account's keys can call it.

  • The same routes, key and request shape
  • Private to your account
  • Catalogue models and your own, side by side

04/Monitor

See what your deployment is doing

Each deployment has a page of its own: requests and failures, replicas as they come and go, the GPU time used, and the logs from every replica in one stream. Usage lands on the same bill as the rest of your account, in dollars.

  • Requests, failures and GPU time
  • Scaling events and live logs
  • One bill with the rest of your usage
Early access

Deployments are in early access

Tell us about the model you want to run and the traffic you expect, and we will get you set up.

Requests go to hello@fluxion-sys.ai. In the meantime, every model in the catalogue is available today.