Serenity*Edge

Models and pricing

Two families on the same deployments. Prices are in US dollars per million tokens.

Open models

Open-weight models served under their original identifiers.

Orion

Serenity’s own model family. At launch, orion-pro and orion-plus apply Serenity’s serving recipe to the open models in the catalogue. Serenity’s own fine-tuned models will follow.

Launch catalogue and prices per million tokens (USD)
ModelFamilyArchitectureContextMax outputInputOutputCached input
Qwen3.8-27Bqwen3.8-27bOpenDenseNVFP4131,072 tokens32,768 tokens0.1441.500.036
Qwen3.6-35B-A3Bqwen3.6-35b-a3bOpenMoENVFP4131,072 tokens32,768 tokens0.1010.600.025
Orion Proorion-proSame base as qwen3.8-27bOrionDenseNVFP4131,072 tokens32,768 tokens0.1441.500.036
Orion Plusorion-plusSame base as qwen3.6-35b-a3bOrionMoENVFP4131,072 tokens32,768 tokens0.1010.600.025

Capabilities

  • Streaming responses
  • Tool calling
  • Structured outputs
  • Prefix caching
  • NVFP4 precision across all models

Notes

  • Cached input applies to prefix tokens reused across requests.
  • Prices may change. Changes are announced in the documentation.
  • The catalogue will grow. The current list is in the documentation.
See the API reference