Cài đặt

Plugin

Cerebrium

Deploy AI on serverless GPUs

Cài đặt plugin

Cerebrium runs Python workloads on serverless GPU and CPU with scale to zero and per second billing: REST endpoints, SSE streaming, WebSockets and async jobs, all described by one cerebrium.toml and driven by one CLI. This plugin helps you go from a Python function to a deployed endpoint, and then keep it healthy. It covers choosing hardware, regions and the right runtime, writing and fixing cerebrium.toml, picking a scaling metric that matches the workload, calling the endpoint over REST, streaming, WebSocket or async, handling secrets and environment variables, wiring CI/CD with a service account, and debugging a failed build or an app that is queueing or returning 5xx. It is built for inference APIs for language, embedding and vision models, real time voice and video apps, bursty traffic that should not hold idle GPUs, multi region deployments, and migrations from Replicate, Hugging Face or Mystic.

Kỹ năng

Thông tin

Tính năng
Read, Write
Nhà phát triển
Cerebrium
Danh mục
Other
Trang web
Phiên bản
0.1.0
Chính sách quyền riêng tư
Điều khoản dịch vụ