Skip to content

Scalable LLM Deployment on AWS with Ray Serve and FastAPI

Last updated Jan 8, 2026
On this page

Deploying Large Language Models (LLMs) effectively requires a robust architecture that can handle high concurrency, manage GPU resources efficiently, and scale dynamically. In this post, I’ll walk through a production-ready setup for hosting open-source models (like Llama 3 or Mistral) on AWS using Ray Serve for orchestration and FastAPI as the interface.

The Architecture

The stack consists of:

  • Infrastructure: AWS EC2 instances (g5.xlarge or similar GPU-optimized instances)
  • Orchestration: Ray Cluster (Head node + Worker nodes)
  • Serving: Ray Serve wrapping a FastAPI application
  • Model Engine: vLLM for high-throughput inference

Why Ray Serve?

Ray Serve excels at “model composition” and scaling. Unlike simple Docker containers, Ray allows us to:

  1. Scale independently: Scale the model replicas separately from the API handling logic.
  2. Batching: Native support for dynamic request batching to maximize GPU utilization.
  3. Pipeline composition: Easily chain multiple models or pre/post-processing steps.

Configuration

Here is a simplified serve_config.yaml to get started. This configuration defines a deployment that autoscales based on request load.

 1proxy_location: EveryNode
 2
 3http_options:
 4  host: 0.0.0.0
 5  port: 8000
 6
 7applications:
 8  - name: llm_app
 9    route_prefix: /
10    import_path: app:deployment
11    runtime_env:
12      pip:
13        - vllm
14        - fastapi
15    deployments:
16      - name: VLLMDeployment
17        autoscaling_config:
18          min_replicas: 1
19          max_replicas: 4
20          target_num_ongoing_requests_per_replica: 10
21        ray_actor_options:
22          num_gpus: 1

The FastAPI Wrapper

We wrap the vLLM engine in a FastAPI app to expose standard REST endpoints. This allows easy integration with existing frontend applications or services.

 1from fastapi import FastAPI
 2from ray import serve
 3from vllm import AsyncLLMEngine, SamplingParams
 4
 5app = FastAPI()
 6
 7@serve.deployment(num_gpus=1)
 8@serve.ingress(app)
 9class VLLMDeployment:
10    def __init__(self):
11        # Initialize vLLM engine
12        self.engine = AsyncLLMEngine.from_engine_args(...)
13
14    @app.post("/generate")
15    async def generate(self, prompt: str):
16        sampling_params = SamplingParams(temperature=0.7, max_tokens=100)
17        results = await self.engine.generate(prompt, sampling_params, ...)
18        return {"text": results[0].outputs[0].text}
19
20deployment = VLLMDeployment.bind()

Deployment on AWS

  1. Cluster Setup: Use the Ray Cluster Launcher to provision EC2 instances. Define your cluster configuration in a cluster.yaml file, specifying the instance types (e.g., g5.xlarge for workers).
  2. Deploy: Run ray up cluster.yaml to start the cluster.
  3. Serve: Submit your serve application using serve run serve_config.yaml.

Monitoring and Optimization

Ray provides a built-in dashboard to monitor actor status, GPU usage, and request latency. For production, integrate this with Prometheus and Grafana to track:

  • Queue Latency: Time requests spend waiting for a replica.
  • GPU Utilization: Ensure you aren’t under-provisioning expensive hardware.
  • Token Throughput: Measure tokens/second to benchmark performance.

By leveraging Ray Serve with AWS, we create a flexible, scalable inference platform that avoids the vendor lock-in of managed services while providing full control over the serving infrastructure.