Generate images and video with vLLM-Omni on SageMaker AI – Part 2
Deploy two generative media models from one AWS vLLM-Omni Deep Learning Container on Amazon SageMaker AI. Generate an image with FLUX.2-klein through real-time inference, then animate it into video with Wan2.1-VACE through asynchronous inference, and retrieve the MP4 from Amazon S3.
Overview
In this post, you turn a text prompt into an image, then animate that image into a short video on Amazon SageMaker AI. You deploy two endpoints from the same AWS vLLM-Omni Deep Learning Container (DLC): a real-time endpoint for FLUX.2-klein-4B image generation and an asynchronous endpoint for Wan2.1-VACE-1.3B video generation. The workflow sends a text prompt to generate a still image, then passes the image and a motion prompt to the video endpoint. It retrieves the MP4 from Amazon Simple Storage Service (Amazon S3) and provides an optional Streamlit interface.
AWS Deep Learning Containers package frameworks and dependencies for training and inference on AWS. The AWS vLLM-Omni DLC packages tracked vLLM-Omni releases and adds routing middleware for SageMaker AI. vLLM-Omni extends vLLM beyond text generation to models that process or generate text, audio, images, and video through OpenAI-compatible APIs.
This post continues a series about specialized AWS DLCs. Part 1 uses vLLM-Omni and SageMaker AI bidirectional streaming for real-time speech. Part 2 covers real-time and asynchronous inference: the image model returns its result inline, while the longer-running video model writes its output to Amazon S3. Keeping image and video generation separate from the text-to-speech walkthrough also makes the different models, payloads, and response patterns clear.
You clone the code sample, deploy FLUX.2-klein-4B for image generation, and pass its output to Wan2.1-VACE-1.3B for image-conditioned video generation. The sample includes a command-line workflow and a Streamlit application.
Solution overview
The solution deploys the same pinned AWS vLLM-Omni DLC image to two SageMaker AI endpoints. The deployment script changes SM_VLLM_MODEL to load FLUX.2-klein on one endpoint and Wan VACE on the other. Keeping a common container image reduces serving-stack variation, while separate endpoints let each model use the instance type and inference option that fits its workload.
- A SageMaker AI real-time endpoint runs FLUX.2-klein and routes requests to
/v1/images/generations. - A SageMaker Asynchronous Inference endpoint runs Wan VACE and routes requests to
/v1/videos/sync. - Amazon S3 stores the multipart video request and the generated MP4.
Figure 1 shows the request path. The application sends the image prompt to FLUX.2-klein and receives a base64-encoded PNG. It resizes the image to the video dimensions, converts it to a compact JPEG data URL, and inserts that reference into the Wan VACE request. SageMaker Asynchronous Inference reads the multipart request from Amazon S3 and writes the MP4 to the returned output location.
Figure 1: A CLI or Streamlit application invokes a FLUX.2-klein real-time endpoint, passes the generated PNG through Amazon S3 to a Wan VACE asynchronous endpoint, and retrieves the generated MP4 from Amazon S3
Figure 2 shows the same flow as a numbered request-response sequence. The image call returns directly to the application. The video call returns an asynchronous output location, and the application retrieves the MP4 from Amazon S3 when generation finishes.
Figure 2: A six-step sequence shows the application invoking the image endpoint, receiving a PNG, uploading a video request to Amazon S3, invoking the asynchronous video endpoint, and retrieving the generated MP4
SageMaker AI sends inference traffic to /invocations. The DLC reads the CustomAttributes header and forwards the request to the selected vLLM-Omni route. The sample pre-builds the multipart video body before uploading it to Amazon S3, so the exact request sent to the vLLM-Omni Videos API remains explicit. It stores successful responses and invocation failures under separate Amazon S3 prefixes.
The sample uses real-time inference for the image because the application expects a direct response before it can construct the next request. Video generation is a longer-running operation and carries an image-conditioned multipart payload. SageMaker Asynchronous Inference queues the request and uses Amazon S3 for the input and output, so the client can poll the result location instead of holding one synchronous request open. This is a workload choice rather than a rule for every video model. Use the inference option that matches your model latency, payload, and client interaction.
The sample uses fixed ml.g6.xlarge and ml.g6e.xlarge instance types. For the real-time image endpoint, SageMaker capacity-aware instance pools can list compatible instance types in priority order. Add a pool only after you validate each candidate type for the image model’s GPU memory and performance requirements. This sample keeps the asynchronous video endpoint on a fixed instance type.
Prerequisites
Before starting, you need:
- An AWS account with credentials configured for the AWS Command Line Interface (AWS CLI) or an AWS SDK.
- A SageMaker AI execution role with access to the sample Amazon S3 bucket.
- Permission to create and invoke SageMaker AI endpoints.
- Endpoint quota for
ml.g6.xlargeandml.g6e.xlargein your selected AWS Region. - Git.
- Python 3.11 or later.
This walkthrough uses the US East (N. Virginia) AWS Region. Review SageMaker AI pricing before deploying the GPU endpoints.
Generate an image and animate it
- Clone the hosting examples repository.Clone the repository and enter the vLLM-Omni image and video sample directory.
- Install the Python dependencies.Create a virtual environment and install the sample requirements.
- Deploy the image and video endpoints.Set your SageMaker AI execution role and run the deployment script. The script pins
omni-sagemaker-cuda-v1.6, creates a real-time image endpoint, creates an asynchronous video endpoint, and writes their resource names to.vllm_omni_media_state.json.The environment variable
SM_VLLM_MODELtells each container which model to load. The video endpoint turns on variational autoencoder (VAE) tiling to reduce peak memory during video decoding. If your local AWS credentials expire while an endpoint starts, refresh them and rerun the command with--resume. The script reuses the saved resources. - Generate an image and animate it.Run the end-to-end command-line workflow with an image prompt and a motion prompt.
FLUX.2-klein returns an OpenAI-compatible JSON response that contains the PNG as base64 data. The shared request helper resizes that image to the target video dimensions and encodes it as JPEG before building the Wan VACE request:
This conversion keeps the JSON-safe image reference below the Videos API multipart parser’s per-part size. The sample uploads the complete request to Amazon S3 and submits it to the asynchronous endpoint with
route=/v1/videos/sync. SageMaker AI returns output and failure locations. The script polls them before validating and saving the MP4 inoutputs/.For smaller requests, InvokeEndpointAsync also accepts inline request data through
Bodyup to 128,000 bytes. Use eitherBodyorInputLocation, not both. This workflow usesInputLocationbecause its image-conditioned multipart request is larger than the 128,000-byte inlineBodydesign parameter documented forInvokeEndpointAsync.The default video settings use 17 frames and 30 diffusion steps. Use four steps only for a quick endpoint smoke test because the lower setting did not preserve the source composition in our review.
In one validation run in the US East (N. Virginia) Region, the
ml.g6.xlargeimage endpoint reachedInServicein 9 minutes 30 seconds. Theml.g6e.xlargevideo endpoint reachedInServicein 8 minutes 40 seconds. A warm default invocation returned the image in 4.7 seconds. Wan VACE reported 8.9 seconds of model latency, and the CLI retrieved the MP4 in 11.8 seconds. These single-run values are reproduction checkpoints, not performance benchmarks. - Run the Streamlit application.Start the Streamlit application after both endpoints reach
InService.Figure 3 shows the complete application flow. Update the image prompt, image size, or seed, then choose Generate image. After reviewing the FLUX.2-klein result, enter a motion prompt and adjust the frame count, frames per second, or inference steps before choosing Generate video. The demo replays artifacts from the validated run. The published application invokes the deployed endpoints.
Figure 3: Configure the image and motion prompts in Streamlit, generate a FLUX.2-klein image, then create the Wan VACE video
Figure 4 shows the generated Wan VACE MP4 without the application interface. The application displays and downloads both artifacts while using the same request helpers as the command-line workflow.
Figure 4: Wan VACE animates the FLUX.2-klein coastal observatory image with the requested camera and environmental motion
Clean up
Delete the endpoints, endpoint configurations, and models when you finish:
The cleanup script retains generated objects under the Amazon S3 output prefix for review. Delete that prefix when you no longer need the artifacts. A running GPU endpoint continues to incur charges.
Conclusion
You deployed one AWS vLLM-Omni DLC release for two generative media models, returned the FLUX.2-klein image through real-time inference, and retrieved the Wan VACE video through asynchronous inference. The workflow shows how a shared serving runtime can support different model families while SageMaker AI applies the response pattern that fits each workload.
Try the code sample, then review the vLLM-Omni DLC documentation and SageMaker Asynchronous Inference developer guide to adapt the pattern to other supported image and video models.
Resources
- vLLM-Omni supported models
- AWS vLLM-Omni DLC tested models
- SageMaker Asynchronous Inference Developer Guide
- Draft image and video code sample
- Code sample pull request
- Build real-time voice applications with vLLM-Omni on SageMaker AI – Part 1
- FLUX.2-dev notebook with the AWS vLLM-Omni DLC
Acknowledgements
The authors thank the AWS Deep Learning Containers and SageMaker AI teams for their technical review and contributions to the sample.
About the authors
Source
Originally published at aws.amazon.com.


