PyTorch Image Captioning BLIP + CLIP Streaming GPU
Model: PyTorch Image Captioning — BLIP (candidate generation) + CLIP (ranking) Accelerator: Tesla T4 GPU Host: 50 × n1-standard-4 (4 vCPUs, 15 GB RAM)
This streaming pipeline performs image captioning using a multi-model open-source PyTorch approach. It first generates multiple candidate captions per image using a BLIP model, then ranks these candidates with a CLIP model based on image-text similarity.
The following graphs show various metrics when running PyTorch Image Captioning BLIP + CLIP Streaming GPU pipeline. See the glossary for definitions.
Full pipeline implementation is available here.
What is the estimated cost to run the pipeline?
RunTime and EstimatedCost

How has various metrics changed when running the pipeline for different Beam SDK versions?
AvgThroughputBytesPerSec by Version

AvgThroughputElementsPerSec by Version

How has various metrics changed over time when running the pipeline?
AvgThroughputBytesPerSec by Date

AvgThroughputElementsPerSec by Date

Last updated on 2026/07/31
Have you found everything you were looking for?
Was it all useful and clear? Is there anything that you would like to change? Let us know!

