PyTorch Image Captioning BLIP + CLIP Streaming CPU

Model: PyTorch Image Captioning — BLIP (candidate generation) + CLIP (ranking) Accelerator: CPU only Host: 50 × n1-standard-4 (4 vCPUs, 15 GB RAM)

This streaming pipeline performs image captioning using a multi-model open-source PyTorch approach. It first generates multiple candidate captions per image using a BLIP model, then ranks these candidates with a CLIP model based on image-text similarity.

The following graphs show various metrics when running PyTorch Image Captioning BLIP + CLIP Streaming CPU pipeline. See the glossary for definitions.

Full pipeline implementation is available here.

What is the estimated cost to run the pipeline?

RunTime and EstimatedCost

RunTime and EstimatedCost

How has various metrics changed when running the pipeline for different Beam SDK versions?

AvgThroughputBytesPerSec by Version

AvgThroughputBytesPerSec by Version

AvgThroughputElementsPerSec by Version

AvgThroughputElementsPerSec by Version

How has various metrics changed over time when running the pipeline?

AvgThroughputBytesPerSec by Date

AvgThroughputBytesPerSec by Date

AvgThroughputElementsPerSec by Date

AvgThroughputElementsPerSec by Date