Topic 09 · Computer Vision
Object Detection, Classification & Counting
Upload an image or use your webcam — DETR detects every object with bounding boxes and counts them.
ViT classifies images into 1000 ImageNet categories. All models run locally in your browser.
Select a tab — model loads on first use and is cached permanently.
🔍 Object Detection
🏷 Classification
📷 Live Webcam
Upload Image
Click or drag & drop
JPG · PNG · WebP
Detect & Count
Sample image
Xenova/detr-resnet-50 · Detection Transformer · ~170MB · cached permanently
Detection Result
Upload an image to detect objects
Objects will appear here…
How DETR Object Detection Works
1 · CNN Backbone
ResNet-50 extracts a 2D feature map encoding what and where objects might be in the image.
2 · Transformer Encoder
DETR's encoder attends across all feature map positions, building a global understanding of the scene.
3 · 100 Object Queries
The decoder runs 100 learned "object queries" in parallel, each predicting one box + class. No NMS needed.
Image Classification
Top-5 predictions from 1000 ImageNet classes using Vision Transformer (ViT).
Classify
Sample image
Xenova/vit-base-patch16-224 · Vision Transformer · ~346MB · ImageNet 1k
Top-5 Predictions
Upload an image to classify…
Live Webcam Detection
Point your camera at any scene — DETR detects and labels every object in real time,
entirely in your browser. No data is sent to any server.
⚡ First use downloads DETR (~170MB) and caches it permanently. Reload is instant.
Start Camera
Stop
Confidence threshold: 50 %
Camera feed will appear here
Bounding boxes are drawn on a canvas overlay · Model: DETR-ResNet50