Skip to main content
In this lesson we’ll explore how to analyze images with Azure AI Vision. You’ll learn what goes into image analysis, how to call the service via the REST API and SDKs (C# and Python), how to configure analysis options, and how to parse the structured responses the service returns. We cover:
  • What the Analyze API returns (captions, detected objects and people, OCR/read, smart crops, etc.)
  • How to select Visual Features to limit and focus the response
  • SDK usage patterns and a full Python example to parse results
  • REST usage patterns and a sample query string for the Analyze endpoint
  • Practical options (smart crops, language, gender-neutral captions, model versioning)
Key aspects of image analysis are summarized in the table below.
An infographic titled "Working with Image Analysis" showing five numbered panels that summarize: Analyze AI Overview, Visual Features Enum, SDK Integration, REST API Usage, and Input Requirements. Each panel includes a short description and an icon explaining the corresponding image-analysis feature.

REST API example

A typical REST Analyze request is performed against the Image Analysis endpoint. Example URL (replace <your-endpoint> and ):
  • Query parameters:
    • features — comma-separated visual features to return (example: caption, people, objects, read, smartCrops).
    • model-name — model to use (e.g., latest or a specific version).
    • language — language for captions / OCR results.
    • api-version — service API version.
You include the image either as:
  • an image URL in the JSON request body, or
  • raw image bytes in the request body (binary upload).
The service responds with structured JSON containing captionResult, objectsResult, peopleResult, smartCropsResult, tagsResult/read results, metadata, and modelVersion.

SDK usage (C# and Python — conceptual)

SDKs simplify calls and return typed objects. Below are conceptual method signatures to illustrate common patterns. C# (conceptual):
Python (conceptual):

Visual features (examples)

Analysis options

You can tune the behavior of the analysis call with these options:
  • Cropping aspect ratios — request smart-crop suggestions for thumbnail generation or fixed aspect ratios.
  • Gender-neutral captioning — enable gender-neutral language for generated captions.
  • Language selection — specify language for OCR and captions.
  • Model versioning — pin to a specific model for reproducible results.
  • Additional flags — options vary between SDKs and REST; consult the model-name and API docs.
A dark-themed infographic titled "Image Analysis options" that lists configurable settings for image analysis. It highlights four features: Cropping Aspect Ratios, Gender-Neutral Captioning, Language Selection, and Model Versioning, each with an icon and short description.

Example: setting analysis options

C# (conceptual):
Python (conceptual):

Image analysis results

Responses from the service are structured and predictable so you can parse them reliably. Typical top-level sections:
  • captionResult — best caption and confidence
  • objectsResult — array of detected objects with bounding boxes and confidence
  • peopleResult — array of people detections with bounding boxes and confidence
  • smartCropsResult — suggested crop boxes for requested aspect ratios
  • tagsResult / tags — label/tag information and confidence
  • read / ocr results — recognized text blocks/lines
  • metadata — image dimensions and format
  • modelVersion — the model used for inference
A dark-themed infographic titled "Image Analysis Result" with four colored panels labeled Caption Result, Object Detection, Smart Crops, and Hierarchical Data, each showing an icon and a short description. It explains that successful image analysis returns structured data (JSON/SDK).
Example JSON structure (illustrative):
Use these fields to:
  • render captions for accessibility,
  • draw bounding boxes for objects and people,
  • select recommended crops for thumbnails, and
  • display detected tags and OCR text in the UI.

Hands-on: Python SDK example

Install the Azure AI Vision package for Python:
Replace endpoint and key values below with the endpoint and key from your Azure AI service (Keys and Endpoint in the Azure portal). Never commit production keys into source control.
A consolidated, practical Python example demonstrating initialization, choosing visual features, calling analysis, and parsing results safely:
This script:
  • Initializes ImageAnalysisClient with your endpoint and key.
  • Chooses the visual features to analyze.
  • Calls analyze_from_url with optional analysis options.
  • Prints the raw JSON response and demonstrates robust parsing of common result sections (people, caption, tags, objects).
Protect your API keys: rotate keys regularly, store secrets in a secure vault (e.g., Azure Key Vault), and avoid hard-coding secrets in source control.

Live demonstration notes and best practices

  • Provision an Azure AI service in the Azure portal. Use the Keys and Endpoint values from the portal for your client.
  • Use blob storage URLs or public URLs for images. For private images, upload binary image bytes in the request body.
  • Gender-neutral captions help avoid gender assumptions in generated text (e.g., “a person hugging a dog”).
  • Smart crops return bounding boxes for the aspect ratios you specify—use these to create thumbnails that preserve important content.
  • Pin model versions for reproducible results; use “latest” for new features and model improvements.
  • Always validate and sanitize service outputs before surface-level display in production applications.
Example parsed output (illustrative):
This guide demonstrates how to call Azure AI Vision to obtain captions, detect objects and people, extract OCR text, and request smart crops. Use these structured outputs to annotate images, drive UI decisions (smart cropping), and provide accessible descriptions for your applications.

Watch Video