Skip to main content
This guide demonstrates how to structure a small web application that uses the OpenAI Python client to talk to a local Ollama backend (or switch later to OpenAI’s hosted API with minimal changes). We’ll build a simple Flask-based AI Poem Generator that sends user prompts to a chat completion endpoint and renders the model’s output. Using a client library (instead of curl) helps keep your app code clean and portable across local development and hosted providers.
A slide titled "What to expect" showing bidirectional arrows between an Ollama icon and the OpenAI logo, paired with a Python logo. Two checked items on the right read "The way of Programming" and "Structuring your code."

What you’ll need

  • Python 3.8+ installed
  • Ollama running locally for local development (optional if you target OpenAI hosted APIs)
  • Familiarity with virtual environments and Flask
Useful references:

Project setup

  1. Create a project folder, for example:
  1. Create and activate a virtual environment:
  1. Install dependencies:
  1. Open your editor and create server.py at the project root.
A screenshot of Visual Studio Code's welcome page showing the "Visual Studio Code — Editing evolved" header with Start options (New File, Open, Clone Git Repository) and a Recent files list. The left sidebar shows an Explorer with a folder named "OLLAMA-APP" and the right side has a "Get Started with VS Code" walkthrough card.

server.py — full implementation

This example exposes a single route (”/”): GET renders a small form, POST sends the prompt to the chat completions endpoint using the OpenAI Python client and displays the returned poem.

Configuration — environment variables

Store runtime configuration in a .env file at the project root. This keeps credentials and endpoints out of your code and makes switching between local Ollama and OpenAI hosted APIs straightforward. Example .env contents:

How the app works (high-level)

  • The HTML form posts the user’s prompt to ”/”.
  • The Flask route builds a messages array: a system message to define role plus the user’s message.
  • The OpenAI Python client sends that to the chat completions endpoint (client.chat.completions.create(...)) and returns a response object.
  • We extract the model output at response.choices[0].message.content and render it inside the page.

Start Ollama’s REST API (local)

Ensure the Ollama local service is running so the OpenAI client can reach it. Typical local start command:
By default Ollama exposes endpoints such as POST /v1/chat/completions and listens on localhost:11434. Representative Ollama server log when started:
output

Run the Flask app

With your virtualenv active and .env in place:
You should see the Flask development server start:
Open http://127.0.0.1:3000, enter a prompt such as “Write a short poem about cats” and click “Generate Poem.” The app will POST the prompt to the model and display the returned poem.
A browser screenshot of a web app titled "AI Poem Generator" with a prompt input box, a "Generate Poem" button, and a panel labeled "Your AI-Generated Poem." The displayed poem is a block of text about felines.

Logs you will see

  • Flask logs GET and POST requests to /.
  • Ollama logs incoming requests to /v1/chat/completions.
Example Flask access log:

Next steps and production considerations

  • This demo illustrates a minimal integration pattern. To prepare for production:
    • Use a production WSGI server (gunicorn/uvicorn) instead of Flask’s dev server.
    • Store secrets securely (e.g., a secret manager or environment variable service).
    • Add robust error handling, rate limiting, and input validation/sanitization.
    • Cache or paginate long-running requests and handle model timeouts gracefully.
    • When switching to OpenAI hosted APIs, change the LLM_ENDPOINT and set a valid OPENAI_API_KEY.
This is a simple demo intended for local development. For production, use a production WSGI server (gunicorn/uvicorn), secure environment secrets properly, and follow best practices for rate-limiting, error handling, and user input sanitization.

Watch Video

Practice Lab