OpenAI and Anthropic API
🎯 Goal
After reading this chapter:
- You know how to work with OpenAI and Anthropic APIs
- You use streaming responses, function calling, vision APIs
- You know how to reduce costs by up to 90% with prompt caching
- You add retry, rate limit, error handling to production
What to learn
- OpenAI SDK — Python client
- Anthropic SDK — Python client
- Chat completions — main API
- Streaming — real-time response
- Function calling / Tool use — structured actions
- Vision — working with images
- Embeddings — for semantic search
- Prompt caching(Anthropic) — 90% cost reduction
- Batching — async parallel calls
- Rate limitingand retry strategies
- Token tracking and observability
Libraries
pip install openai anthropic
pip install instructor # structured output
pip install tenacity # retry logic
pip install backoff # exponential backoff
Code examples
OpenAI — basic chat
from openai import OpenAI
client = OpenAI(api_key="sk-...") # or os.getenv("OPENAI_API_KEY")
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Hello! What is list comprehension in Python?"},
],
temperature=0.7,
max_tokens=500,
)
print(response.choices[0].message.content)
print(f"Tokens: in={response.usage.prompt_tokens}, out={response.usage.completion_tokens}")
Anthropic — basic message
from anthropic import Anthropic
client = Anthropic(api_key="sk-ant-...")
response = client.messages.create(
model="claude-sonnet-4-6",
max_tokens=1024,
system="You are a helpful assistant.",
messages=[
{"role": "user", "content": "What is list comprehension in Python?"},
],
)
print(response.content[0].text)
print(f"Tokens: in={response.usage.input_tokens}, out={response.usage.output_tokens}")
Streaming — real-time
OpenAI streaming
stream = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Write a long story"}],
stream=True,
)
for chunk in stream:
if chunk.choices[0].delta.content is not None:
print(chunk.choices[0].delta.content, end="", flush=True)
Anthropic streaming
with client.messages.stream(
model="claude-sonnet-4-6",
max_tokens=1024,
messages=[{"role": "user", "content": "Write a long story"}],
) as stream:
for text in stream.text_stream:
print(text, end="", flush=True)
Function Calling / Tool Use
OpenAI function calling
tools = [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Returns the weather for a given city",
"parameters": {
"type": "object",
"properties": {
"city": {"type": "string", "description": "City name"},
"unit": {"type": "string", "enum": ["celsius", "fahrenheit"]},
},
"required": ["city"],
},
},
}]
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "What's the weather in Toshkent?"}],
tools=tools,
)
# Execute the tool call
tool_call = response.choices[0].message.tool_calls[0]
if tool_call.function.name == "get_weather":
args = json.loads(tool_call.function.arguments)
weather = get_weather(args["city"], args.get("unit", "celsius"))
# Send the result back to the LLM
response2 = client.chat.completions.create(
model="gpt-4o-mini",
messages=[
{"role": "user", "content": "What's the weather in Toshkent?"},
response.choices[0].message,
{"role": "tool", "tool_call_id": tool_call.id, "content": str(weather)},
],
tools=tools,
)
print(response2.choices[0].message.content)
Anthropic tool use
tools = [{
"name": "get_weather",
"description": "Returns the weather for a given city",
"input_schema": {
"type": "object",
"properties": {
"city": {"type": "string"},
"unit": {"type": "string", "enum": ["celsius", "fahrenheit"]},
},
"required": ["city"],
},
}]
response = client.messages.create(
model="claude-sonnet-4-6",
max_tokens=1024,
tools=tools,
messages=[{"role": "user", "content": "What's the weather in Toshkent?"}],
)
# Execute the tool use
for block in response.content:
if block.type == "tool_use":
if block.name == "get_weather":
result = get_weather(**block.input)
# Send result back
response2 = client.messages.create(
model="claude-sonnet-4-6",
max_tokens=1024,
tools=tools,
messages=[
{"role": "user", "content": "What's the weather in Toshkent?"},
{"role": "assistant", "content": response.content},
{"role": "user", "content": [{
"type": "tool_result",
"tool_use_id": block.id,
"content": str(result),
}]},
],
)
Vision API
OpenAI vision
import base64
def encode_image(image_path: str) -> str:
with open(image_path, "rb") as f:
return base64.b64encode(f.read()).decode()
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "What do you see in this image?"},
{
"type": "image_url",
"image_url": {"url": f"data:image/jpeg;base64,{encode_image('photo.jpg')}"},
},
],
}],
)
Anthropic vision
response = client.messages.create(
model="claude-sonnet-4-6",
max_tokens=1024,
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "What do you see in this image?"},
{
"type": "image",
"source": {
"type": "base64",
"media_type": "image/jpeg",
"data": encode_image("photo.jpg"),
},
},
],
}],
)
Prompt Caching (Anthropic) — 90% cheaper!
# Large system prompt is cached, not paid repeatedly
LARGE_SYSTEM = open("docs.md").read() # 50K token docs
response = client.messages.create(
model="claude-sonnet-4-6",
max_tokens=1024,
system=[
{
"type": "text",
"text": LARGE_SYSTEM,
"cache_control": {"type": "ephemeral"}, # CACHE!
},
],
messages=[{"role": "user", "content": "Question about the docs..."}],
)
# First time: full price + cache write (1.25x) # Next 5 minutes: 0.1x price (90% cheaper!)
Embeddings
OpenAI embeddings
response = client.embeddings.create(
model="text-embedding-3-small", # 1536-dim, $0.02 / 1M tokens
input=["Salom dunyo", "Machine learning"],
)
embeddings = [d.embedding for d in response.data]
# Shape: [(1536,), (1536,)]
Anthropic embeddings? — none
Anthropic has no embeddings API. Options:
- OpenAI text-embedding-3-small
- Voyage AI (recommended by Anthropic)
- Cohere embeddings
- Sentence Transformers (local)
Retry + Rate Limiting
from tenacity import retry, stop_after_attempt, wait_exponential
from openai import RateLimitError, APIError
@retry(
stop=stop_after_attempt(5),
wait=wait_exponential(multiplier=1, min=2, max=60),
retry=lambda e: isinstance(e, (RateLimitError, APIError)),
)
async def call_llm_with_retry(messages: list, model: str = "gpt-4o-mini"):
response = await async_client.chat.completions.create(
model=model,
messages=messages,
)
return response.choices[0].message.content
Async batching
import asyncio
from openai import AsyncOpenAI
async_client = AsyncOpenAI()
async def process_one(text: str):
response = await async_client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": f"Summarize:{text}"}],
)
return response.choices[0].message.content
async def process_batch(texts: list[str], max_concurrent: int = 10):
sem = asyncio.Semaphore(max_concurrent)
async def bounded(text):
async with sem:
return await process_one(text)
return await asyncio.gather(*[bounded(t) for t in texts])
# 100 texts with 10 concurrent
results = asyncio.run(process_batch(texts, max_concurrent=10))
Cost tracking middleware
import logging
from contextlib import contextmanager
logger = logging.getLogger("llm_costs")
PRICES = {
"gpt-4o-mini": (0.15, 0.60),
"claude-sonnet-4-6": (3.00, 15.00),
"claude-haiku-4-5": (0.80, 4.00),
}
@contextmanager
def track_llm_call(model: str, user_id: int = None):
"""Usage: with track_llm_call("gpt-4o-mini"): ..."""
response_holder = {}
def hook(response):
response_holder["response"] = response
yield hook
response = response_holder.get("response")
if response and hasattr(response, "usage"):
u = response.usage
in_price, out_price = PRICES[model]
cost = (u.prompt_tokens * in_price + u.completion_tokens * out_price) / 1_000_000
logger.info(f"model={model}in={u.prompt_tokens}out={u.completion_tokens}"
f"cost=${cost:.6f}user={user_id}")
Backend integration
Streaming chat endpoint in FastAPI (SSE)
from fastapi import FastAPI
from fastapi.responses import StreamingResponse
from openai import AsyncOpenAI
app = FastAPI()
client = AsyncOpenAI()
class ChatRequest(BaseModel):
message: str
session_id: str
async def stream_chat(messages: list):
stream = await client.chat.completions.create(
model="gpt-4o-mini",
messages=messages,
stream=True,
)
async for chunk in stream:
if chunk.choices[0].delta.content:
text = chunk.choices[0].delta.content
yield f"data:{json.dumps({'text': text})}\n\n"
yield "data: [DONE]\n\n"
@app.post("/chat/stream")
async def chat_stream(req: ChatRequest):
history = await get_history(req.session_id)
messages = history + [{"role": "user", "content": req.message}]
return StreamingResponse(
stream_chat(messages),
media_type="text/event-stream",
)
WebSocket chat
from fastapi import WebSocket
@app.websocket("/ws/chat")
async def chat_ws(websocket: WebSocket):
await websocket.accept()
try:
while True:
data = await websocket.receive_json()
messages = data["messages"]
async with client.messages.stream(
model="claude-sonnet-4-6",
max_tokens=1024,
messages=messages,
) as stream:
async for text in stream.text_stream:
await websocket.send_json({"type": "delta", "text": text})
await websocket.send_json({"type": "done"})
except Exception as e:
await websocket.send_json({"type": "error", "message": str(e)})
await websocket.close()
Multi-provider abstraction
from abc import ABC, abstractmethod
class LLMProvider(ABC):
@abstractmethod
async def chat(self, messages: list, **kwargs) -> str: ...
class OpenAIProvider(LLMProvider):
def __init__(self, model="gpt-4o-mini"):
self.client = AsyncOpenAI()
self.model = model
async def chat(self, messages, **kwargs):
response = await self.client.chat.completions.create(
model=self.model, messages=messages, **kwargs)
return response.choices[0].message.content
class AnthropicProvider(LLMProvider):
def __init__(self, model="claude-sonnet-4-6"):
from anthropic import AsyncAnthropic
self.client = AsyncAnthropic()
self.model = model
async def chat(self, messages, **kwargs):
# Separate out the system message
system = next((m["content"] for m in messages if m["role"] == "system"), None)
msgs = [m for m in messages if m["role"] != "system"]
response = await self.client.messages.create(
model=self.model,
max_tokens=kwargs.pop("max_tokens", 1024),
system=system,
messages=msgs,
**kwargs,
)
return response.content[0].text
# Usage
provider = OpenAIProvider("gpt-4o-mini")
# or
provider = AnthropicProvider("claude-haiku-4-5")
response = await provider.chat([{"role": "user", "content": "Hello"}])
Resources
- OpenAI docs — platform.openai.com/docs
- Anthropic docs — docs.anthropic.com
- OpenAI Cookbook — cookbook.openai.com
- Anthropic Cookbook — GitHub
- LiteLLM — universal LLM wrapper: litellm.ai
- OpenRouter — one API for many models: openrouter.ai
🏋️ Exercises
🟢 Easy
- “Hello World” with OpenAI and Anthropic APIs — 5 Q&A.
- Get a streaming response, output each char separately.
- Embedding for similarity between 2 sentences.
🟡 Medium
- Function calling: weather, calculator, search — agent with 3 tools.
- Vision: Upload an image, extract structured data from it (Instructor + vision).
- Prompt caching: 10 questions with a large system prompt — observe the cost difference.
🔴 Hard
- Multi-provider chat: OpenAI/Anthropic/Google — one abstraction, auto-fallback.
- Cost-aware router: Automatically choose the right model by input complexity and context size.
- Streaming chatbot: FastAPI + WebSocket + Postgres history + Redis caching.
Capstone
notebooks/month-05/03_llm_apis.ipynb:
- Get fully familiar with 3 providers (OpenAI, Anthropic, OpenRouter)
- Multi-turn chatbot with streaming
- Function calling — 5 tools
- Vision — image classification
- Cost tracking dashboard
✅ Checklist
- I know the OpenAI and Anthropic APIs
- I use streaming responses
- Function calling / tool use
- Working with the Vision API
- Computing and storing embeddings
- Prompt caching (Anthropic)
- Async batching
- Retry and rate limit handling
- Cost tracking and observability
Moving on to LangChain and LlamaIndex.