API RAG

An API documentation agent that uses retrieval-augmented generation to answer endpoint questions with citations.

API RAG

What changed

  • Turned API docs search into a chat where every answer shows the snippets it came from, so it can be checked.
  • Served the same retrieval pipeline to people through a chat UI and to coding assistants through an MCP server.

What I worked on

  • I built the pipeline end to end, from OpenAPI ingestion to retrieval with Vertex AI Search and cited answers from Gemini.
  • I added the MCP tools so coding assistants can query API docs directly.

API docs are long, and their search boxes match words, not questions. API RAG lets you ask about an API in plain English and answers from the docs themselves.

How it works

Point it at an OpenAPI spec, by URL or as a local file. It splits the spec into endpoint and schema chunks, saves them as JSONL, and indexes them in Vertex AI Search. When you ask a question, it retrieves the relevant chunks and Gemini writes the answer from them instead of guessing.

There are two ways in: a Gradio chat with streaming answers and a sources panel, or an MCP server with search_docs and get_schema tools for coding assistants.

Built in Python with Vertex AI Search, Gemini, Gradio and MCP.