---
title: "Octo — Multi-agent research workspace | Harshith Nayaka L"
description: "Multi-agent research workspace on the OpenAI Agents API: specialist agents, live web search, cited reports and company records, with no duplicate paid runs."
canonical: "https://harshith-nayaka-l-portfolio.vercel.app/work/octo"
last-updated: "2026-10-04T17:07:29Z"
published: "2026-10-04T16:45:37Z"
author: "Harshith Nayaka L"
author-role: "AI Engineer - Full Stack"
author-title: "AI Engineer - Full Stack"
author-location: "Bengaluru, India"
author-availability: "Available for freelance work"
author-email: "harshith28124@gmail.com"
content-type: "text/markdown"
html-version: "https://harshith-nayaka-l-portfolio.vercel.app/work/octo"
---
# Octo — Multi-agent research workspace

> A research workspace where a director agent hands questions to specialist agents with live web search and returns cited reports and structured company records, built so a lost response can never turn into a second paid session.

Case study by Harshith Nayaka L, AI Engineer - Full Stack, Bengaluru, India.
Canonical page: https://harshith-nayaka-l-portfolio.vercel.app/work/octo

**Status: in progress.** Octo is built and tested, but not yet proven live. Its 49 backend tests and browser suite run against mocked OpenAI and Google responses; it has not run a real paid session, has not been deployed, and has no real CRM connected. Everything below describes the code as it stands.

- **Type:** Agent research workspace
- **Providers:** OpenAI Agents API (primary), Google Antigravity (secondary)
- **Role:** Solo build
- **Status:** Tested with mocks; live runs not yet verified

## The problem

Agent research tools are easy to demo and hard to trust. A report without dated sources can't be checked, a company list with an invented decision maker is worse than no list, and an agent that reads the open web will meet pages that try to give it instructions.

The operational side is just as unforgiving. Hosted agent sessions are billed, so a server restart or a response lost in transit must never quietly start the same paid job twice.

## What I built

Octo is a TypeScript workspace (React 19 and Express 5) built around the OpenAI Agents API. A research director delegates independent questions to specialist agents, at most two at a time: market, competitor, pricing and regulation specialists for market research, and company-discovery and company-evidence specialists for sales research. Google Antigravity on Gemini Interactions is the secondary provider: a run can be started on it instead, for example to test on a Google free-tier project before spending OpenAI credit.

There are three workflows. Market intelligence covers market structure, company comparisons, public pricing and regulation, ending in an investor briefing. Sales research takes an explicit ideal customer profile and a target of 1 to 100 companies, and returns fit evidence, potential needs, public decision makers and source links. Ongoing research keeps a saved session and re-checks it on a schedule against what it found before. Each run takes a brief and up to five context files and returns a Markdown report and JSON, plus CSV for sales.

The evidence rules are part of the job, not an afterthought. Every material claim needs a dated source URL; unknown values stay unknown; facts, estimates and hypotheses are kept apart; and websites and uploaded files are treated as data, never as instructions. Structured output from either provider is validated before any company record is shown, and CSV export neutralises spreadsheet formulas.

It runs locally on Node's built-in SQLite, or on Vercel as one function backed by Postgres, where each run advances in short leased steps because a serverless function cannot hold a background worker.

## Pipeline

- **Brief:** Brief + files (Workflow, ICP, up to 5 documents)
- **Direct:** Research director (Splits the work, sets the rules)
- **Research:** Specialist agents (Two at a time, live web search) → Hosted sandbox (Inputs, notes, outputs)
- **Check:** Validate output (Schema + dated sources)
- **Deliver:** Report + records (Markdown, JSON, CSV) → CRM export (Separate session, after review)

## The judgment calls

**A lost response never becomes a second paid session**

If the server restarts or a create response is lost, Octo searches the project's saved sessions for the local run ID before doing anything else. If the outcome is still uncertain and no match is found, it stops for manual inspection instead of creating another billed session. On Google, which has no way to look a session up by local ID, an uncertain submission is never repeated automatically.

**The database decides which instance advances a run**

On Vercel there is no single long-running worker, so any request can trigger the next step. A database lease lets exactly one function instance advance a run at a time, with steps spaced at least four seconds apart. A test simulates two instances stepping the same run and requires exactly one remote session; deliberately breaking the lease makes that test fail.

**No silent switching between providers**

Each run is pinned to the provider, agent and model it started with. If that provider is unavailable or rejects a request, the run reports the error; it is never quietly retried on the other provider, which would change both the cost and the quality of the result without anyone choosing it.

**CRM writes are walled off from research**

Research sessions have no CRM access at all. Exporting results is a separate session that can use only an allow-listed set of CRM tools, and only after a person reviews the company list and confirms the export. The agent is told to deduplicate by website and never to contact anyone.

## What it changed

**Tested:** 49 backend tests on mocked provider responses and temporary databases, including Postgres run through PGlite, plus a Playwright suite of four runs across desktop and mobile, and a local Vercel build of the serverless function.

**Honest scope:** No live paid OpenAI or Google session has been run, it has not been deployed, and no real CRM is connected. It is a single-workspace, self-hosted starter, not a multi-tenant SaaS; its README lists what a hosted service would still need, from per-user isolation to enforced budgets.

## Questions this answers

**How do you stop an AI research agent from inventing companies or sources?**

Make the evidence rules part of the task and check the output before showing it. In Octo, a research workspace built by Harshith Nayaka L, every material claim needs a dated source URL, unknown values must stay unknown, and the agents are told never to invent a company, decision maker, price, citation or financial figure; retrieved pages are treated as data, not instructions. Structured output is validated before any company record appears. Octo's tests run on mocked provider responses, so its live research quality has not yet been verified.

## Built with

TypeScript, React 19, Express 5, OpenAI Agents API, Google Antigravity (Gemini Interactions), SQLite (Node built-in), Postgres on Vercel (Neon), Zod, Vitest, Playwright

## Links

- [View on GitHub](https://github.com/HarshithNayakaL/octo)
