Bello Agent

Mac native · AI conversations & tools

Bello Agent.
Room to think.

AI conversations, coding tools and independent side discussions in a native Mac app, with request details close at hand.

Download Bello Agent 0.1.6
6.35 MiB · Apple Silicon · macOS 14 or later
Developer ID signed and notarized · Updates through Sparkle

An inspectable project

From a question to a complete tool run.

Built around the Mac

Native message composers, familiar selection and undo, saved drafts, searchable chat history and a readable Markdown transcript. Group chats in collapsible projects, each with one or more trusted folders.

Keep the main thread moving

Pin frequent chats, rename sessions, and archive conversations without deleting their history. Delete an archived chat when you choose to. Open a saved child side conversation, fork the same context into another chat, queue follow-ups or steer between complete tool turns.

Tools with clear outcomes

Work with files, shell commands, selected Codex skills and configured MCP tools. Disable skills only in Bello Agent without changing Codex. Interrupted operations remain visible and aren’t automatically repeated.

Inspect the exchange

Open a message’s details to see its model, cost, timing, headers and retained request and response bodies directly. Capture is on by default for 30 days. Authentication headers are masked, and known credentials in request bodies use labeled hashes.

A clear view of usage

Open the Usage Report page for requests, reported cost, reasoning-token cost, cache use and response timing. See requested and final models side by side. Expand filters and details when needed, then return to your chat with its draft and place preserved.

Activity in your menu bar

Left or right click the menu bar icon for running work, queues, unread replies and live output speed. Switch to Usage for consumed tokens, reported cost with its reasoning portion, and model distribution. Auto-router aliases stay visible alongside the models resolved by the gateway. Missing reports remain explicit.

Choose the model for each chat

Switch models and reasoning effort in the composer. A configured catalog supplies the model list, supported efforts and context limits, and keeps its last successful list if a refresh fails. Gateway fallback models are off by default and can be enabled per connection.

Copy and revise messages

Copy individual code blocks or Markdown sections, or edit and resend from a previous user message. The new reply uses the earlier context, while replaced replies remain preserved in the local journal.

Keep context in view

Click the circular context indicator to inspect prepared instructions, messages, tools and request JSON. Compaction progress and its successful result appear above the composer; context usage remains an explicitly labeled estimate.

Get started

Your connection.
Your project.

Bello Agent requires your own LiteLLM gateway. Model availability, routing, reported telemetry and usage charges depend on that gateway.

  1. Open the DMG and drag Bello Agent into Applications.
  2. Configure your LiteLLM Responses endpoint, API key and model alias. Add a model catalog URL if your gateway’s model list is incomplete.
  3. Choose and trust your project. Test & Start checks the selected model with a small request before opening your first chat.
  4. Open a message’s Details or Requests link, or the Usage Report page, to inspect captured bodies and headers. Adjust capture and retention in Settings.

Local storage

Credentials and app configuration are kept in the native Keychain vault. HTTP request and response capture is enabled by default with 30-day body retention, subject to your storage quota and settings. New bodies are stored locally without encryption. Authentication headers are masked; longer request tokens retain only their final four characters. Known credentials found in request bodies are replaced with SHA-256 fingerprints. Bodies can still contain your prompts, files and tool data. Existing encrypted captures remain readable.

Full-text search covers chat history; searching inside captured HTTP bodies is not included. Gateway response-cache hits and provider prompt-cache tokens are shown separately. Cost totals include only values the gateway reports, with coverage counts.