<!-- Generated by `just docs` from catalog/toolkits/*.yaml, catalog/evals/scorecard.json. Edit the source, not this file. -->

# Toolkits

The shared catalog ships 65 toolkits: 2296 action tools and 11 triggers. 346 of those actions are classed destructive, which the mutation gate holds until the call's arguments carry `"confirm": true` beside the tool's own fields. A project can add private toolkits of its own, visible to that project alone; see [Private tools](../../using/private-tools.md).

An agent does not read this list at runtime. It calls `search_tools` with an intent and gets back a ranked slate, so the catalog's size costs the model nothing in context.

| Toolkit | Slug | Tools | By class | Triggers | Auth | top-1 | top-8 |
| --- | --- | --- | --- | --- | --- | --- | --- |
| [Adyen](./adyen_checkout.md) | `adyen_checkout` | 27 | 3 read, 21 write, 3 destructive | 0 | `api_key` | 30/40 (75.0%) | 39/40 (97.5%) |
| [Airtable](./airtable.md) | `airtable` | 28 | 11 read, 13 write, 4 destructive | 0 | `oauth2`, `api_key` | 24/42 (57.1%) | 39/42 (92.9%) |
| [Algolia Ingestion](./algolia_ingestion.md) | `algolia_ingestion` | 56 | 17 read, 33 write, 6 destructive | 0 | `api_key` | 47/69 (68.1%) | 65/69 (94.2%) |
| [Algolia Search](./algolia_search.md) | `algolia_search` | 58 | 22 read, 23 write, 13 destructive | 0 | `api_key` | 45/74 (60.8%) | 70/74 (94.6%) |
| [Amplitude](./amplitude.md) | `amplitude` | 32 | 16 read, 14 write, 2 destructive | 0 | `api_key` | 37/51 (72.5%) | 47/51 (92.2%) |
| [Asana](./asana.md) | `asana` | 33 | 15 read, 14 write, 4 destructive | 0 | `oauth2` | 10/65 (15.4%) | 41/65 (63.1%) |
| [atmon](./atmon.md) | `atmon` | 9 | 8 read, 1 write | 0 | `api_key` | 24/29 (82.8%) | 28/29 (96.6%) |
| [Box](./box.md) | `box` | 36 | 11 read, 17 write, 8 destructive | 2 | `oauth2` | 31/54 (57.4%) | 52/54 (96.3%) |
| [Calendly](./calendly.md) | `calendly` | 31 | 23 read, 3 write, 5 destructive | 0 | `oauth2`, `api_key` | 28/62 (45.2%) | 56/62 (90.3%) |
| [ClickUp](./clickup.md) | `clickup` | 39 | 14 read, 17 write, 8 destructive | 0 | `oauth2`, `api_key` | 29/98 (29.6%) | 72/98 (73.5%) |
| [Confluence](./confluence.md) | `confluence` | 30 | 16 read, 12 write, 2 destructive | 0 | `oauth2` | 18/42 (42.9%) | 36/42 (85.7%) |
| [Discord](./discord.md) | `discord` | 39 | 14 read, 19 write, 6 destructive | 0 | `api_key`, `oauth2` | 27/60 (45.0%) | 52/60 (86.7%) |
| [DocuSign](./docusign.md) | `docusign` | 29 | 13 read, 12 write, 4 destructive | 0 | `oauth2` | 28/44 (63.6%) | 37/44 (84.1%) |
| [Dropbox](./dropbox.md) | `dropbox` | 31 | 13 read, 11 write, 7 destructive | 1 | `oauth2` | 17/49 (34.7%) | 42/49 (85.7%) |
| [Dropbox Sign](./dropbox_sign.md) | `dropbox_sign` | 31 | 11 read, 14 write, 6 destructive | 0 | `api_key`, `oauth2` | 28/62 (45.2%) | 57/62 (91.9%) |
| [Figma](./figma.md) | `figma` | 49 | 38 read, 8 write, 3 destructive | 0 | `oauth2`, `api_key` | 38/57 (66.7%) | 53/57 (93.0%) |
| [Freshdesk](./freshdesk.md) | `freshdesk` | 39 | 18 read, 13 write, 8 destructive | 0 | `api_key` | 29/79 (36.7%) | 65/79 (82.3%) |
| [Front](./front.md) | `front` | 39 | 15 read, 19 write, 5 destructive | 0 | `api_key` | 28/55 (50.9%) | 51/55 (92.7%) |
| [GitHub](./github.md) | `github` | 36 | 19 read, 14 write, 3 destructive | 3 | `oauth2`, `api_key` | 14/30 (46.7%) | 24/30 (80.0%) |
| [GitLab](./gitlab.md) | `gitlab` | 36 | 20 read, 13 write, 3 destructive | 0 | `oauth2`, `api_key` | 25/51 (49.0%) | 44/51 (86.3%) |
| [Gmail](./gmail.md) | `gmail` | 28 | 9 read, 12 write, 7 destructive | 1 | `oauth2` | 6/27 (22.2%) | 16/27 (59.3%) |
| [Google Calendar](./google_calendar.md) | `google_calendar` | 26 | 10 read, 12 write, 4 destructive | 1 | `oauth2` | 13/40 (32.5%) | 29/40 (72.5%) |
| [Google Docs](./google_docs.md) | `google_docs` | 34 | 1 read, 27 write, 6 destructive | 0 | `oauth2` | 40/54 (74.1%) | 51/54 (94.4%) |
| [Google Drive](./google_drive.md) | `google_drive` | 27 | 10 read, 10 write, 7 destructive | 1 | `oauth2` | 22/48 (45.8%) | 46/48 (95.8%) |
| [Google Maps Platform](./google_maps.md) | `google_maps` | 13 | 13 read | 0 | `api_key` | 17/26 (65.4%) | 21/26 (80.8%) |
| [Google Sheets](./google_sheets.md) | `google_sheets` | 31 | 2 read, 23 write, 6 destructive | 0 | `oauth2` | 34/47 (72.3%) | 46/47 (97.9%) |
| [HubSpot](./hubspot.md) | `hubspot` | 34 | 13 read, 16 write, 5 destructive | 0 | `oauth2` | 19/50 (38.0%) | 44/50 (88.0%) |
| [Intercom](./intercom.md) | `intercom` | 40 | 17 read, 19 write, 4 destructive | 0 | `oauth2`, `api_key` | 44/79 (55.7%) | 72/79 (91.1%) |
| [Jira](./jira.md) | `jira` | 31 | 14 read, 13 write, 4 destructive | 0 | `oauth2` | 22/65 (33.8%) | 48/65 (73.8%) |
| [Linear](./linear.md) | `linear` | 38 | 17 read, 18 write, 3 destructive | 0 | `oauth2`, `api_key` | 28/56 (50.0%) | 47/56 (83.9%) |
| [Mailchimp](./mailchimp.md) | `mailchimp` | 39 | 14 read, 18 write, 7 destructive | 0 | `oauth2`, `api_key` | 43/72 (59.7%) | 68/72 (94.4%) |
| [Mailgun](./mailgun.md) | `mailgun` | 37 | 16 read, 14 write, 7 destructive | 0 | `api_key` | 37/74 (50.0%) | 59/74 (79.7%) |
| [Microsoft Outlook](./microsoft_outlook.md) | `microsoft_outlook` | 37 | 11 read, 19 write, 7 destructive | 0 | `oauth2` | 36/57 (63.2%) | 53/57 (93.0%) |
| [Microsoft Teams](./microsoft_teams.md) | `microsoft_teams` | 38 | 19 read, 12 write, 7 destructive | 0 | `oauth2` | 39/59 (66.1%) | 57/59 (96.6%) |
| [monday.com](./monday.md) | `monday` | 32 | 13 read, 14 write, 5 destructive | 0 | `oauth2`, `api_key` | 23/64 (35.9%) | 53/64 (82.8%) |
| [Notion](./notion.md) | `notion` | 31 | 12 read, 17 write, 2 destructive | 0 | `oauth2`, `api_key` | 19/49 (38.8%) | 40/49 (81.6%) |
| [OpenAI](./openai.md) | `openai` | 32 | 13 read, 13 write, 6 destructive | 0 | `api_key` | 35/54 (64.8%) | 45/54 (83.3%) |
| [Ory Hydra](./ory_hydra.md) | `ory_hydra` | 36 | 17 read, 18 write, 1 destructive | 0 | `api_key` | 35/62 (56.5%) | 58/62 (93.5%) |
| [Ory Identities](./ory_kratos.md) | `ory_kratos` | 49 | 35 read, 13 write, 1 destructive | 0 | `api_key` | 51/63 (81.0%) | 63/63 (100.0%) |
| [PagerDuty](./pagerduty.md) | `pagerduty` | 35 | 14 read, 15 write, 6 destructive | 0 | `api_key` | 35/48 (72.9%) | 46/48 (95.8%) |
| [PayPal](./paypal.md) | `paypal` | 38 | 14 read, 16 write, 8 destructive | 0 | `oauth2`, `api_key` | 36/51 (70.6%) | 49/51 (96.1%) |
| [Pipedrive](./pipedrive.md) | `pipedrive` | 39 | 20 read, 12 write, 7 destructive | 0 | `api_key` | 35/78 (44.9%) | 66/78 (84.6%) |
| [Plaid](./plaid.md) | `plaid` | 28 | 16 read, 10 write, 2 destructive | 0 | `api_key` | 29/42 (69.0%) | 39/42 (92.9%) |
| [Qdrant](./qdrant.md) | `qdrant` | 65 | 17 read, 36 write, 12 destructive | 0 | `api_key` | 65/83 (78.3%) | 81/83 (97.6%) |
| [QuickBooks](./quickbooks.md) | `quickbooks` | 39 | 19 read, 14 write, 6 destructive | 0 | `oauth2` | 40/53 (75.5%) | 46/53 (86.8%) |
| [Salesforce](./salesforce.md) | `salesforce` | 40 | 17 read, 17 write, 6 destructive | 0 | `oauth2` | 28/60 (46.7%) | 51/60 (85.0%) |
| [Segment](./segment.md) | `segment` | 35 | 16 read, 11 write, 8 destructive | 0 | `api_key` | 36/47 (76.6%) | 47/47 (100.0%) |
| [SendGrid](./sendgrid.md) | `sendgrid` | 40 | 17 read, 16 write, 7 destructive | 0 | `api_key` | 38/80 (47.5%) | 56/80 (70.0%) |
| [SendGrid Suppressions](./sendgrid_suppressions.md) | `sendgrid_suppressions` | 22 | 18 read, 4 write | 0 | `api_key` | 22/34 (64.7%) | 31/34 (91.2%) |
| [Shopify](./shopify.md) | `shopify` | 39 | 16 read, 18 write, 5 destructive | 0 | `api_key` | 37/72 (51.4%) | 69/72 (95.8%) |
| [Slack](./slack.md) | `slack` | 35 | 15 read, 16 write, 4 destructive | 2 | `oauth2` | 6/33 (18.2%) | 23/33 (69.7%) |
| [Spotify](./spotify.md) | `spotify` | 40 | 18 read, 17 write, 5 destructive | 0 | `oauth2` | 25/58 (43.1%) | 47/58 (81.0%) |
| [Square](./square.md) | `square` | 34 | 16 read, 12 write, 6 destructive | 0 | `api_key` | 21/45 (46.7%) | 39/45 (86.7%) |
| [Stripe](./stripe.md) | `stripe` | 31 | 13 read, 12 write, 6 destructive | 0 | `api_key` | 29/46 (63.0%) | 40/46 (87.0%) |
| [Telegram](./telegram.md) | `telegram` | 28 | 6 read, 20 write, 2 destructive | 0 | `api_key` | 19/40 (47.5%) | 31/40 (77.5%) |
| [Todoist](./todoist.md) | `todoist` | 29 | 12 read, 12 write, 5 destructive | 0 | `oauth2`, `api_key` | 7/52 (13.5%) | 30/52 (57.7%) |
| [Trello](./trello.md) | `trello` | 40 | 11 read, 21 write, 8 destructive | 0 | `api_key` | 25/73 (34.2%) | 55/73 (75.3%) |
| [Twilio](./twilio.md) | `twilio` | 34 | 16 read, 9 write, 9 destructive | 0 | `api_key` | 45/80 (56.2%) | 60/80 (75.0%) |
| [Typeform](./typeform.md) | `typeform` | 28 | 14 read, 8 write, 6 destructive | 0 | `api_key` | 26/56 (46.4%) | 51/56 (91.1%) |
| [Typesense](./typesense.md) | `typesense` | 69 | 33 read, 23 write, 13 destructive | 0 | `api_key` | 61/81 (75.3%) | 80/81 (98.8%) |
| [Webflow](./webflow.md) | `webflow` | 30 | 13 read, 11 write, 6 destructive | 0 | `api_key` | 24/42 (57.1%) | 42/42 (100.0%) |
| [Xero](./xero.md) | `xero` | 38 | 19 read, 12 write, 7 destructive | 0 | `oauth2` | 41/61 (67.2%) | 56/61 (91.8%) |
| [YouTube](./youtube.md) | `youtube` | 34 | 16 read, 13 write, 5 destructive | 0 | `oauth2` | 21/49 (42.9%) | 38/49 (77.6%) |
| [Zendesk](./zendesk.md) | `zendesk` | 34 | 19 read, 11 write, 4 destructive | 0 | `oauth2` | 30/52 (57.7%) | 42/52 (80.8%) |
| [Zoom](./zoom.md) | `zoom` | 31 | 15 read, 12 write, 4 destructive | 0 | `oauth2` | 36/77 (46.8%) | 54/77 (70.1%) |

## Measured routing accuracy

3652 golden cases over the whole catalog, measured over corpus `ea4f12ad2948` (65 toolkits, 2283 tools indexed and 13 declared uncallable). Each toolkit's own row is on its page, and the rows add up to this one: every case belongs to exactly one toolkit.

| Split | Cases | top-1 | top-8 |
| --- | --- | --- | --- |
| `all` | 3652 | 1937/3652 (53.0%) | 3155/3652 (86.4%) |
| `hand` | 2217 | 1379/2217 (62.2%) | 2099/2217 (94.7%) |
| `paraphrase` | 1412 | 542/1412 (38.4%) | 1035/1412 (73.3%) |
| `paraphrase-clean` | 1132 | 398/1132 (35.2%) | 806/1132 (71.2%) |
| `paraphrase-tuned` | 280 | 144/280 (51.4%) | 229/280 (81.8%) |
| `context` | 23 | 16/23 (69.6%) | 21/23 (91.3%) |

`hand` cases are written with the toolkit; `paraphrase` cases are written by a second author to test whether the tuning generalizes. `paraphrase-clean` is the half whose wording shares no four-word run with a retrieval phrase indexed for its own gold tool, so it is the generalization number and the only paraphrase row to quote as one. `context` cases name no app at all and are decided by the session's recent turns.

The sweep is offline: the reranker is a deterministic identity fake that returns candidates in the order retrieval produced them, so top-1 measures retrieval order rather than a reranked slate. `just eval-live` measures the same cases through the live reranker.

## Coverage caveat

Every toolkit here is authored and the catalog clears the routing eval gate. None has been smoke-tested against the live provider API yet, because no OAuth applications are registered: the system runs against fake providers. Read the tool contracts as the shape automaton will send, not as a verified round trip. Two known gaps: a write call that needs a multipart body cannot be expressed in the current request template, and a tool that needs one is declared uncallable rather than left to fail, so it is indexed nowhere and the row above scores its toolkit over the rest; and toolkits promoted from an OpenAPI import carry mechanical descriptions until they are tuned.
