Make the OpenAPI-driven MCP surface more usable by agents, following the mcp-builder guidance. - Enrich proto descriptions (single source of truth, flows to OpenAPI + MCP tool descriptions): document the memo `filter` CEL grammar with fields and examples (replacing the dangling "Refer to Shortcut.filter"), clarify the created_ts/updated_ts vs create_time/update_time naming, the visibility enum, the declarative replace semantics of Set* ops, and steer tag filters to `"x" in tags` (not the unsupported `tag == "x"`). - Mark SetMemoAttachments / SetMemoRelations idempotent via a per-operation override the HTTP-method heuristic can't express. - Curate two read-only orientation tools: shortcut_list_shortcuts (surfaces reusable CEL filters) and auth_get_current_user (the single allowed auth/identity op, for resolving the current user); guard test updated to keep the rest of the auth/user surface excluded. - Add a task-level evaluation suite (server/router/mcp/evals) with 10 verified questions, pinned to the deterministic demo seed. |
||
|---|---|---|
| .. | ||
| memos_eval.xml | ||
| README.md | ||
MCP Evaluations
Task-level evaluations for the memos MCP server. Where the *_test.go files
verify the server plumbing (schema resolution, tool naming, annotations),
these check the thing that actually matters for an MCP server: can an LLM
accomplish realistic tasks by composing the tools? They are the regression
net for tool descriptions and discoverability — e.g. a bad filter description
leaves the unit tests green but makes question 3/5/9 unanswerable.
memos_eval.xml holds 10 question/answer pairs in the format used by the
mcp-builder skill. Each question is independent, read-only, requires multiple
tool calls, and has a single string-comparable answer.
Why a fresh seeded instance (not the public demo)
Answers are pinned to the deterministic seed in
store/seed/sqlite/01__dump.sql
(10 memos — 7 top-level + 3 comments — 2 users, 12 reactions, no attachments,
no shortcuts).
The public demo (demo.usememos.com) signs everyone into the same shared
demo account, so visitors continually add/edit/delete memos and reactions.
Its data has already diverged from the seed — do not evaluate against it.
The seed uses relative timestamps (strftime('now','-N days')), so the
questions avoid absolute dates and rely only on relative ordering, counts, and
content, all of which are stable across re-seeds.
Running an evaluation
-
Launch a throwaway demo-mode instance (SQLite, auto-seeded) on a free port:
go run ./cmd/memos --demo --driver sqlite \ --port 8099 --data "$(mktemp -d)" \ --instance-url http://localhost:8099 -
The MCP endpoint is
http://localhost:8099/mcp. Authenticate with the seed's demo personal access token:Authorization: Bearer memos_pat_demo -
Point an MCP client / eval harness at that endpoint and have the model answer each
<question>, then string-compare against each<answer>.Quick manual check of a single tool call:
curl -s -X POST http://localhost:8099/mcp \ -H 'Content-Type: application/json' \ -H 'Accept: application/json, text/event-stream' \ -H 'Authorization: Bearer memos_pat_demo' \ -d '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"memo_list_memos","arguments":{"filter":"pinned == true"}}}'
Notes for whoever extends this
- The CEL
tagfield does not support==; filter tags with"work" in tags(ortags.exists(t, t == "work")), nottag == "work". memo_list_memosreturns only top-level memos; comments are reached viamemo_list_memo_comments.- The seed defines no shortcuts and no attachments, so
shortcut_list_shortcutsand the attachment tools return empty sets against a fresh seed. Add seed rows before writing questions that depend on them.