Use cases
These are nine of our favorite ways to use Opyt, written out as the exchange you would actually have with your client. We will keep adding to this page as we and the people using it find more.
Everything on this page works today with the twelve tools the docs describe. Nothing here is a roadmap.
you › I'm about to spend a week building our own RAG evaluation harness. Has anyone already built it, and did anyone I trust call it a dead end?
claude › Built, in part. Ragas covers the scoring half and assumes you already have a labeled set. The paper argues reference-free metrics drift once contexts get long, and the thread you saved says why the author stopped trusting them. The open gap is the labels, and that is where your week should go.
Repos, papers and posts rank in one index, so "someone built it" and "someone credible thinks it fails" arrive in the same list. The check costs one query, before the first commit.
you › My sources have argued about long-context models versus retrieval for a year. Where does that disagreement actually stand now?
claude › Two camps, and the split moved. Nine of your sources argue a long enough context window removes the need for retrieval. Six argue cost per token makes that a benchmark result rather than a production one. Three authors changed side during 2026, and the read hands you the post where each of them turned.
A sitting reads every document on the topic, oldest to newest. That order is what separates one person changing their mind from two people disagreeing, and it is why a reversal shows up as a date.
you › I was heads-down for ten days. What arrived?
claude › Fourteen, and three are worth opening now. A repo implementing x402 settlement, which four of your standing questions reached independently. A paper on who is liable when an agent pays the wrong party, reached by three. A paper on delegated spend limits, reached by two. Separately, four of your people published nine posts on agentic payments while you were out.
Opyt keeps working while you are away, and holds what it finds until you ask. The top of the queue is what several of your standing questions reached independently, so the ranking is already your own judgment.
you › My sources have been on agentic payments for two years. How did that thinking actually develop?
claude › Three phases. Through 2025 the argument was custody: whether an agent can hold a card at all. From March 2026 it moved to settlement, once stablecoin fees dropped under card interchange. Since June it has been liability, and that is where it still sits: who pays when an agent buys the wrong thing.
What comes back is summaries labeled with the dates they cover, so the answer says what changed and when. Each stretch is summarized once ever, so asking again, or asking something else of it, is free.
you › What do the people I trust actually think about letting agents write to production?
claude › Six of your people have written on it, and they agree on one thing: never a standing credential. Four want a human approving each write. Two argue short-lived scoped tokens make that unnecessary, and one of those has run it in production for a year, which is the post to read before you pick a side.
One query returns where your people landed, and who dissents. The split comes from six sources you chose, so the minority view here is somebody you already trust rather than a loud stranger.
you › Whose work do I keep saving without noticing?
claude › Six people, ranked by what you already did. The top one you liked six times and bookmarked twice and never followed. The second wrote three of the papers you saved. The endorsement already happened, one save at a time. One confirmation brings in the whole archive.
Every name shows the count that put it there, so you can check the ranking before you trust it. The evidence is clicks you already made, which is how somebody you forgot about still lands on the list.
you › We are moving our API from per-seat to per-token pricing next quarter. What have my sources said that bears on it?
claude › Eleven things, and they split on one variable: whether a customer's usage tracks their own revenue. Four operators who shipped per-token report support load up and churn down. Two argue the opposite where usage is lumpy, and one of those is a post you saved and forgot. The decision stays yours. The homework ran against people you chose.
Your client reasons over full source text, on the subscription you already pay for. When an answer changes your mind, you can open the exact post that did it.
you › Draft the argument that open-weight models win the enterprise, cited only from what I saved.
claude › Nine citations, every one resolving to a live URL from someone you confirmed. Two places where your store holds nothing are marked as gaps rather than filled in from general knowledge, so you can see what you still have to read.
No claim without a source, and no source you never chose. The gaps stay visible, which is the part a model working from general knowledge will not show you.
you › Ana sent me an Opyt invite. I have to pick an eval setup for retrieval this week. Has she already been through this?
claude › She has. Ana read the framework and its repo back in March, and she saved a thread arguing that reference-free metrics drift once contexts get long, which is the one that bears on your week. All four citations are hers and say so. You hold nothing on this yourself.
A friend's reading becomes searchable the moment you accept. Ana built that store from people she chose, and every hit says it came from hers, so you can weigh it as her judgment and open the source yourself.