Let an AI read your Obsidian vault without uploading it anywhere
Searches for Obsidian vault AI come from two different places. One is the wish to stop rereading four years of notes by hand and have something summarise, connect and answer instead. The other is the worry that turning that on means shipping a personal archive to a server. Most of the material on the subject picks one of those and ignores the other: plugin roundups on one side, arguments for keeping AI out of a vault on the other.
The two are separable, and the reason is the file format. What has to be decided is not whether an AI touches the notes, but which of four arrangements is in use, and which notes are in scope for each one. Those are settings, not a philosophy.
The format is what makes this a choice rather than a bargain
A vault is not a database with an export function. Obsidian's own help puts it plainly:
A vault is a folder on your local file system where Obsidian stores your notes. Source: obsidian.md
The product page is equally direct about the format and the location, stating that Obsidian "stores your notes locally as plain text Markdown files", that it "uses open file formats, so you're never locked in", and that notes stay on the device and can be reached "even offline".
Three consequences follow, and they are the whole reason this is easier than it looks. First, anything that can read a folder can read the vault, with no plugin, no API and no export step. A shell, a search index, a script and a coding agent all see the same files. Second, reading is separable from sending. A tool can open every note on the disk and still transmit nothing, because opening a file and making a network request are different operations. Third, the unit of control is the folder. A directory path is a boundary that every tool on the system already understands, which means scope can be enforced by the file layout rather than by trusting a setting inside an application.
That last point is the one worth sitting with. Debates about AI and notes tend to be framed as all or nothing, because the software people reach for first, a plugin inside the app, has the whole vault in scope by default. The folder is a dial, and it was there the entire time.
Four arrangements, and what each one sends
The distinction that matters is where the model runs, not what the interface looks like. Arranged by that, the field is small.
| Arrangement | What leaves the machine | Good for | Cost of being wrong |
|---|---|---|---|
| Plugin calling a hosted model | The text of whatever notes are put in context, plus the question | Long synthesis, good writing quality | Notes sent before the scope is understood |
| Plugin calling a local model | Nothing | Private drafts, journals, client material | Slower, and weaker on long reasoning |
| Agent in a terminal, scoped to a directory | Whatever the agent chooses to read, if the model is hosted | Bulk work across many files, renaming, restructuring | The agent can also write, so edits land in the vault |
| Search and indexing with no model at all | Nothing | Finding the ten notes that matter | Answers nothing, only locates |
The fourth row is the one most write-ups skip, and it is usually the right first move. A large part of what people want from AI over a vault is retrieval: find where this was written about, gather the scattered mentions, show what connects. That is a search problem, and macOS already indexes note contents.
The second row deserves a word of caution against overselling it. A local model sends nothing, which settles the privacy question completely. It also has less capacity than a hosted one at the same moment in time, so the honest way to use it is for the tasks where the input is short and the output is a judgment rather than an essay.
Sync encryption is a different question
One point gets conflated often enough to be worth separating. Obsidian's paid sync service is described as letting notes be reached "on any device, secured with end-to-end encryption", with the accompanying claim that "No one else can read them, not even us". That is a statement about notes in transit and at rest on the sync service, and it is a good reason to trust the sync path.
It says nothing about the AI question. Encryption between two devices does not constrain what a plugin running on one of those devices sends to a third party, because at that moment the plugin holds the decrypted text. The two protections sit on different segments of the journey. A vault can be synced end to end and still be read by a hosted model, and it can be unsynced and still be private. Keeping the two claims apart is what makes it possible to reason about either.
Scope the vault before handing any of it over
The practical move, whichever arrangement is chosen, is to stop treating the vault as one object. Selecting the relevant notes first and passing only those is faster, cheaper, produces better answers because the context is less diluted, and makes the privacy question mostly moot.
The system search does this well and needs no setup, because Spotlight already stores note contents as a metadata attribute. man mdfind documents the interesting part: -onlyin directory limits the search to one directory, and the content attribute is kMDItemTextContent, which is what the -interpret option expands a bare word into. The manual page shows the expansion for the word "search" as (* = search* cdw || kMDItemTextContent = search* cdw), which is the metadata attribute and the content attribute checked together.
So the shape of a scoped query is a directory plus a content match, and the result is a list of paths. -0 prints a NUL after each result, documented as "useful when used in conjunction with xargs -0", which is how the list becomes the input to something else rather than something to read and copy by hand. -count returns the number of matches instead of the paths, which is the cheap way to find out whether a query is too broad before running it for real.
Two habits follow from this. Check the count first and narrow until it is a number worth reading, then pass the paths onward. And keep the hidden configuration folder inside the vault out of scope, because plugin settings and workspace state are not notes and add nothing but noise to a context window.
Plain text search has a place next to the index, too. The index answers questions about words, and it is only as current as the last time it caught up. A recursive text search over the folder is slower but reads the files as they are on disk right now, which matters when the last five minutes of editing is the part in question.
Giving a coding agent a directory instead of a vault
The third arrangement is the one that has changed most recently, and it is worth understanding in file terms rather than in marketing terms. A terminal agent reaches a folder through a local server process, declared in a configuration file, and the folder path in that declaration is the scope. The declaration for a local process names the command to run and the arguments to pass it:
{
"mcpServers": {
"notes": {
"type": "stdio",
"command": "npx",
"args": ["-y", "some-filesystem-server", "/path/to/a/subfolder"]
}
}
}
Three things about that arrangement are easy to miss.
The scope is a path, so the vault's own folder structure becomes the permission model. Pointing the server at one project folder rather than the vault root is a one-word change with a large effect, and it costs nothing in capability for a task that only concerns that project.
The file the declaration lives in decides who else gets it. Configuration written at project level lands in a file at the root of the directory, which is the one intended for version control. Configuration written at the machine level lands in a file in the home directory. A vault is usually personal, so the machine-level file is the right home for this, and a path to a personal notes folder should not end up in a repository that other people clone.
And a file server that can read can usually also write. That is the actual reason to be careful here, and it has nothing to do with privacy. An agent asked to tidy front matter across two hundred notes will do it, including on the notes where it guessed wrong. Committing the vault to version control before any bulk pass is the cheap insurance, and it converts a bad run from a loss into a diff.
What breaks, and what to check before starting
Four failure modes come up repeatedly, and none of them are about the model.
Sync fighting an external writer. A vault that syncs between devices is a folder under continuous observation by another process. A tool rewriting many files at once produces a burst of changes, and on a second device that arrives as a burst of conflicts. Doing bulk work on one machine, with sync paused, and letting it settle before starting the next device is the difference between a clean pass and an afternoon of merge files.
Link syntax that standard tools do not understand. Obsidian's double bracket links and its embeds are not part of standard Markdown. A tool that rewrites a note through a generic Markdown parser can normalise those into something the app no longer resolves. Anything doing bulk edits should be treating notes as text rather than parsing and re-emitting them.
Front matter as a silent casualty. The block at the top of a note is structured data that other things depend on: queries, dashboards, sorting. Rewrites that regenerate a file from its prose lose it. This is the single most common way a helpful cleanup pass becomes a restoration job.
Attachments in the context window. A vault usually holds images and PDFs alongside the notes. A scope that includes the attachments folder feeds binary files to something built for text, which wastes the context and sometimes produces noise that looks like content. Excluding it is a one-line change made once.
None of these are reasons to avoid the arrangement. They are reasons to version the vault first, to scope by folder, and to prefer tools that treat notes as files rather than as documents to be re-rendered. A window that holds the folder tree, the shell and the model in one place makes the check cheap, since the scope being handed over is visible in the same view as the command that hands it over. The Features page sets out what that arrangement covers, and Compared with other file managers sets out where tools differ on it.
What to change first
Put the vault under version control before anything else, because that single step turns every later mistake into something reversible. Then pick one subfolder, not the vault root, and point whichever tool is being trialled at that folder only. Run the retrieval case first, finding the notes rather than summarising them, and add a model to the loop only once the scoping works. Keeping the folder, the shell and the model in one window is what Atriens is built around.
Frequently asked questions
Does using AI with an Obsidian vault mean uploading the notes?
Only in the arrangements that call a hosted model. A vault is a folder of plain text Markdown files on the local disk, so reading it and sending it are separate operations. A local model, or a search index, reads every note and transmits nothing. What decides the answer is where the model runs, not whether AI is involved.
What is the safest first step for a large vault?
Put the vault in version control, then scope to one subfolder rather than the root. Version control makes any bulk edit reversible, which is the real risk with an agent that can write as well as read. Scoping to a subfolder also improves answers, since a smaller and more relevant context beats a larger one.
Is a plugin or a terminal agent the better route?
They suit different work. A plugin is better for asking questions while writing, because it lives where the writing happens. A terminal agent is better for bulk work across many files, such as normalising front matter or restructuring folders, because it can act on a list of paths rather than one open note. Neither replaces the other.
How can the right notes be found before passing anything to a model?
Use the system index. man mdfind documents -onlyin for limiting a search to one directory and -count for returning the number of matches instead of the paths, which is the cheap way to test whether a query is too broad. The content attribute is kMDItemTextContent, and -0 makes the result list usable as input to another command.
Will a bulk edit break Obsidian's links or front matter?
It can. The double bracket link syntax and embeds are not standard Markdown, and front matter is structured data that queries and dashboards depend on. A tool that parses a note and re-emits it can normalise both into something the app no longer resolves. Prefer tools that treat notes as text, and commit before any pass that touches many files.