The Archive as AI Context Folder
Is there anyone here using The Archive as the local file storage and organisation tool that allows their local files to be queried by an AI harness, such as Claude Code or ChatGPT Desktop?
If so, is that working for you? Are there any cautionary aspects of the setup of the archive for this use case that you have identified? For example, linking causing potential issues with the queryable database.
Howdy, Stranger!
Categories
- 3K All Categories
- 153 Research & Reading
- 696 The Zettelkasten Method
- 10 Knowledge Work
- 100 Writing
- 467 Software & Gadgets
- 154 Workflows
- 731 The Archive
- 15 Plug-In Showcase
- 88 Resolved Issues
- 225 Projects Logs and Journals
- 83 Project: Zettelkasten.de
- 53 Critique my Zettel
- 173 Random
- 374 Introduce Yourselves!

Comments
Related: Introducing 'ta', a The Archive-compatible Zettelkasten exploration tool for coding agents
@jameslongley, I have implemented ta and Codex as a local AI digester of my archive. I've not been using it much after install and setup. There is one output at Introducing 'ta', a The Archive-compatible Zettelkasten exploration tool for coding agents. I was inspired by the output and have also queried Codex with the prompt: "Research my notes, writing in detail and depth about savoring life's experiences and how this relates to flourishing." Also inspiring to see the depth of the connections.
One practical use of this marriage of Codex and ta is that I've been able to search my archive for all the books I’ve read with titles or authors with reference to Japanese culture. It found a list of 36 books. I was able to print this and share it with a friend who is a Japanese math professor. Including the extremely relevant and recommended book, The Housekeeper and the Professor — Yōko Ogawa.
Will Simpson
My peak cognition is behind me. One day soon, I will read my last book, write my last note, eat my last meal, and kiss my sweetie for the last time.
My Internet Home — My Now Page
I'm using another software (Obsidian), but I think some of the issues should be similar in The Archive.
Searching your Zettelkasten (
ta)It took me a while to understand that an LLM can't process all files at once. The LLM's context window has a limited size. It cannot hold the entire Zettelkasten, it can only hold selected data.
One challenge is to design a process that filters out just the right amount of relevant data, so that the LLM can do its magic.
The bottleneck isn't AI, it's search.
In theory a RAG setup should work, where you use a separate tool that walks through your entire zettelkasten to build a specialized database. The LLM would talk with that tool. In practice I wasn't able to set it up reliably or get meaningful results.
I also tried agentic workflows that directly accesses my Markdown files. The agent builds custom commands that are supposed to find the right stuff. It works for simple full-text search, but doesn't recognize more advanced content.
That's where an app-specific command line tool like
tacomes in. It allows for a much more targeted search. It makes the LLM happy, because the results are smaller and follow-up queries are easier to ask.Understanding your Zettelkasten (Skills)
However,
taknows only The Archive, it doesn't know your personal setup.For example, do you strictly separate atomic notes and structure notes? If yes, how would you explain the different to your agent? How do you expect the agent to understand the difference? How do you expect the agent to interpret structure notes? Should a nested list imply a hierarchical relationship between the linked notes?
I struggled most with writing skills that correctly describe my conventions. Over the years I made many design decisions that work well for my human brain. I can recognize them instantly. But they are surprisingly difficult to describe as a skill that produces reliable results.
Skill writing is where I gave up. I'm bad at it and I don't get useful results. :-)
I rediscovered that I like to interact directly with my Zettelkasten. The back and forth during manual search is where I find zettels that prompt thinking or need some fine-tuning.
Recommendations
1. Be clear about the goals of AI. What do you expect to learn from AI? Do you just need a more convenient search? Or do you have another workflow in mind?
2. Revisit the content of your ZK. What conventions do you have? Do you follow them consistenly? Could you explain them to another person? Could you write them down? Could you describe them as a skill?
3. Optimize the content for the tools. Use syntax that
tarecognizes. Use syntax that an agent can easily query withgreporrg. If you use other syntax, make sure that you provide your agent with a tool that recognizes it.Thanks @Will. So would you say it's essentially been a curiosity project and a "nice to have" in your implementation to date?
Some great insights here, thanks. I continue to find this aspect of LLMs an irritation - even when they have large context windows they often seem to be built in ways that prevent them querying the full data set available.
I think the context size limitation is a feature of the chatbot interface, not a basic LLM limitation. I envision a specialized LLM that gets trained on one's ZK. Then a chatbot interface could be added.
I can't experiment with this idea using open-source LLMs because the heavy computation makes my laptop CPUs jump up to or beyond their maximum safe temperatures (I also suspect I would need much higher computing power).
If there were such a beast, a challenge would be to prevent it from adapting too closely to the practitioner, at least if one hoped to get surprises from it.
@jameslongley, my experience is evolving. I love computers and programming. I love The Archive. When @ctietze announces a new capacity, I’m excited and all in. So far I've thrown five queries at it. The latest one queried for the authors of books that I've read and ranked them by the number of reads. This is in response to a friend's question about my most frequent authors. It reminds me that there are some great authors that I should read more of. See a fragment of the results below.
Author Summary
ta is used to limit the scope of what the agent can do based on your query. For example, you can ask "Use my notes for ENGL501 to answer the question of how this is a form of literature compared to ENGL4730."
Student Note Results
Assuming “this” refers to the food writing studied in ENGL501, and “ENGL4730” means ENGL473:
Food writing is a form of literature because food serves as more than a subject. Like the Northwestern literature studied in ENGL473, it uses a concrete feature of human life to investigate larger questions of identity, culture, memory, place, power, and moral responsibility. ENGL473 begins with landscape: its readings ask how people are shaped by the Northwest and how they understand their relationships with land, animals, communities, and regional histories. ENGL501 begins with food, but your notes show that it reaches similarly far. “When we write about food,” one note says, “we write about class struggle”—who eats, who does not, and why. Food writing can therefore become memoir, cultural criticism, journalism, philosophy, political protest, or environmental advocacy.
The courses differ chiefly in their organizing lens. In ENGL473, place and landscape connect the works. The readings use such things as salmon, forests, distance, journeys, and encounters with animals to explore ecological identity and the stories people inhabit. In ENGL501, meals, recipes, hunger, kitchens, and taste perform the same literary work. A meal can reveal a family history; a recipe can carry a culture; a food shortage can expose political conditions; and eating a lobster can raise questions about animal suffering. In both courses, an ordinary material thing becomes a way into an entire moral and imaginative world.
ENGL501 also demonstrates its literary status through its range of forms. The syllabus includes memoir, autobiography, poetry, essays, journalism, experimental writing, cookbooks, and literary criticism—from Hemingway and Orwell to Gertrude Stein, M. F. K. Fisher, and David Foster Wallace. These works do not simply provide information about food or instructions for preparing it. They use voice, scene, metaphor, structure, sensory language, and narrative to change how readers perceive human experience.
Thus, food writing relates to ENGL473’s environmental and regional writing almost as table relates to landscape. Both begin with the physical world, but neither remains there. ENGL473 asks what our relationship with place reveals about us; ENGL501 asks what our relationship with food reveals about us. Both become literature when the subject is transformed into an instrument for examining how people live, remember, belong, exploit, care, and make meaning.
Sources consulted: “U-ENGL501 Traditions of Food Writing” (§202211270741), “U-ENGL473” (§202012212010), “What We Write About When We Write About Food” (§202301241914), and “The Philosophy of Food” (§202302211651).
Will Simpson
My peak cognition is behind me. One day soon, I will read my last book, write my last note, eat my last meal, and kiss my sweetie for the last time.
My Internet Home — My Now Page
@Will Can you tell us a bit more about the notes, skills and prompts used in the examples? How does the AI know that you read a book? How extensive are notes like "ENGL501"?
In the Author Summary example, the notes are my reading goals for the year, each year, since 2012. It is my habit to track my reading goals. The prompt I used was: "Look at all the Bookography notes and list all the Author/book titles where the author has been read two or more times. Group them by the number of reads, the author, and the book title. Sort them Alphabetically."
In the Student Notes Example, I referenced class notes that were tagged with the class name. ENGL501 and ENG L4730. One note per class session. 93 notes in total. Here is the prompt I used: "Use my student notes for ENGL501 to answer the question of how this is a form literature compared to ENGL4730."
Unless you are a compulsive life tracker like I am, students would benefit more from this as a way to question their notes and support their learning. Something I missed out on during my early college years.
Will Simpson
My peak cognition is behind me. One day soon, I will read my last book, write my last note, eat my last meal, and kiss my sweetie for the last time.
My Internet Home — My Now Page
I just asked Codex, "How many books have you read?" and it told me, "You have recorded 885 book reads through July 2026." Easy peasy!
This ta/codex is the cat's meow. It is only the tip of the proverbial iceberg. It opens a new way of seeing into my notes.
Will Simpson
My peak cognition is behind me. One day soon, I will read my last book, write my last note, eat my last meal, and kiss my sweetie for the last time.
My Internet Home — My Now Page
One solution is to teach AI to perform properly search your ZK (if your Zettelkasten is build properly).
I am a Zettler
Do you have real life examples of what has worked for you? What prompts and skills did you use? How did the relevant parts of the relevant zettels look like, so that they could be found? Any other experiences with your AI setup that might help your readers?
tabasically wraps a search in structured output.The LLM needs to come up with queries and searches. The searching on the file system alone would work; but with structured output, you get a list of e.g. all links and backlinks for a found note without the agent having to find a matching file, then grep its content for links, then grep its content or filename for an ID, then resolve outgoing links and search for incoming links, etc.
Agents with mediocre models can do this, but it takes longer and is not token-efficient and can lead to failures when the agent tries to converge on some result ('laziness')
The structured output makes everything a "find good matches" call, including a depth-of-links traversal, and the results contains matches plus their related notes in one go to reduce round-trips.
The tool encodes the convention of linking of The Archive, of course, but if you need changes, point your agent to the codebase and tell it to create a fork with your own conventions.
Author at Zettelkasten.de • https://christiantietze.de/
I hadn't thought of using AI for writing custom tools in this way. A customized
tashould cover many search-related use cases. Thanks for the tip!