ContextGem: Effortless LLM extraction from documents
-
Updated
Jul 28, 2026 - Python
ContextGem: Effortless LLM extraction from documents
Karpathy’s LLM Wiki, 100% local with Ollama. Drop Markdown notes → AI extracts concepts → your Obsidian wiki auto-links and grows. Zero sharing. Your notes stay yours.
More than just Karpathy’s LLM Wiki, 100% local with Ollama. Drop Markdown notes → AI extracts concepts → your Obsidian wiki auto-links and grows. Zero sharing. Your notes stay yours.
An agentic extractor for messaging platforms. Noria receives messages & URLs via chat; scrapes web data to evaluate the content against custom scoring matrices (e.g., jobs, scholarships); and sends a structured summary with a calculated match score directly to the assigned messenger number.
Bounded, inspectable LLM inference pipelines from declared YAML — runs offline against Ollama or any OpenAI-compatible local server, emitting JSONL traces, inspect reports, and stable exit codes.
VL-JEPA inspired pipeline — compress images/text locally via Ollama, send compact payloads to any LLM API. Cut token costs by ~80%.
Agent skills that turn Claude Code (or any agent) into an award-winning multi-modal creative pipeline — art, music, video, story. Panel-of-Experts debate, a deterministic state graph, and an immutable continuity block that makes creative drift structurally impossible. Prose contracts, not framework code.
Document Summarization App using large language model (LLM) and Langchain framework. Used a pre-trained T5 model and its tokenizer from Hugging Face Transformers library. Created a summarization pipeline to generate summary using model.
Simplified Technical English for Code - extract, adapt, and enforce ASD-STE100 rules for software documentation. 9-agent pipeline, 59-test benchmark (96.6%), multi-language scaffolding.
A powerful CLI tool for extracting text from documents using DeepSeek OCR and generating high-quality datasets with LLM assistance.
Local-first pipeline for a daily AI news digest
Build a personal LLM wiki using Claude Code with a two-layer cache architecture that integrates with Logseq and Obsidian.
Trade-Alert is a notification system that keeps users updated on critical news impacting their stock portfolios. It simplifies staying informed by delivering timely notifications for important articles, eliminating the need to monitor multiple platforms in today’s fast-paced market.
Self-hosted batch LLM pipeline for analyzing customer feedback from Excel. Upload xlsx, describe the task, configure output fields — get structured results. Works with any OpenAI-compatible API and Ollama.
Long-form novel analysis pipeline with chapter extraction, report export, and quality review
OpenSCAD 3D model generation & refinement pipeline with a silly website for comparing LLMs.
信息高度图名的时代,真正稀缺的不是信息,而是结构化的判断力。每天面对 arXiv、Hacker News、TechCrunch 等渠道的海量信息资源,真正需要的是一个能在信息噪音中提取信号、在碎片中构建体系的洞察引擎。
This Project, I worked on the development of an LLM-powered AI chatbot using Gemini 2.5 Flash, LangChain, Streamlit, and LangSmith observability. While building the system, I analyzed the LLM run logs to better understand how prompts flow through the pipeline and how responses are generated
Podcast 🎧 Search 🔎 platform with semantic text 📃retrieval & LLM 🤖 pipeline for long-form audio transcripts
CLI tool for LLM prompt pipelines. Reusable. Shareable. Scriptable.
Add a description, image, and links to the llm-pipeline topic page so that developers can more easily learn about it.
To associate your repository with the llm-pipeline topic, visit your repo's landing page and select "manage topics."