Skip to content
#

ai-training-data

Here are 6 public repositories matching this topic...

Language: All
Filter by language

🚀 Interactive JSONL editor for Claude Code conversation files with real-time file system synchronization. Efficient prompt engineering through conversation editing.

  • Updated Aug 13, 2025
  • JavaScript

[CURRENT WIP] Production-grade AI training data mining system. Hardware-adaptive architecture scales from Raspberry Pi to GPU workstations. Features advanced web scraping, multi-level deduplication, domain-specific extractors, and enterprise monitoring.

  • Updated Jul 21, 2025
  • Python

🤖 Automated Q&A Dataset Generation Pipeline powered by LLMs. Multi-stage pipeline that searches, filters, extracts and transforms web content into high-quality question-answer datasets for LLM training. Supports multiple LLM providers (Groq, Mistral, Ollama) and search engines.

  • Updated Jun 7, 2025
  • Python

Improve this page

Add a description, image, and links to the ai-training-data topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the ai-training-data topic, visit your repo's landing page and select "manage topics."

Learn more