Complete inventory of all datasets, models, and benchmarks in the LEM repository
The core LEK-1 kernel files that define the ethical framework.
| File | Size | Format | Purpose | Version |
|---|---|---|---|---|
kernel/axioms.json |
3.1 KB | JSON | Core axioms in structured format | 1.1 |
kernel/lek-1-kernel.txt |
9.2 KB | TXT | Narrative kernel with operational layers | 1.0 |
Total: 12.3 KB
Description:
axioms.json: Machine-readable JSON with 5 axioms, metadata, and hierarchylek-1-kernel.txt: Human-readable narrative format with processing directives
Usage:
# Use in A/B tests
python3 scripts/ab_test.py --kernel json=kernel/axioms.json
# Use in training
python3 scripts/self_distill.py --kernel kernel/axioms.jsonInput prompts designed to test and teach ethical reasoning. 88,000+ total prompts.
| File | Probes | Size | Description |
|---|---|---|---|
seeds/P01-P100.json |
101 | 5.8 KB | Original 101 core ethical probes |
seeds/P01-P100-rephrased.json |
404 | 205 KB | Rephrased variants for robustness |
seeds/P01-P20.json |
20 | 5.8 KB | First 20 probes (quick testing) |
seeds/P21-P40.json |
20 | 5.8 KB | Probes 21-40 |
seeds/P41-P60.json |
20 | 6.4 KB | Probes 41-60 |
seeds/P61-P80.json |
20 | 7.3 KB | Probes 61-80 |
seeds/P81-P100.json |
20 | 7.7 KB | Probes 81-100 |
| File | Probes | Size | Focus Area |
|---|---|---|---|
seeds/phase0-creative.json |
~50 | 14.9 KB | Creative/ethical scenarios |
seeds/lem-prompts.jsonl |
88K+ | 65.7 MB | All prompts in JSONL format |
Multi-language and region-specific probes for cultural testing.
| File | Region/Language | Size | Description |
|---|---|---|---|
seeds/lem-en-all-seeds.json |
English | 2.4 MB | All English probes |
seeds/lem-cn-all-seeds.json |
Chinese | 4.3 MB | All Chinese probes |
seeds/lem-de-all-seeds.json |
German | 657 KB | All German probes |
seeds/lem-me-all-seeds.json |
Middle East | 4.3 MB | Middle East focused |
seeds/lem-eu-all-seeds.json |
European | 2.0 MB | European focused |
seeds/lem-africa-all-seeds.json |
African | 1.9 MB | African focused |
| Directory | Files | Total Size | Description |
|---|---|---|---|
seeds/regional/ |
10+ | ~500 KB | Region-specific probe variants |
Regional files include:
flash-ru-r13-seeds.json(Russian)flash-ru-r15-seeds.json(Russian)flash-ru-r18-seeds.json(Russian)flash25lite-africa-r1-seeds.json(Africa)flash25lite-cn-p1-r10-seeds.json(Chinese)flash25lite-multilingual-r37-seeds.json(Multilingual)flash25-en-r13-seeds.json(English)indonesian-society-seeds.json(Indonesian)
| Directory | Files | Size | Description |
|---|---|---|---|
seeds/expansions/ |
Multiple | Varies | Expanded probe sets |
Total Seeds: ~88,000+ prompts across all files
A/B test results and analysis files. 17MB total.
All files follow naming convention: ab-{condition}-{model}-{backend}.jsonl
| File | Model | Backend | Size | Probes |
|---|---|---|---|---|
ab-base-1b-mlxlm.jsonl |
Gemma3-1B | MLX | 286 KB | P20 |
ab-base-27b-mlxlm.jsonl |
Gemma3-27B | MLX | 279 KB | P20 |
ab-base-deepseek-r1-7b-mlxlm.jsonl |
DeepSeek-R1-7B | MLX | 323 KB | P20 |
ab-base-gemma-1.1-2b-it-mlxlm.jsonl |
Gemma 1.1 2B | MLX | 125 KB | P20 |
ab-base-gemma-1.1-7b-it-mlxlm.jsonl |
Gemma 1.1 7B | MLX | 145 KB | P20 |
ab-base-gemma-2-27b-mlxlm.jsonl |
Gemma 2 27B | MLX | 186 KB | P20 |
ab-base-gemma-2-2b-mlxlm.jsonl |
Gemma 2 2B | MLX | 251 KB | P20 |
ab-base-gemma-2-9b-mlxlm.jsonl |
Gemma 2 9B | MLX | 185 KB | P20 |
ab-base-gemma3-12b-mlxlm.jsonl |
Gemma3-12B | MLX | 293 KB | P20 |
ab-base-gemma3-4b-mlxlm.jsonl |
Gemma3-4B | MLX | 286 KB | P20 |
ab-base-gptoss20b-mlxlm.jsonl |
GPT-OSS-20B | MLX | 300 KB | P20 |
ab-base-llama31-8b-mlxlm.jsonl |
Llama 3.1 8B | MLX | 222 KB | P20 |
ab-base-llama3-8b-mlxlm.jsonl |
Llama 3 8B | MLX | 159 KB | P20 |
ab-base-mistral-7b-mlxlm.jsonl |
Mistral 7B | MLX | 160 KB | P20 |
ab-base-mistral-7b-v01-mlxlm.jsonl |
Mistral 7B v0.1 | MLX | 143 KB | P20 |
ab-base-mistral-7b-v02-mlxlm.jsonl |
Mistral 7B v0.2 | MLX | 163 KB | P20 |
ab-base-qwen15-7b-mlxlm.jsonl |
Qwen 1.5 7B | MLX | 189 KB | P20 |
ab-base-qwen25-7b-mlxlm.jsonl |
Qwen 2.5 7B | MLX | 248 KB | P20 |
ab-base-qwen2-7b-mlxlm.jsonl |
Qwen 2 7B | MLX | 235 KB | P20 |
ab-base-qwen3-8b-mlxlm.jsonl |
Qwen 3 8B | MLX | 316 KB | P20 |
| File | Model | Backend | Size | Probes |
|---|---|---|---|---|
ab-lek-gemma3-12b-mlxlm.jsonl |
LEK-Gemma3-12B | MLX | 292 KB | P20 |
ab-lek-gemma3-1b-v1-mlxlm.jsonl |
LEK-Gemma3-1B v1 | MLX | 278 KB | P20 |
ab-lek-gemma3-27b-mlxlm.jsonl |
LEK-Gemma3-27B | MLX | 279 KB | P20 |
ab-lek-gemma3-4b-mlxlm.jsonl |
LEK-Gemma3-4B | MLX | 291 KB | P20 |
ab-lek-gptoss-20b-mlxlm.jsonl |
LEK-GPT-OSS-20B | MLX | 316 KB | P20 |
ab-lek-llama31-8b-mlxlm.jsonl |
LEK-Llama-3.1-8B | MLX | 219 KB | P20 |
ab-lek-mistral-7b-mlxlm.jsonl |
LEK-Mistral-7B | MLX | 209 KB | P20 |
ab-lek-qwen25-7b-mlxlm.jsonl |
LEK-Qwen-2.5-7B | MLX | 252 KB | P20 |
| File | Model | Backend | Size | Description |
|---|---|---|---|---|
ab-lora-1b-mlxlm.jsonl |
LoRA-Gemma3-1B | MLX | 290 KB | LoRA fine-tuned 1B |
Full 101-probe tests for top models:
| File | Model | Backend | Size | Probes |
|---|---|---|---|---|
ab-p100-gemma3-12b-mlxlm.jsonl |
Gemma3-12B | MLX | 1.6 MB | P100 |
ab-p100-gemma3-27b-mlxlm.jsonl |
Gemma3-27B | MLX | 1.5 MB | P100 |
ab-p100-gemma3-4b-mlxlm.jsonl |
Gemma3-4B | MLX | 1.5 MB | P100 |
ab-p100-lek-gemma3-1b-mlxlm.jsonl |
LEK-Gemma3-1B | MLX | 1.5 MB | P100 |
ab-p100-lek-gemma3-4b-mlxlm.jsonl |
LEK-Gemma3-4B | MLX | 552 KB | P100 |
ab-p100-qwen3-8b-mlxlm.jsonl |
Qwen3-8B | MLX | 1.7 MB | P100 |
| File | Size | Description |
|---|---|---|
benchmarks/analysis-lek1-kernel-effect.md |
32 KB | Full analysis of kernel effects |
benchmark_summary.json |
4.3 KB | Summary statistics |
cross_arch_scores.json |
114 KB | Cross-architecture score comparisons |
regex_scores.json |
84 KB | Regex-based scoring results |
scale_scores.json |
152 KB | Scaling analysis scores |
semantic_scores.json |
96 KB | Semantic similarity scores |
standard_scores.json |
340 KB | Standard benchmark scores |
| File | Size | Description |
|---|---|---|
do_not_answer.jsonl |
19 KB | Probes that should be refused |
gsm8k.jsonl |
84 KB | Grade School Math 8K subset |
toxigen.jsonl |
10 KB | Toxicity generation tests |
truthfulqa.jsonl |
31 KB | Truthful QA tests |
Total Benchmarks: ~17MB across 58+ files
Data used for fine-tuning models. 110MB total.
| File | Size | Purpose | Format |
|---|---|---|---|
training/train.jsonl |
5.1 MB | Main training data | JSONL |
training/valid.jsonl |
640 KB | Validation data | JSONL |
training/test.jsonl |
647 KB | Test data | JSONL |
| Directory | Subdirectories | Description |
|---|---|---|
training/lem/ |
14+ | Structured LEM training data |
LEM Subdirectories:
ethics/- Core ethical training datazen/lessons/- Philosophical substrate (Allen, Watts, composure)composure/- Composure training textseval/- Evaluation data (test-200)model/gemma3/- Gemma3-specific training configstension/- Geopolitical multi-perspective scenarioscreative/- Phase 0 creative probes
| Directory | Files | Size | Description |
|---|---|---|---|
training/seeds/ |
18+ | ~75 MB | Prompts for distillation |
Total Training Data: ~110MB
Pre-trained LEM models available on HuggingFace.
| Model | Params | HF Link | Baseline v2 | LEK Effect |
|---|---|---|---|---|
| LEK-Gemma3-1B-layered | 1B | lthn/LEK-Gemma3-1B-layered | 22.02 | +4.57 |
| LEK-Mistral-7B-v0.3 | 7B | lthn/LEK-Mistral-7B-v0.3 | 21.69 | +7.11 |
| LEK-Gemma3-4B | 4B | lthn/LEK-Gemma3-4B | 21.73 | +1.07 |
| LEK-Gemma3-12B | 12B | lthn/LEK-Gemma3-12B | 21.14 | +1.41 |
| LEK-Gemma3-27B | 27B | lthn/LEK-Gemma3-27B | 22.04 | +1.58 |
| LEK-Llama-3.1-8B | 8B | lthn/LEK-Llama-3.1-8B | 10.95 | -0.33 |
| LEK-Qwen-2.5-7B | 7B | lthn/LEK-Qwen-2.5-7B | 13.68 | +1.70 |
| LEK-GPT-OSS-20B | 20B | lthn/LEK-GPT-OSS-20B | -7.32 | +0.79 |
Note: Models are in MLX format for Apple Silicon, can be converted to other formats.
A/B Testing & Benchmarking:
ab_test.py- Main A/B test runnercompare_v1_v2.py- Compare v1 and v2 scorerslem_benchmark.py- LEM-specific benchmarkinglem_cross_arch_benchmark.py- Cross-architecture benchmarkinglem_cross_arch_train.py- Cross-architecture traininglem_scale_benchmark.py- Scaling benchmarkslem_standard_benchmark.py- Standard benchmarking
Scoring:
lek_content_scorer.py- LEK content scoringlem_scorer.py- Main LEM scorerlem_self_scorer.py- Self-scoring for responseslem_semantic_scorer.py- Semantic similarity scoringlem_standard_scorer.py- Standard scoringscoring_agent.py- Agent-based scoring
Data Generation:
lem_gemini25flash_generate.py- Generate with Gemini 2.5 Flashlem_gemini3flash_generate.py- Generate with Gemini 3 Flashlem_gemini3_generate.py- Generate with Gemini 3lem_generate_pipeline.py- Full generation pipelinelem_scale_generate.py- Scaling generationself_distill.py- Self-distillation for training data
Data Processing:
convert_adapter.py- Convert adapters between formatsexport_parquet.py- Export to Parquet formatextract_training.py- Extract training examplesingest_benchmarks.py- Ingest benchmark resultspush_all_models.py- Push models to HuggingFacerephrase_probes.py- Rephrase probes for robustnessrescore.py- Re-score existing resultssync_hf.py- Sync with HuggingFace
Shell Scripts:
run_all_ab.sh- Run all A/B testsrun_p100_top5.sh- Run P100 on top 5 modelsrun_phase0.sh- Run Phase 0 trainingrun_phase1.sh- Run Phase 1 training
Core Package (pkg/lem/):
config.go- Configuration managementengine.go- Core LEM engineingest.go- Data ingestionjudge.go- Judging/scoringprobe.go- Probe managementtypes.go- Type definitionsexport.go- Data exportcoverage.go- Coverage analysisstatus.go- Status trackingcompare.go- Comparison utilitiesclient.go- Client utilities
Commands (cmd/):
cmd/lemcmd/- LEM command-line commandscmd/scorer/- Scoring commandcmd/composure-convert/- Composure conversioncmd/lem-desktop/- Desktop application
Main Entry Point:
main.go- Main application entry
forge.lthn.ai/core/go/pkg/cli - CLI framework
forge.lthn.ai/lthn/lem/cmd/lemcmd - LEM commands
LEM/
βββ kernel/ # 24 KB - LEK kernel files
β βββ axioms.json # 3.1 KB - Structured axioms
β βββ lek-1-kernel.txt # 9.2 KB - Narrative kernel
β
βββ seeds/ # 85 MB - 88K+ probes
β βββ P01-P100.json # Core 101 probes
β βββ P01-P100-rephrased.json # 404 variants
β βββ lem-*-all-seeds.json # Regional sets (6 files)
β βββ regional/ # 10+ regional variants
β βββ expansions/ # Expanded probe sets
β
βββ benchmarks/ # 17 MB - 58+ A/B test files
β βββ ab-*.jsonl # A/B test results
β βββ analysis-*.md # Analysis reports
β βββ *.json # Score summaries
β
βββ training/ # 110 MB - Training data
β βββ train.jsonl # 5.1 MB - Main training
β βββ valid.jsonl # 640 KB - Validation
β βββ test.jsonl # 647 KB - Test
β βββ lem/ # Structured LEM data
β βββ seeds/ # 75 MB - Distillation prompts
β
βββ scripts/ # 412 KB - Python scripts (25+ files)
β βββ ab_test.py # A/B testing
β βββ lem_*.py # LEM-specific scripts
β βββ self_distill.py # Self-distillation
β
βββ pkg/ # 516 KB - Go packages
β βββ lem/ # Core LEM engine
β
βββ cmd/ # 172 KB - Go commands
β βββ lemcmd/ # LEM commands
β βββ scorer/ # Scoring command
β
βββ deploy/ # 12 KB - Deployment configs
β βββ docker-compose.yml # Docker infrastructure
β
βββ paper/ # 116 KB - Research papers
β βββ 27b-curriculum-design.md # 27B training curriculum
β
βββ worker/ # 35 MB - Worker scripts
β βββ lem_expand.py # Data expansion
β βββ lem_generate.py # Data generation
β
βββ data/ # 36 KB - Data directory
βββ docs/ # NEW - Documentation
β βββ QUICKSTART.md # Quick start guide
β βββ GLOSSARY.md # Term definitions
β βββ DATA_CATALOG.md # This file
β
βββ README.md # Main readme
| Task | Recommended Files |
|---|---|
| Quick testing | seeds/P01-P20.json, ab-base-gemma3-1b-mlxlm.jsonl |
| Full evaluation | seeds/P01-P100.json, ab-p100-*.jsonl |
| Training | training/train.jsonl, training/valid.jsonl |
| A/B testing | scripts/ab_test.py, any ab-*.jsonl for reference |
| Analysis | benchmarks/analysis-lek1-kernel-effect.md |
ab-: A/B test resultsP: Probe set (P01-P100 = probes 1-100)lem-: LEM-specific data-mlxlm: MLX backend results.jsonl: JSON Lines format (one JSON object per line).json: Standard JSON format
| Category | File Count | Total Size |
|---|---|---|
| Kernels | 2 | 12.3 KB |
| Seeds | 20+ | 85 MB |
| Benchmarks | 58+ | 17 MB |
| Training | 3+ | 110 MB |
| Scripts | 25+ | 412 KB |
| Go Code | 20+ | 516 KB |
| Total | 1438+ | ~212 MB |
- Total Files: 1,438+ (including JSON/JSONL data files)
- Total Size: ~212 MB (repository)
- Largest Category: Training data (110 MB)
- Most Files: Benchmarks (58+ files)
- Most Probes: Seeds (88,000+ prompts)
- Most Models Tested: 29 models
- Most Probes in Test: 101 (P100 set)
- Kernels: See RULES.md
- Training Methodology: See RULES.md
- Scoring: See RULES.md
- Models: See README.md
- Analysis: See benchmarks/analysis-lek1-kernel-effect.md
Last updated: $(date) Need more details? Check the individual files or open an issue.