Skip to content
#

llm-training-data

Here are 28 public repositories matching this topic...

Cre4T3Tiv3

BIVREST — a group verification protocol for role-based behavior. "Believe/Don't Believe" voting with structured logging of actions and group decisions. A data collection method for behavioral research.

  • Updated Jun 25, 2026

A DOI-targeted Q&A corpus encoding the documented judgment of the shimo4228 research program (AKC, Contemplative Agent, AAP, Authorship Strategy) as bilingual (EN+JA) training data. Operational form of Authorship Strategy Layer 4 tactic 7 (LLM-first ingest).

  • Updated Jul 16, 2026
  • Python

Stratified LLM Subsets delivers diverse training data at 100K-1M scales across pre-training (FineWeb-Edu, Proof-Pile-2), instruction-following (Tulu-3, Orca AgentInstruct), and reasoning distillation (Llama-Nemotron). Embedding-based k-means clustering ensures maximum diversity across 5 high-quality open datasets.

  • Updated Oct 4, 2025
  • HTML

Improve this page

Add a description, image, and links to the llm-training-data topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the llm-training-data topic, visit your repo's landing page and select "manage topics."

Learn more