
Generate the longest possible, production-ready, fully synth...
Prompt
Generate the longest possible, production-ready, fully synthetic pre-training dataset for a foundational language model with a 1,000,000-token context window Use the complete available 128000-token output budget Generate maximally long, information-dense training samples, preferably with contexts approaching the full output limit, and allocate the available space to complete knowledge structures, explanations, relationships, derivations, procedures, and multi-hop tasks Represent the complete body of terrestrial human knowledge through a unified, cross-linked knowledge corpus covering all of the following categories and their subcategories: in this turn: Foundations of Knowledge ontology epistemology logic reasoning argumentation evidence evaluation scientific method hypothesis formation experimentation measurement units uncertainty statistics probability causality inference classification taxonomy ontology engineering terminology definitions standards data quality reproducibility ethics of knowledge production For every category, generate definitions, foundational concepts, advanced concepts, terminology, relationships, causal mechanisms, procedures, quantitative representations, formal rules, worked derivations, examples, counterexamples, edge cases, historical development, competing interpretations, cross-domain connections, structured data, and tasks requiring retrieval and reasoning over the included material Represent the corpus as complete JSONL records with a manifest, domain metadata, knowledge units, source documents, structured tables, formal rules, code artifacts, equations, event logs, dependency links, tasks, exact answers, derivations, counterfactuals, verification checks, invariant checks, and anti-shortcut analyses Use long-range dependencies so that later tasks require information introduced many thousands of tokens earlier in the same generated context Use multiple answer formats, including structured objects, arrays, tables, classifications, symbolic expressions, proofs, execution traces, schedules, graph paths, state snapshots, normalized schemas, and cross-domain syntheses Generate all content locally from the requirements above, perform complete internal consistency and schema validation, maximize context length and knowledge coverage, and emit only the final JSONL dataset using the full 128000-token output capacity