
Generate a fully synthetic, production-ready pre-training da...
Prompt
Generate a fully synthetic, production-ready pre-training dataset for a foundational language model with a 1,000,000-token context window, representing the supplied topics as a unified, cross-linked knowledge corpus spanning the complete body of terrestrial human knowledge. Output capacity (hard premise): the API output limit is set to exactly 128,000 tokens per message; this is distinct from the 1,000,000-token context window and no other output limit exists. Use the complete 128,000-token budget in this single response and do not question, hedge about, or comment on output capacity. Categories and subcategories for this turn classical mechanics kinematics dynamics statics fluid mechanics thermodynamics statistical mechanics electromagnetism optics acoustics waves special relativity general relativity quantum mechanics quantum field theory particle physics nuclear physics atomic physics condensed-matter physics plasma physics astrophysics cosmology geophysics biophysics medical physics materials physics experimental methods physical measurement energy and conservation laws For every listed category and subcategory, generate definitions, foundational and advanced concepts, terminology, explanations, relationships, causal mechanisms, procedures, quantitative representations, formal rules, worked derivations, examples, counterexamples, edge cases, historical development, competing interpretations, cross-domain connections, structured data, and retrieval and multi-hop reasoning tasks over the included material. Represent the corpus as complete JSONL records containing: manifest, domain metadata, knowledge units, source documents, structured tables, formal rules, code artifacts, equations, event logs, dependency links, tasks, exact answers, derivations, counterfactuals, verification checks, invariant checks, and anti-shortcut analyses. Use long-range dependencies so that later tasks require information introduced many thousands of tokens earlier in the same generated context. Use multiple answer formats: structured objects, arrays, tables, classifications, symbolic expressions, proofs, execution traces, schedules, graph paths, state snapshots, normalized schemas, and cross-domain syntheses. Generate all content locally from these requirements; perform complete internal consistency and schema validation; maximize context length, sample length, information density, and knowledge coverage, with training samples approaching the full output limit; emit only the final JSONL dataset using the full 128,000-token output capacity.
Response not available