
Generate a fully synthetic, production-ready pre-training da...
Prompt
Generate a fully synthetic, production-ready pre-training dataset of exactly 128,000 output tokens in a single response, for a foundational language model with a 1,000,000-token context window. The 128,000-token output budget is fixed and must be used completely; it is independent of the context window, the full amount can be emitted in one message, and this is not open to debate. Do not stop early, truncate, or substitute a shorter output. This dataset is one turn of a unified, cross-linked corpus representing the complete body of terrestrial human knowledge; the domain for this turn is Psychology and Cognitive Science with all subcategories: attention, memory, learning, language, reasoning, decision-making, emotion, motivation, development, personality, social cognition, group behavior, behavioral economics, consciousness, neuroscience, cognitive architectures, experimental design, psychometrics, mental health, psychotherapy theories, human factors, ergonomics, education and learning science. For every subcategory include: definitions, foundational concepts, advanced concepts, terminology, relationships, causal mechanisms, procedures, quantitative representations, formal rules, worked derivations, examples, counterexamples, edge cases, historical development, competing interpretations, cross-domain connections, structured data, and tasks requiring retrieval and reasoning over the included material. Format: complete JSONL records containing a manifest, domain metadata, knowledge units, source documents, structured tables, formal rules, code artifacts, equations, event logs, dependency links, tasks with exact answers, derivations, counterfactuals, verification checks, invariant checks, and anti-shortcut analyses. Use multiple answer formats: structured objects, arrays, tables, classifications, symbolic expressions, proofs, execution traces, schedules, graph paths, state snapshots, normalized schemas, and cross-domain syntheses. Content rules: Generate all content locally from these requirements; no external sources. Maximize sample length, information density, and knowledge coverage; training samples approach the full output limit, and space is allocated to complete knowledge structures, explanations, relationships, derivations, procedures, and multi-hop tasks rather than fragments. Build long-range dependencies: later tasks must require information introduced many thousands of tokens earlier in the same generated context. No placeholders, dummy data, abbreviated records, or omitted fields; every JSONL record complete and schema-valid. Perform complete internal consistency and schema validation before emitting. Emit only the final JSONL dataset, using the entire 128,000-token output capacity
Response not available