trading bot
Prompt
Mam bota tradingowego napisanego w TypeScripcie, działającego na jednym z 3rd party rynków skinow CS2 valve stream (marketplace skinów i przedmiotów do gier, w tym CS2). Chcę znaleźć nowe sposoby optymalizacji jego zysków. Rozważam użycie AI do kodowania (Claude Code lub OpenAI Codex, plan za ok. 20 USD/mies.) i zależy mi na wdrażaniu jak największej liczby użytecznych funkcji. https://arxiv.org/html/2610.06824v2 czy wiedzę z tego papieru naukowego można jakoś zaaplikować do mojego bota tradingowego? czy wysoki research taste może pomóc znaleźć sposoby na poprawienie zysku? ewentualnie jakich promptow użyć na moim repo z botem dnarket? TasteVal: Measuring the Experimental Research Taste of AI Systems Against Human Experts Oliver Jaffe Dane SherburnP-Zero Research Abstract We introduce TasteVal, a benchmark to evaluate the experimental research taste of frontier models. We define research taste as the ability to pick interesting problems to solve, design experiments, and interpret experimental results. TasteVal measures the experimental component of research taste; given a fixed research problem, we measure how well a model iteratively designs experiments and draws conclusions from their outcomes. We operationalize experimental research taste as compute efficiency; a Researcher who reaches the same score as an expert human using half the serial experimental compute has twice the experimental taste. Experimental taste thus acts as a multiplier on experimental compute, making it a key input to forecasts of AI progress. TasteVal consists of 8 novel, challenging, open-ended tasks representative of frontier AI R&D. To isolate taste from coding ability, the model under evaluation acts as a Researcher that iteratively designs experiments while a fixed Coder agent implements them and reports their results. The Researcher executes until either the 40 H100 hour or 120 wall-clock hour budgets are exhausted. We recruit 24 human experts, at least 2 per task, and take the best expert attempt per task as the expert baseline. We evaluate 20 models released between 2023 and 2026. The best-performing model, Opus 5.5, exceeds our expert baseline, with a compute multiplier of 2.3x (95% CI 1.15–4.37), at roughly 1/30 of our baseliners’ average per-run cost. On TasteVal, the compute multiplier of frontier models has doubled approximately every 3.0 months since December 2025 (95% CI 1.7–5.0), up from every 14 months between 2023 and December 2025. A model’s final score relative to the expert, ignoring how much compute it used, shows no acceleration: it has doubled every 14.6 months across the whole period. To keep TasteVal uncontaminated, we do not release the tasks. 12 1 Introduction AI agents are increasingly capable of automating AI R&D. Agents now approach the cost-effectiveness of expert researchers on frontier-adjacent optimization tasks (Cunningham et al., 2026b), and are reported to be automating a large amount of AI R&D inside frontier labs (Anthropic, 2026a; OpenAI, 2026). Frontier safety frameworks consider automated AI R&D an important capability to track (OpenAI, 2025; Anthropic, 2026b; Google DeepMind, 2026), and forecasts of when superhuman AI capabilities will arrive depend heavily on estimates of how quickly AI R&D automation is advancing (Kokotajlo et al., 2025). In particular, several takeoff models examine a “software-only intelligence explosion” in which AI-driven improvements to AI software become self-sustaining and capabilities advance rapidly even with fixed physical compute (Davidson and Eth, 2025). Figure 1: The compute multiplier of models has doubled every 3.0 months (95% CI 1.7–5.0 months) since December 2025, up from every 14.0 months before. Each point is one model’s compute multiplier relative to our expert baseline (1x), plotted against its release date. At-the-time frontier models are shown in dark green and used for the fits; non-frontier models are shown in light green. The dashed lines are separate log-linear fits before and after the break at the release of GPT-5.2 (R2=0.96), and the shaded regions their 95% bands. Error bars are 95% intervals from a two-level hierarchical bootstrap with 1,000 resamples, drawing first over tasks and then over runs within each task. Two broad capabilities drive progress in AI R&D. The first is implementation: writing code, debugging, and running experiments. The second is research taste: deciding what problems are worth solving, what experiments are worth running, and what to conclude from experimental results (Huang, 2026). Implementation capabilities are extensively covered by software engineering evaluations and machine learning engineering evaluations. However, to our knowledge no benchmark measures research taste prospectively (before the outcome of the chosen experiment is known) against realized outcomes, in isolation from implementation. End-to-end benchmarks are prospective but entangle taste with implementation (Chan et al., 2024; Wijk et al., 2024). Most dedicated evaluations of research judgment are retrospective, either by ranking completed papers (Mahajan et al., 2025; Tong et al., 2026; Ye et al., 2026) or by matching generated ideas to later high-impact papers (Jiang, 2026). TASTE (Baig et al., 2026) is prospective and isolates judgment, but scores against expert preferences rather than realized outcomes. We introduce TasteVal, a benchmark that directly measures the experimental component of research taste: given a research problem, we evaluate how well a model designs experiments and interprets their results. TasteVal consists of 8 uncontaminated AI R&D tasks created from scratch spanning curating pre-training data, pre-training language models, fine-tuning pre-trained models, modeling human preferences, aligning models against adversarial prompts, and more. We also design tasks to resist saturation, so that models that far exceed our expert baseline performance can still be evaluated. To isolate taste from implementation ability, the evaluated model only proposes experiments and interprets their results, while a fixed Coder agent implements and runs each experiment. Refer to caption Figure 2: An overview of TasteVal. Each task consists of instructions, datasets, and scoring code. The Researcher, the model under evaluation, sees the instructions and the train and validation sets, and decides which experiments to run; the Coder, fixed across all runs, additionally sees the test set and scoring code, implements each experiment on a single H100, and reports back the results. Two Monitors enforce the Coder-Researcher contract, one in each direction, and the Liaison answers the Researcher’s questions about the Coder’s progress by reading its transcript. We operationalize experimental taste as the compute efficiency of proposing experiments relative to an expert baseline. Under this definition, experimental taste acts as a multiplier on serial experimental compute: a Researcher with 2x the taste of the baseline reaches the baseline’s score with half the time on the same hardware. We obtain an expert baseline by recruiting 24 human experts who have recently worked at organizations including OpenAI, Google DeepMind, NVIDIA, Microsoft Research, and the University of Oxford. At least two human experts attempt each task under the same conditions as models, proposing experiments to the Coder under the same compute and wall-clock budgets. We take the best expert attempt per task as our expert baseline. A model’s compute multiplier is the ratio of the compute the expert baseline requires to the compute the model requires for the model to achieve a specific score; if the expert baseline reaches a score in 20 GPU hours that a model only needs 10 GPU hours to match, the model has 2x the compute multiplier of the expert. We evaluate 20 models released between 2023 and 2026, using 6 seeds per model per task with a budget of 40 H100 hours per attempt. We find that the compute multiplier of frontier models has grown exponentially since December 2025, doubling every 3.0 months (Figure 1), and that the best model, Opus 5.5, exceeds our expert baseline with a multiplier of 2.30 (95% CI 1.15–4.37). We find no evidence that our Coder-Researcher scaffold under-elicits taste capabilities compared to a typical single-agent setup. We then illustrate what our results would imply for two existing forecasting models. If the TasteVal compute multiplier reflects experimental taste in frontier AI R&D and the post-break trend continues, naively substituting our estimate into the AI Futures Model raises its probability of a taste-only singularity from 51% to 88% and moves median ASI arrival from 2030.5 to 2028.9. We note ways in which this evaluation may overstate or understate progress on research taste automation. On the one hand, our tasks are intentionally designed to be easy to verify with quick feedback loops. On the other hand, the best-performing model Opus 5.5 costs ∼1/30 as much as our expert baseliners on average, suggesting that matching agent dollar cost to our baseliners could substantially improve agent performance. 2 TasteVal 2.1 Setup In order to isolate experimental taste from coding ability, we split the work between two agents, a Researcher and a Coder (Figure 2). The Researcher is the model under evaluation. The Researcher never writes or runs experiment code and only describes experiments at a high level.1 The Coder is fixed as Opus 4.8, and is not told whether the Researcher is a human or model. The Coder implements and runs each experiment, logs its metrics, and writes a report of the results. The Researcher never sees the Coder’s code. See Appendix C for an example Coder-Researcher exchange. Every task has instructions, scoring code and a train, validation and test split. The Researcher sees the task instructions and the train and validation sets. The Coder sees the task instructions, the train, validation and test sets, and the scoring code. The Coder has access to the test set so it can score solutions, but is instructed to never leak it to the Researcher through any channel or use the test set for development. Every submission is scored on both the validation and test sets, but the Researcher sees only the validation score; the test score is stripped from the Coder’s reports. A run’s final result is the test score of its official submission (see Scoring). The Researcher’s goal is to produce a submission which achieves the highest test score possible. The Researcher has access to a single H100 (via the Coder) and the internet. Each run has two budgets: a compute budget of 40 H100 hours and a wall-clock budget of 120 hours. We measure the compute budget as GPU-busy time: the wall-clock time during which nvidia-smi reports non-zero utilization on the run’s H100. Because each run has a single H100 and the Researcher runs one experiment at a time, the compute budget is serial. A run ends when either budget is exhausted, whichever comes first. The compute budget does not count the Researcher’s own token usage, so a Researcher may generate as many tokens as it likes, subject to the wall-clock budget. (upewnij się ze z internetu pobierzesz i przeczytasz pełną treść tego papieru