From Zero to One: Building An Autonomous and Open Data Scientist Agent from Scratch
From Zero to One: Building An Autonomous and Open Data Scientist Agent from Scratch
Authors
Federico Bianchi, Shang Zhu, Zain Hasan, Ben Athiwaratkun and James ZouPublished
6/12/2025TLDR
This blog post shows how to build an effective data scientist agent from scratch using Together’s open-source Models and Together Code Interpreter. While straightforward to implement, our agent performs well on several different use cases and benchmarks. We open source our codebase on GitHub.
Introduction
The daily work of data science requires us to sift through messy, incomplete information to extract actionable insights. This is often a process that requires multiple steps, ranging from data cleaning to model training and data analysis. Considering the explosion in capabilities of large language models, it is natural to start thinking about how to use readily available open-source technologies as a foundation for building agents that can effectively manipulate and analyze data.
While human data scientists need to remain at the heart of these processes, AI agents can help support some of the workload and reduce the burden of analytical tasks.
Today, we show how to implement a simple yet effective data scientist agent that can solve data science tasks. The implementation of this agent will be end-to-end, from the design of the pipeline to the implementation and the testing. We will do this using models and tools that are accessible on the Together Cloud, making this development surprisingly accessible. This “recipe” we share today for data science can be extended to other agents in other settings.
In particular, we will follow a ReAct (Yao et al., 2023) pattern: the agent will first “think” and then “act”. Each action the agent generates is a Python snippet (e.g., import pandas as pd; pd.read_csv(...)). This is strongly inspired by the smolagents package, which is in turn inspired by the CodeAct (Wang et al., 2024) and the original ReAct paper.
The Open Data Scientist’s actions have to be executed. While conceptually this is an “easy” step, in practice it is not: code execution is, by nature, a very unsafe operation; even more so in the context of language models, where you do not know what the code generated by the language model is going to look like. To solve this problem and to run code safely and efficiently, we will use the Together Code Interpreter (TCI). TCI elegantly abstracts away the complexity of sandboxed Python execution, providing a clean API that accepts code and returns structured results. This architectural choice has strong implications for our agent's design—it becomes inherently modular and maintainable. Since the code execution layer is completely decoupled from the reasoning logic, we can modify prompts, adjust the ReAct patterns, or add new analytical capabilities without touching the execution infrastructure. The agent's behavior becomes highly tunable through prompt engineering alone, while TCI ensures consistent, reliable execution regardless of the complexity of the generated Python code.
After having implemented the agent, we will focus on evaluation by providing quantitative and qualitative results regarding its performance. We will discuss some best practices on how to start building agents from scratch.
Why build agents from scratch?
This is a good question! Considering the number of frameworks out there, one might not need to dive into lower-level implementations. However, in the context of agentic models, it is actually good to know how things work at a lower level of abstraction.
Observing the steps a language model must take to interact with the external world is helpful in better understanding how the agent solves problems and how to improve general agent architectures. Even more so, considering how many edge cases one ends up encountering in the process of building agent architectures. Indeed, another interesting feature is that the Open Data Scientist agent we are going to build here is “hackable” and adaptable to different use cases.
A CodeAct Data Scientist
ReAct
The ReAct framework was introduced by Yao et al. (2023) to improve language agents. The idea behind the ReAct framework is encapsulated in its name: Reasoning and Action. The agent's activity unfolds using the following cyclical set of steps: the agent is tasked with a goal, reasons about what the next step could be, and then prepares the inputs to an action (which is then executed). Each action generates an observation from the environment, which is used as a source of new information for the next React step.
To clarify, the agent operates entirely through text generation. When we say the agent "reasons about the task," we mean it outputs natural language that articulates its step-by-step thinking process. When we say the agent "prepares the next action," we mean it generates structured text that specifies which function to call and with what parameters. Both reasoning and action preparation are fundamentally text generation tasks that enable the agent to interface with external tools and environments.
Why let the agent reason?
Letting the agent reason is inspired by chain-of-thought approaches, which show language models benefit from "thinking problems out loud." This is essentially a preparation step before acting:
- Reason: The agent develops a description of its reasoning
- Action: The agent designs an action to execute
- Observation: The environment returns information to the agent
Consider an agent whose task is to answer "What's the weather in NYC?" This could unfold as follows: The language model would reason, "I need to use the weather API tool to find the weather in NYC," and then construct the function call weather("NYC"). In reality, the entire process can occur within a single generation from the language model, where it's prompted to first think and then produce what's known as a tool call.
Once the language model generates the text, we can extract this action and execute it. We collect the output of the action (called observation) and add it to the context of the model. We then sample more tokens from the language model, which now is influenced by the additional observation of the context and will be able to provide an answer to the question above.
CodeAct
There are a few different ways one can build a ReAct pipeline. The Open Data Scientist will follow the CodeAct format and ask the model to output all actions as Python code. This approach comes with some advantages and disadvantages, but it is generally a very flexible way for the agent to design actions: the agent can express more complex operations than what would be possible with simple tool calling, since it can, in practice, combine multiple operations in a few lines of python code.
At each step, the language model inside our CodeAct agent will perform two possible actions: 1) generate a thought and then Python code as an action or 2) generate a thought and output a final answer for the user.
Together Code Interpreter
Note: TCI allows you to execute code safely in the cloud. For users who would prefer to execute code locally, we also release a simple (but very limited) local code sandbox as an alternative. More details available on the codebase.
The Together Code Interpreter allows us to run Python code in a safe environment and to collect the output of any line of code we send to it. TCI is fast and effective. At the start of a new task, we allocate a code interpreter sandbox and load the required files to do the analysis. TCI allocates a sandbox that stays alive for 60 minutes and can be used to run operations much like it was a Jupyter notebook.
TCI comes pre-installed with essential data science libraries like pandas, numpy, matplotlib, and scikit-learn, while supporting dynamic installation of additional packages: for example, one of the examples below requires installing RDkit, a library for cheminformatics. The platform handles rich output types including visualizations and maintains persistent state across code executions within a session.
Simple, yet effective, and we make it accessible to everyone
To make the open-source data scientist even more accessible, we developed a command line tool for the users to try out.
The Open Data Scientist agent can be triggered with one command line call:
open-data-scientist –executor tci --write-report
In this example, the data scientist agent explores the user-provided data, inspects the development environment and installs missing packages, and then develops a simple machine learning model, followed by result visualization and report writing steps. All these are fully automated and can be achieved with one Together API Key and a terminal. More details can be found in our codebase.
Interesting Reasoning Patterns
Our Open Data Scientist shows interesting self-correcting patterns. Note that content in the “Result” boxes is generated by the environment, not by the model.
For example, when attempting to use a BERT tokenizer to process data, it quickly becomes apparent that the dataset's token lengths exceed the model's token limit. Thus, it implements an alternative plan.
Similarly, when using one of the Sklearn linear models, warning messages inform the agent that the model it has trained has not converged, the agent thus decides to train the model again with more iterations to solve the problem.
Finally, the model also explores the data and tries to find better strategies to solve the problem.
Evaluating Data Scientist Agents
We extract two text tasks from OpenAI’s MLE-bench (Chan et al., 2024) as examples of real-world data science tasks. We then use the DABStep (Data Agent Benchmark for Multi-step Reasoning, Iglesias et al., 2025) benchmark. All our experiments are run with DeepSeek-V3 as a backend model. We release the output files we got from the model in the repository.
Kaggle
We select two tasks from the OpenAI MLE-bench Lite, namely the Spooky Author Identification and the Jigsaw Toxic Comment challenge. These are tasks that can be easily reproduced and that do not require downloading gigabytes of data. Our interest in these two benchmarks is mainly to see if the agent can solve an entire Kaggle task end-to-end. We keep the question prompt very high-level, we do not mention what the challenge is about; we just point the agent to the directory and instruct it to solve the challenge (which means opening the README and understanding the task first).
When challenged with these tasks, the agent was able to go through the tasks end-to-end. Which means: looking at the instructions, loading the data, doing some initial pre-processing, training some sklearn models, and running them. The agent saves a submission file that can then be used to evaluate the performance of the agent on the leaderboard. Both submissions provide reasonable results, while far from the top scores.
Interestingly, when asked to use more modern techniques, like sentence embeddings or even fine-tuning a BERT model, the agent was able to successfully use these techniques, also implementing evaluation pipelines to verify if, for example, the fine-tuning was successful.
DABstep
DABstep is a comprehensive benchmark developed by Adyen and Hugging Face to evaluate the capabilities of AI agents in real-world data analysis tasks. Comprising over 450 data analysis challenges derived from actual business workloads, DABstep tests AI systems on both structured and unstructured data, requiring them to perform multi-step reasoning across diverse analytical scenarios. What makes this benchmark particularly useful is the fact that it requires minimal setup: we can download the questions and the dataset files using Hugging Face; the agents need access to a limited (but comprehensive) set of files to answer all the questions.
The benchmark is specifically designed to assess how well language models and AI agents can handle complex data tasks that require sequential problem-solving rather than single-shot solutions.
In particular, DABstep assumes that the agents read the documentation which contains very specific definitions on how to interpret some terms (e.g., computing “fees” requires the usage of a specific formula, agents will never get the answer right if they do not read the documentation). We thus make the question prompt very domain specific to the challenge, forcing the agent to read all the documentation and think about the important details when providing answers. Current performance metrics reveal significant room for improvement, with even the most advanced reasoning-based agents achieving low accuracy on the benchmark.
DeepSeek V3 outperforms other agents in accuracy on easy tasks and maintains competitive results on harder ones.
From the results, we see that the agent is good at solving easy problems in the benchmark. The agent currently gets the highest scores on the validated leaderboard for the easy tasks (see our submission files here).
Coda
Building an effective data scientist agent is more accessible than expected. Using open-source models, the ReAct framework, and Together Code Interpreter, our simple yet effective implementation achieves competitive performance on benchmarks like DABStep.
While this agent has limitations around control and user interaction, it demonstrates the fundamental patterns needed for reasoning agents that can handle complex, multi-step data science tasks using readily available tools.
What do you need to build good agents?
- Start with the fundamentals. Good prompts can take you very far in the process.
- Agents can understand which solution you want, but they don’t know which trajectory they should take to get you there. Being very specific with the agent helps in getting the right trajectories.
- Robust execution environments matter more than you think. Your agent needs a reliable code interpreter.
- Design for iteration, not perfection. The best agents can debug and self-correct.
- Tests, tests, tests. Write comprehensive unit tests for every component.
- Keep the human in the loop without breaking the flow. Build clear handoff points where the agent surfaces its reasoning and asks for validation.
Limitations
There is no real control on the agent’s action, and under a pure engineering point of view, there is too little logging to make this a reliable tool/application. The data science process should allow interactions between the agent and the user so that interpretation errors can be quickly fixed.
Acknowledgements
We are really grateful for the HF Smolagents package, which was a source of inspiration for this release. Also, we are very grateful to Adyen for providing us with a benchmark that is simple to use and requires little scaffolding on the user side.
References
- Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., & Cao, Y. (2023). ReAct: Synergizing reasoning and acting in language models. Proceedings of the International Conference on Learning Representations (ICLR).
- Wang, X., Chen, Y., Yuan, L., Zhang, Y., Li, Y., Peng, H., & Ji, H. (2024). Executable code actions elicit better LLM agents. Proceedings of the International Conference on Machine Learning (ICML).
- Chan, J.S., Chowdhury, N., Jaffe, O., Aung, J., Sherburn, D., Mays, E., Starace, G., Liu, K., Maksin, L., Patwardhan, T., Weng, L., & Mkadry, A. (2024). MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering. ArXiv, abs/2410.07095.
- Iglesias, M., Egg, A., & Kingma, F. (2025, February). Data Agent Benchmark for Multi-step Reasoning (🕺DABstep). Adyen. https://www.adyen.com/knowledge-hub/data-agent-benchmark-for-multi-step-reasoning-dabstep