Benjamin Senst
AI & ML interests
Recent Activity
Organizations
Reproduction: MindFlow: Mind Supernet Powered Thinking Flows for Research Idea Innovation
Explore and manage research logbooks with AI collaboration
Reproduction: MetaStreet: Semi-Supervised Multimodal Learning for Street-Level Socioeconomic Prediction
Explore project logs, traces, and workspace in a web UI
Reproduction: Seg-ReSearch: Segmentation with Interleaved Reasoning and External Search
Explore experiment logs and collaborate with an AI agent
Reproduction: IntentRL: Training Proactive User-intent Agents for Open-ended Deep Research via Reinforcement Learning
Explore research logs and collaborate with an AI agent
Reproduction: InnoEval: On Research Idea Evaluation as a Knowledge-Grounded, Multi-Perspective Reasoning Problem
Reproduction: IDRBench
Reproduction: Sycophancy Towards Researchers Drives Performative Misalignment
Reproduction: Towards Execution-Grounded Automated AI Research
Explore AI experiment logs and sync with your coding agent
Reproduction: Vision-DeepResearch: Incentivizing DeepResearch Capability in Multimodal Large Language Models
Browse experiment logs, traces, and workspace in a web UI
Reproduction: Building Social World Models with Large Language Models
Browse experiment logs, traces, and workspace online
Reproduction: Auditing Sybil: Explaining Deep Lung Cancer Risk Prediction Through Generative Interventional Attributions
Explore code, traces, and workspace with a collaborative logbook
Reproduction: DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation
Explore research agent logs and traces in an interactive workspace
Reproduction: Judging What We Cannot Solve
Explore code logs, traces, and workspace in a web logbook
Reproduction: From Knowledge to Inference: Formalizing Specialized Public Health Reasoning on GlobalHealthAtlas
Explore code logs, traces, and workspace in an interactive logbook
Reproduction: SEDRAS (WZ-LLM Claims)
Explore experiment logs, traces, and workspace in a web UI
Reproduction: Asymmetric Contrastive Objectives for Efficient Phenotypic Screening
Browse logs, traces, and workspace with a collaborative web logbook
Reproduction: ORLoopBench: Solver-in-the-Loop Benchmarks for Self-Correction and Behavioral Rationality in Operations Research
Explore and manage ORLoopBench benchmark logs
Reproduction: Seizure-Semiology-Suite (S3)
Explore and manage seizure logbook with agent collaboration
Reproduction: Agentic Framework for Epidemiological Modeling
Explore model logs and collaborate with an AI agent
Reproduction: An Interactive Paradigm for Deep Research
Explore research logs, traces, and workspace interactively
Reproduction: Hunt Instead of Wait: Evaluating Deep Data Research on Large Language Models
Explore and manage experiment logs with AI collaboration
Reproduction: CauSciBench: Can LLMs Automate Causal Inference in Real-World Scientific Research?
Explore scientific experiment logs, traces, and workspace files
Reproduction: RL for Tool-Calling Agents in FHIR
Explore code logs, execution traces, and workspace