ayushozha commited on
Commit
e407bd2
·
1 Parent(s): 06b9d69

Add master blueprint and architecture SVG

Browse files
ReplicaLab_Architecture.svg ADDED

Git LFS Details

  • SHA256: 03bcb908e384bf7064f2252b21ff427b511706c727040b57155c512b137bafd4
  • Pointer size: 130 Bytes
  • Size of remote file: 23.6 kB
ReplicaLab_Master_Blueprint.md ADDED
@@ -0,0 +1,1097 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # ReplicaLab Master Blueprint
2
+
3
+ ## 1. Executive summary
4
+
5
+ **ReplicaLab** is an OpenEnv based scientific replication environment.
6
+
7
+ In each episode, the system creates:
8
+
9
+ 1. An original experiment or paper summary
10
+ 2. A lab with real constraints such as budget, equipment, reagent stock, staffing, and time
11
+ 3. A negotiation task where a **Scientist agent** and a **Lab Manager agent** must agree on a valid replication plan
12
+
13
+ The core idea is simple:
14
+
15
+ **One agent knows what the science needs. One agent knows what the lab can actually do. They must negotiate a replication plan that is scientifically valid and realistically feasible.**
16
+
17
+ This becomes a true environment because it has state, actions, observations, transitions, rewards, and episode termination. It is not just a chatbot prompt. It is a structured, trainable world.
18
+
19
+ ---
20
+
21
+ ## 2. The real world problem we are targeting
22
+
23
+ ReplicaLab targets the gap between **ideal scientific protocols** and **real lab constraints**.
24
+
25
+ In the real world, many experiments are hard to replicate because:
26
+
27
+ 1. Papers describe ideal methods
28
+ 2. Labs lack the full equipment or materials
29
+ 3. Budgets and schedules are limited
30
+ 4. Some substitutions are acceptable, but some break the science
31
+ 5. Teams must decide what is essential and what can change
32
+
33
+ So the real question ReplicaLab asks is:
34
+
35
+ **How do we adapt an experiment without breaking the science?**
36
+
37
+ This is the practical version of the replication crisis problem.
38
+
39
+ ---
40
+
41
+ ## 3. One line pitch
42
+
43
+ **ReplicaLab is an OpenEnv environment where a Scientist agent and a Lab Manager agent negotiate how to replicate scientific experiments under realistic lab constraints, and RL trains the Scientist to make better replication decisions over time.**
44
+
45
+ ---
46
+
47
+ ## 4. Which hackathon tracks we are following
48
+
49
+ ReplicaLab touches **4 out of the 5** hackathon problem statements.
50
+
51
+ ### 4.1 Primary tracks
52
+
53
+ #### A. Multi Agent Interactions
54
+
55
+ This is the strongest fit.
56
+
57
+ Why:
58
+
59
+ 1. The Scientist and Lab Manager hold different private information
60
+ 2. Neither can solve the task alone
61
+ 3. They must negotiate, exchange information, and converge
62
+
63
+ #### B. World Modeling, Professional Tasks
64
+
65
+ This is the second strongest fit.
66
+
67
+ Why:
68
+
69
+ 1. The environment simulates a real scientific workflow
70
+ 2. The agent must reason inside a partially observable professional world
71
+ 3. It must infer what the lab can and cannot do before making a good plan
72
+
73
+ ### 4.2 Supporting tracks
74
+
75
+ #### C. Long Horizon Planning and Instruction Following
76
+
77
+ Why:
78
+
79
+ 1. The task takes several rounds
80
+ 2. The agent must ask, revise, recover from mistakes, and plan ahead
81
+ 3. Reward is delayed until a protocol is good enough
82
+
83
+ #### D. Self Improvement
84
+
85
+ Why:
86
+
87
+ 1. The same environment is used for RL training
88
+ 2. The Scientist improves across repeated episodes
89
+ 3. The environment supports curriculum and replay later on
90
+
91
+ ### 4.3 Track summary
92
+
93
+ **Tracks touched technically:** 4
94
+
95
+ **Tracks we should lead with in the pitch:** 2
96
+
97
+ 1. Multi Agent Interactions
98
+ 2. World Modeling
99
+
100
+ **Tracks we should mention as supporting evidence:**
101
+
102
+ 1. Long Horizon Planning
103
+ 2. Self Improvement
104
+
105
+ ---
106
+
107
+ ## 5. Sponsor and partner alignment
108
+
109
+ ### 5.1 Best sponsor fits
110
+
111
+ #### Halluminate
112
+
113
+ Best fit because ReplicaLab is a true **multi actor environment**.
114
+
115
+ 1. The Scientist is one actor
116
+ 2. The Lab Manager is another actor
117
+ 3. The Judge can later act as a third oversight actor
118
+
119
+ #### Snorkel AI
120
+
121
+ Best fit because ReplicaLab behaves like **simulated experts in the loop**.
122
+
123
+ 1. The Scientist acts like a domain expert
124
+ 2. The Lab Manager acts like an operations expert
125
+ 3. The learning model improves through repeated expert style interactions
126
+
127
+ ### 5.2 Good optional fit
128
+
129
+ #### Fleet AI
130
+
131
+ This becomes stronger if the Judge is framed as an **oversight agent** that monitors, explains, and audits the decisions of the Scientist and Lab Manager.
132
+
133
+ ### 5.3 Resource fit
134
+
135
+ 1. **Hugging Face** for Spaces deployment and credits
136
+ 2. **Unsloth** for RL notebooks and simpler training setup
137
+ 3. **Northflank** for H100 access if faster training is needed
138
+ 4. **Cursor** for coding speed only
139
+
140
+ ---
141
+
142
+ ## 6. Why this is truly an environment
143
+
144
+ ReplicaLab is an environment because it contains the full RL loop.
145
+
146
+ ### 6.1 State
147
+
148
+ The state contains:
149
+
150
+ 1. The paper or experiment description
151
+ 2. The hidden minimum viable replication spec
152
+ 3. The lab constraints
153
+ 4. The round number
154
+ 5. The negotiation history
155
+ 6. The current proposed protocol
156
+ 7. The current score state
157
+ 8. Whether the episode is done
158
+
159
+ ### 6.2 Actions
160
+
161
+ The Scientist can:
162
+
163
+ 1. Propose a protocol
164
+ 2. Revise a protocol
165
+ 3. Request information
166
+ 4. Accept
167
+
168
+ The Lab Manager can:
169
+
170
+ 1. Report feasibility
171
+ 2. Suggest alternatives
172
+ 3. Reject
173
+ 4. Accept
174
+
175
+ ### 6.3 Observations
176
+
177
+ Each role sees a different view of the world.
178
+
179
+ The Scientist sees scientific requirements and negotiation state.
180
+
181
+ The Lab Manager sees operational constraints and negotiation state.
182
+
183
+ ### 6.4 Transitions
184
+
185
+ Each step updates:
186
+
187
+ 1. The conversation history
188
+ 2. The current protocol
189
+ 3. The round counter
190
+ 4. Budget usage if needed
191
+ 5. The done status if agreement happens or time runs out
192
+
193
+ ### 6.5 Reward
194
+
195
+ The environment returns a score based on:
196
+
197
+ 1. Scientific rigor
198
+ 2. Feasibility
199
+ 3. Fidelity to the original experiment
200
+
201
+ That is what makes it a trainable environment instead of a static task.
202
+
203
+ ---
204
+
205
+ ## 7. The core environment loop
206
+
207
+ ### 7.1 One episode
208
+
209
+ 1. `reset(seed=42)` creates a paper, a lab context, and a hidden evaluation rubric
210
+ 2. The Scientist receives its observation
211
+ 3. The Lab Manager receives its observation
212
+ 4. The Scientist acts first
213
+ 5. The Lab Manager responds
214
+ 6. This repeats for up to a fixed number of rounds
215
+ 7. If both accept, the episode ends successfully
216
+ 8. If time runs out, the episode ends with a penalty
217
+ 9. The Judge computes the final reward
218
+
219
+ ### 7.2 Environment methods
220
+
221
+ The environment should implement:
222
+
223
+ 1. `reset()`
224
+ 2. `step()`
225
+ 3. `state()`
226
+ 4. `close()`
227
+
228
+ These are the core methods that make the system compatible with OpenEnv serving and RL rollouts.
229
+
230
+ ---
231
+
232
+ ## 8. Scenario environments inside ReplicaLab
233
+
234
+ For the MVP, we should use **3 scenario families**.
235
+
236
+ ### 8.1 MVP scenario families
237
+
238
+ #### A. Cell Biology
239
+
240
+ Example:
241
+ Drug effect on cell proliferation using MTT or WST1 style assay
242
+
243
+ Why it is good:
244
+
245
+ 1. Easy to explain
246
+ 2. Has obvious lab constraints
247
+ 3. Good match between rigor and feasibility tradeoffs
248
+
249
+ #### B. Machine Learning Benchmark Replication
250
+
251
+ Example:
252
+ Reproducing a benchmark result with limited GPU budget and compute time
253
+
254
+ Why it is good:
255
+
256
+ 1. Easier to simulate
257
+ 2. Good for judges who understand ML
258
+ 3. Strong world modeling story around compute, time, and reproducibility
259
+
260
+ #### C. Behavioral Psychology Survey Study
261
+
262
+ Example:
263
+ Replicating a survey study with participant limits, time limits, and platform constraints
264
+
265
+ Why it is good:
266
+
267
+ 1. Gives variety beyond wet lab work
268
+ 2. Shows broader scientific replication use case
269
+ 3. Easy to explain ethical and logistical constraints later on
270
+
271
+ ### 8.2 Stretch scenario families
272
+
273
+ 1. Biochemistry
274
+ 2. Materials Science
275
+ 3. Chemistry
276
+
277
+ ---
278
+
279
+ ## 9. How each model interacts with the others
280
+
281
+ ### 9.1 Scientist agent
282
+
283
+ Role:
284
+ Protect scientific validity
285
+
286
+ Knows:
287
+
288
+ 1. The paper goal
289
+ 2. Important methodological elements
290
+ 3. Hidden scientific priorities through the environment design
291
+ 4. The negotiation history
292
+
293
+ Does not directly know:
294
+
295
+ 1. Full budget
296
+ 2. Full inventory
297
+ 3. Full equipment schedule
298
+ 4. Full staffing details
299
+
300
+ Main job:
301
+ Design a protocol that still counts as a meaningful replication.
302
+
303
+ ### 9.2 Lab Manager agent
304
+
305
+ Role:
306
+ Protect operational feasibility
307
+
308
+ Knows:
309
+
310
+ 1. Budget
311
+ 2. Equipment availability
312
+ 3. Booking conflicts
313
+ 4. Reagent stock
314
+ 5. Personnel constraints
315
+ 6. Safety restrictions
316
+ 7. The negotiation history
317
+
318
+ Does not directly know:
319
+
320
+ 1. Which scientific elements are absolutely critical
321
+ 2. Which substitutions are scientifically acceptable unless told
322
+
323
+ Main job:
324
+ Tell the Scientist what is actually possible and suggest realistic alternatives.
325
+
326
+ ### 9.3 Judge agent
327
+
328
+ Role:
329
+ Audit the final plan and score it
330
+
331
+ Knows:
332
+
333
+ 1. Original paper summary
334
+ 2. Minimum viable replication rubric
335
+ 3. Final protocol
336
+ 4. Actual constraints
337
+ 5. Full conversation history
338
+
339
+ Main job:
340
+ Compute the final reward and optionally explain it in plain English.
341
+
342
+ ---
343
+
344
+ ## 10. How the agents should be implemented
345
+
346
+ ### 10.1 MVP implementation choice
347
+
348
+ For the hackathon MVP:
349
+
350
+ 1. **Scientist** should be the only trained LLM policy
351
+ 2. **Lab Manager** should be rule based and deterministic
352
+ 3. **Judge** should be a deterministic rubric engine with optional LLM explanation
353
+
354
+ This is the safest and most realistic build path.
355
+
356
+ ### 10.2 Why only one agent should be trained first
357
+
358
+ 1. It reduces instability
359
+ 2. It makes reward improvement easier to show
360
+ 3. It makes the environment more deterministic and judge friendly
361
+ 4. It gives a clean before versus after story
362
+
363
+ ### 10.3 Scientist creation
364
+
365
+ The Scientist can be built from a small instruct model with structured JSON output.
366
+
367
+ The prompt should instruct it to:
368
+
369
+ 1. Protect scientific validity
370
+ 2. Ask for missing information before committing
371
+ 3. Output only valid schema fields
372
+ 4. Avoid invalid or impossible protocols
373
+
374
+ ### 10.4 Lab Manager creation
375
+
376
+ The Lab Manager should be implemented as a deterministic policy layer that:
377
+
378
+ 1. Checks budget
379
+ 2. Checks equipment availability
380
+ 3. Checks stock and restock timing
381
+ 4. Checks staff limits
382
+ 5. Returns templated natural language plus structured feasibility data
383
+
384
+ ### 10.5 Judge creation
385
+
386
+ The Judge should be implemented as:
387
+
388
+ 1. A rubric based scoring engine
389
+ 2. An audit note generator
390
+ 3. Optionally, an explanation layer that converts scores into readable comments for the frontend
391
+
392
+ ---
393
+
394
+ ## 11. How the judge agent is integrated
395
+
396
+ The Judge is integrated **inside the environment**.
397
+
398
+ It is called:
399
+
400
+ 1. At the end of the episode for final reward computation
401
+ 2. Optionally after each round for intermediate score previews
402
+
403
+ ### 11.1 What the Judge evaluates
404
+
405
+ 1. Whether critical controls were preserved
406
+ 2. Whether sample size is sufficient
407
+ 3. Whether substitutions are scientifically acceptable
408
+ 4. Whether the plan fits budget and inventory
409
+ 5. Whether the plan is faithful enough to the original design
410
+
411
+ ### 11.2 What the Judge returns
412
+
413
+ 1. `rigor_score`
414
+ 2. `feasibility_score`
415
+ 3. `fidelity_score`
416
+ 4. `total_reward`
417
+ 5. `judge_notes`
418
+
419
+ ### 11.3 Important design rule
420
+
421
+ The Judge should not be the entire reward source through free form opinions.
422
+
423
+ The Judge should primarily be a **deterministic rubric engine**.
424
+
425
+ That makes training, replay, and scoring much more stable.
426
+
427
+ ---
428
+
429
+ ## 12. Reward structure
430
+
431
+ The reward should be easy to explain and hard to game.
432
+
433
+ ### 12.1 Core reward dimensions
434
+
435
+ #### A. Rigor
436
+
437
+ Questions:
438
+
439
+ 1. Did the final plan preserve critical scientific elements?
440
+ 2. Are the controls present?
441
+ 3. Is sample size good enough?
442
+ 4. Is the technique valid?
443
+ 5. Is the study duration acceptable?
444
+
445
+ #### B. Feasibility
446
+
447
+ Questions:
448
+
449
+ 1. Is the plan within budget?
450
+ 2. Is the equipment actually available?
451
+ 3. Are the reagents in stock or restockable in time?
452
+ 4. Is the timeline realistic?
453
+ 5. Is staffing sufficient?
454
+
455
+ #### C. Fidelity
456
+
457
+ Questions:
458
+
459
+ 1. How close is the proposed protocol to the original experiment?
460
+ 2. Did the core method stay intact?
461
+ 3. Did the control logic stay intact?
462
+ 4. Is the sample size close enough?
463
+
464
+ ### 12.2 Composite reward
465
+
466
+ Use a multiplicative core so the agent cannot cheat.
467
+
468
+ ```text
469
+ base_reward = rigor * feasibility * fidelity * 10
470
+ bonus = efficiency_bonus + communication_bonus
471
+ penalty = timeout_penalty + invalid_action_penalty + over_budget_penalty
472
+ final_reward = base_reward + bonus - penalty
473
+ ```
474
+
475
+ ### 12.3 Why this is good
476
+
477
+ 1. High rigor but impossible protocol still scores poorly
478
+ 2. Cheap but scientifically broken protocol still scores poorly
479
+ 3. Fast, thoughtful negotiation gets rewarded
480
+ 4. The score is intuitive for judges
481
+
482
+ ---
483
+
484
+ ## 13. How RL works in ReplicaLab
485
+
486
+ ### 13.1 Simple explanation
487
+
488
+ RL works like this:
489
+
490
+ 1. The Scientist tries an action in the environment
491
+ 2. The environment responds through the Lab Manager and Judge logic
492
+ 3. The Scientist gets a reward at the end
493
+ 4. Training pushes the Scientist toward behaviors that earn higher rewards
494
+
495
+ ### 13.2 What behavior should improve
496
+
497
+ Over time, the Scientist should learn to:
498
+
499
+ 1. Ask better questions before proposing
500
+ 2. Avoid impossible protocols
501
+ 3. Preserve critical scientific details
502
+ 4. Choose better substitutions
503
+ 5. Reach agreement faster
504
+ 6. Reduce invalid actions
505
+
506
+ ### 13.3 What model should be trained
507
+
508
+ For the MVP, train only the Scientist.
509
+
510
+ That gives the clearest reward curve and the cleanest training narrative.
511
+
512
+ ---
513
+
514
+ ## 14. How self improvement works
515
+
516
+ ### 14.1 MVP self improvement
517
+
518
+ Self improvement in the MVP simply means:
519
+
520
+ **The Scientist gets better after repeated episodes.**
521
+
522
+ That is enough to satisfy the track.
523
+
524
+ ### 14.2 Stretch self improvement ideas
525
+
526
+ 1. Curriculum learning from easy to medium to hard scenarios
527
+ 2. Post episode self critique before retry
528
+ 3. Later training of both Scientist and Lab Manager
529
+ 4. Automatic scenario difficulty scaling
530
+
531
+ ---
532
+
533
+ ## 15. How world modeling is being done
534
+
535
+ World modeling means the agent must reason about a hidden world and update its internal understanding over time.
536
+
537
+ In ReplicaLab, that world includes:
538
+
539
+ 1. What equipment exists
540
+ 2. What equipment is missing
541
+ 3. Which items are booked
542
+ 4. What is in stock
543
+ 5. What can be substituted
544
+ 6. What is scientifically critical
545
+ 7. What tradeoffs hurt future feasibility
546
+
547
+ The Scientist does not see all of this at once.
548
+
549
+ So it must build a mental model of the lab through dialogue, feedback, and revision.
550
+
551
+ That is why ReplicaLab fits the world modeling track strongly.
552
+
553
+ ---
554
+
555
+ ## 16. How long horizon planning is being done
556
+
557
+ Long horizon planning appears because the task is multi step.
558
+
559
+ A good Scientist should:
560
+
561
+ 1. Understand the experimental goal
562
+ 2. Ask for missing constraints
563
+ 3. Propose an initial protocol
564
+ 4. Revise after operational feedback
565
+ 5. Trade off rigor against feasibility
566
+ 6. Converge before timeout
567
+
568
+ This is not one shot generation. It is multi round planning with delayed reward.
569
+
570
+ ---
571
+
572
+ ## 17. How constraints work
573
+
574
+ Constraints come from a seeded scenario generator.
575
+
576
+ ### 17.1 Constraint categories
577
+
578
+ 1. Budget
579
+ 2. Time limit
580
+ 3. Equipment availability
581
+ 4. Equipment booking calendar
582
+ 5. Reagent stock
583
+ 6. Reagent restock timelines
584
+ 7. Personnel count
585
+ 8. Safety restrictions
586
+
587
+ ### 17.2 Difficulty levels
588
+
589
+ #### Easy
590
+
591
+ The lab has most of what is needed.
592
+
593
+ #### Medium
594
+
595
+ The lab is missing some important pieces and requires thoughtful substitutions.
596
+
597
+ #### Hard
598
+
599
+ The lab is missing major pieces and forces serious protocol redesign.
600
+
601
+ ### 17.3 How constraints should change
602
+
603
+ For the MVP, keep each episode deterministic once the seed is fixed.
604
+
605
+ That means:
606
+
607
+ 1. `reset(seed=42)` always produces the same paper and constraint world
608
+ 2. The world only changes because of the agents’ actions
609
+ 3. No random hidden shocks should happen inside an episode yet
610
+
611
+ This makes testing and replay much stronger.
612
+
613
+ ---
614
+
615
+ ## 18. What the end result should be
616
+
617
+ The end result is **not** a full system that proves whether a paper is true or false.
618
+
619
+ The end result should be:
620
+
621
+ 1. A working OpenEnv environment
622
+ 2. A trained Scientist agent
623
+ 3. A stable Lab Manager policy
624
+ 4. A Judge rubric engine
625
+ 5. A public Hugging Face Space
626
+ 6. A training notebook that shows reward improvement
627
+ 7. A visual demo that clearly shows untrained versus trained behavior
628
+
629
+ The final result we are trying to fit is:
630
+
631
+ **a trainable benchmark and demo for scientific replication planning under constraints**
632
+
633
+ ---
634
+
635
+ ## 19. What the interface should look like
636
+
637
+ ### 19.1 Frontend choice
638
+
639
+ **React + Vite** is the right choice.
640
+
641
+ It is faster and cleaner than trying to build a full Cursor style IDE interface.
642
+
643
+ ### 19.2 UI layout
644
+
645
+ #### Left panel
646
+
647
+ 1. Original paper summary
648
+ 2. Key scientific requirements
649
+ 3. Seed
650
+ 4. Scenario type
651
+ 5. Round counter
652
+
653
+ #### Middle panel
654
+
655
+ 1. Negotiation log
656
+ 2. Scientist messages in blue
657
+ 3. Lab Manager messages in green
658
+ 4. Judge summary at the end
659
+
660
+ #### Right panel
661
+
662
+ 1. Current proposed protocol
663
+ 2. Budget bar
664
+ 3. Inventory summary
665
+ 4. Score bars for rigor, feasibility, and fidelity
666
+ 5. Final composite score
667
+
668
+ #### Bottom controls
669
+
670
+ 1. New episode
671
+ 2. Seed selector
672
+ 3. Scenario selector
673
+ 4. Replay slider
674
+ 5. Before versus after training toggle
675
+
676
+ ### 19.3 Fallback option
677
+
678
+ If the custom UI slips, use the OpenEnv web interface as a fallback and polish only the essential display panels.
679
+
680
+ ---
681
+
682
+ ## 20. Architecture overview
683
+
684
+ ```mermaid
685
+ flowchart TD
686
+ A[Scenario Templates] --> B[Scenario Engine]
687
+ B --> C[ReplicaLabEnv]
688
+ C --> D[Scientist Policy]
689
+ C --> E[Lab Manager Policy]
690
+ C --> F[Judge Rubric Engine]
691
+ D --> C
692
+ E --> C
693
+ F --> G[Step Result and Logs]
694
+ C --> G
695
+ G --> H[FastAPI and WebSocket Server]
696
+ H --> I[React Vite Frontend]
697
+ H --> J[Colab Training Client]
698
+ J --> K[TRL or Unsloth RL Training]
699
+ K --> L[Reward Curves and Evaluation]
700
+ ```
701
+
702
+ ---
703
+
704
+ ## 21. How exactly we are using the hackathon tools
705
+
706
+ ### 21.1 OpenEnv 0.2.1
707
+
708
+ Used for:
709
+
710
+ 1. Defining the environment interface
711
+ 2. Creating the stateful RL world
712
+ 3. Serving the environment over FastAPI and WebSocket
713
+ 4. Enabling clients to connect locally or remotely
714
+
715
+ ### 21.2 Hugging Face Spaces
716
+
717
+ Used for:
718
+
719
+ 1. Public deployment
720
+ 2. Judge accessible demo hosting
721
+ 3. Satisfying the official submission requirement
722
+
723
+ ### 21.3 Docker
724
+
725
+ Used for:
726
+
727
+ 1. Packaging the backend and optional frontend
728
+ 2. Ensuring the app runs on port 7860 in HF Spaces
729
+
730
+ ### 21.4 Colab
731
+
732
+ Used for:
733
+
734
+ 1. The required minimal training script
735
+ 2. Running rollouts against the environment
736
+ 3. Plotting reward improvement
737
+
738
+ ### 21.5 TRL or Unsloth
739
+
740
+ Used for:
741
+
742
+ 1. Training the Scientist policy
743
+ 2. Applying RL against the environment reward
744
+ 3. Producing visible reward curves and before versus after behavior
745
+
746
+ ### 21.6 Matplotlib
747
+
748
+ Used for:
749
+
750
+ 1. Reward curve visualization
751
+ 2. Component score plots
752
+ 3. Training summary charts
753
+
754
+ ### 21.7 GitHub
755
+
756
+ Used for:
757
+
758
+ 1. Public source code
759
+ 2. README
760
+ 3. Notebook storage
761
+ 4. Architecture documentation
762
+
763
+ ### 21.8 YouTube
764
+
765
+ Used for:
766
+
767
+ 1. The one minute demo video required by the hackathon
768
+
769
+ ---
770
+
771
+ ## 22. Scope of work
772
+
773
+ ### 22.1 In scope for the hackathon MVP
774
+
775
+ 1. OpenEnv environment implementation
776
+ 2. 3 scenario families
777
+ 3. Scientist as the trainable policy
778
+ 4. Rule based Lab Manager
779
+ 5. Deterministic Judge rubric engine
780
+ 6. FastAPI and WebSocket server
781
+ 7. Docker deployment
782
+ 8. Hugging Face Space
783
+ 9. Colab training notebook
784
+ 10. Reward curve
785
+ 11. React Vite frontend or clean fallback UI
786
+ 12. Public GitHub repo
787
+ 13. Demo video
788
+ 14. README
789
+
790
+ ### 22.2 Stretch scope if ahead of schedule
791
+
792
+ 1. LLM based Lab Manager
793
+ 2. Judge explanation LLM
794
+ 3. Live replay mode
795
+ 4. Before versus after split screen
796
+ 5. More scientific domains
797
+ 6. Difficulty curriculum
798
+
799
+ ### 22.3 Out of scope
800
+
801
+ 1. Proving a real paper is factually true or false
802
+ 2. Full autonomous laboratory automation
803
+ 3. Real wet lab execution
804
+ 4. Arbitrary paper ingestion from the internet
805
+ 5. Full self play between multiple LLM agents
806
+ 6. Complex enterprise integrations unrelated to the core demo
807
+
808
+ ---
809
+
810
+ ## 23. Folder structure
811
+
812
+ ```text
813
+ replicalab/
814
+ ├── README.md
815
+ ├── pyproject.toml
816
+ ├── openenv.yaml
817
+ ├── .dockerignore
818
+ ├── replicalab/
819
+ │ ├── __init__.py
820
+ │ ├── models.py
821
+ │ ├── client.py
822
+ │ ├── prompts/
823
+ │ │ ├── scientist.txt
824
+ │ │ ├── lab_manager.txt
825
+ │ │ └── judge.txt
826
+ │ ├── scenarios/
827
+ │ │ ├── templates.py
828
+ │ │ ├── cell_biology.py
829
+ │ │ ├── ml_benchmark.py
830
+ │ │ └── behavioral_psych.py
831
+ │ ├── scoring/
832
+ │ │ ├── rubric.py
833
+ │ │ ├── rigor.py
834
+ │ │ ├── feasibility.py
835
+ │ │ └── fidelity.py
836
+ │ ├── agents/
837
+ │ │ ├── scientist_policy.py
838
+ │ │ ├── lab_manager_policy.py
839
+ │ │ └── judge_policy.py
840
+ │ ├── env/
841
+ │ │ └── replicalab_env.py
842
+ │ ├── utils/
843
+ │ │ ├── seed.py
844
+ │ │ ├── validation.py
845
+ │ │ └── logging.py
846
+ │ └── outputs/
847
+ │ ├── logs/
848
+ │ ├── replays/
849
+ │ └── plots/
850
+ ├── server/
851
+ │ ├── app.py
852
+ │ ├── requirements.txt
853
+ │ └── Dockerfile
854
+ ├── frontend/
855
+ │ ├── package.json
856
+ │ ├── vite.config.ts
857
+ │ └── src/
858
+ │ ├── App.tsx
859
+ │ ├── components/
860
+ │ └── pages/
861
+ ├── notebooks/
862
+ │ └── train_colab.ipynb
863
+ └── tests/
864
+ ├── test_env.py
865
+ ├── test_reward.py
866
+ ├── test_scenarios.py
867
+ └── test_server.py
868
+ ```
869
+
870
+ ---
871
+
872
+ ## 24. How the judges are likely to judge the project
873
+
874
+ The hackathon judging criteria emphasize:
875
+
876
+ 1. Environment innovation
877
+ 2. Storytelling
878
+ 3. Training improvement
879
+ 4. Reward and pipeline coherence
880
+
881
+ ### 24.1 Why ReplicaLab scores well
882
+
883
+ #### Environment Innovation
884
+
885
+ Strong because this is a partially observable scientific negotiation world, not a toy single prompt task.
886
+
887
+ #### Storytelling
888
+
889
+ Strong because the Scientist versus Lab Manager framing is intuitive and memorable.
890
+
891
+ #### Training Improvement
892
+
893
+ Strong because the Scientist can visibly improve through RL and reward curves.
894
+
895
+ #### Reward and Pipeline Coherence
896
+
897
+ Strong because the scoring dimensions are simple and explainable.
898
+
899
+ ### 24.2 Ideal judge demo flow
900
+
901
+ 1. Show the problem in one sentence
902
+ 2. Start a seeded episode
903
+ 3. Show the paper and lab constraints
904
+ 4. Show the back and forth negotiation
905
+ 5. Show the score breakdown
906
+ 6. Replay the same seed with the trained Scientist
907
+ 7. Show higher reward and better decision quality
908
+
909
+ ---
910
+
911
+ ## 25. Completion rate expectations
912
+
913
+ ### 25.1 Project completion reality
914
+
915
+ With a focused 4 person team, we should aim to complete:
916
+
917
+ **90 percent of the judge critical MVP**
918
+
919
+ Even if that is only around **60 percent of the full dream vision**, that is completely fine.
920
+
921
+ ### 25.2 Environment success metrics
922
+
923
+ Track these metrics:
924
+
925
+ 1. Average reward
926
+ 2. Agreement rate
927
+ 3. Average rounds to agreement
928
+ 4. Invalid action rate
929
+ 5. Reward by scenario difficulty
930
+
931
+ A strong demo should show:
932
+
933
+ 1. Higher reward after training
934
+ 2. Higher agreement rate after training
935
+ 3. Fewer invalid proposals after training
936
+ 4. Faster convergence after training
937
+
938
+ ---
939
+
940
+ ## 26. Team split for 4 people
941
+
942
+ ### Person 1: Environment and scoring owner
943
+
944
+ Owns:
945
+
946
+ 1. Scenario generation
947
+ 2. Environment state and transitions
948
+ 3. Constraint system
949
+ 4. Reward logic
950
+ 5. Tests
951
+
952
+ ### Person 2: RL and model owner
953
+
954
+ Owns:
955
+
956
+ 1. Scientist prompts and action schema
957
+ 2. Training notebook
958
+ 3. TRL or Unsloth integration
959
+ 4. Reward curves
960
+ 5. Before versus after evaluation
961
+
962
+ ### Person 3: Backend and deployment owner
963
+
964
+ Owns:
965
+
966
+ 1. FastAPI server
967
+ 2. WebSocket protocol
968
+ 3. Docker image
969
+ 4. HF Spaces deployment
970
+ 5. Logs and replay endpoints
971
+
972
+ ### Person 4: Frontend and story owner
973
+
974
+ Owns:
975
+
976
+ 1. React Vite UI
977
+ 2. Visual score panels
978
+ 3. Demo polish
979
+ 4. README
980
+ 5. One minute YouTube demo
981
+
982
+ ---
983
+
984
+ ## 27. Workflow for the team
985
+
986
+ ### 27.1 Build order
987
+
988
+ 1. Freeze environment schema and reward structure
989
+ 2. Build one scenario end to end
990
+ 3. Add deterministic Lab Manager
991
+ 4. Add Judge rubric engine
992
+ 5. Connect FastAPI and WebSocket serving
993
+ 6. Add basic frontend
994
+ 7. Add Colab training notebook
995
+ 8. Deploy to HF Space
996
+ 9. Add remaining scenarios
997
+ 10. Record demo and finish README
998
+
999
+ ### 27.2 Runtime workflow
1000
+
1001
+ 1. User starts a new episode
1002
+ 2. The environment generates a seeded paper and lab
1003
+ 3. The Scientist receives its observation
1004
+ 4. The Lab Manager receives its observation
1005
+ 5. The Scientist proposes or asks a question
1006
+ 6. The Lab Manager replies with feasibility data
1007
+ 7. The environment updates state
1008
+ 8. The Judge computes intermediate or final scores
1009
+ 9. The episode ends on agreement or timeout
1010
+ 10. The replay is stored for demo and evaluation
1011
+
1012
+ ---
1013
+
1014
+ ## 28. Revenue model
1015
+
1016
+ This is not needed for judging, but it is useful for investor or product framing.
1017
+
1018
+ ### 28.1 Possible revenue paths
1019
+
1020
+ #### A. Enterprise experiment planning assistant
1021
+
1022
+ Sell a planning and auditing tool to biotech and research organizations.
1023
+
1024
+ #### B. Scientific AI benchmark licensing
1025
+
1026
+ Offer ReplicaLab as a benchmark for labs or AI teams evaluating scientific agents.
1027
+
1028
+ #### C. Simulation API
1029
+
1030
+ Charge for API access to scenarios, scoring, and replay infrastructure.
1031
+
1032
+ #### D. Workflow software expansion
1033
+
1034
+ Expand later into experiment design, lab operations support, and protocol adaptation copilots.
1035
+
1036
+ ---
1037
+
1038
+ ## 29. Five year old explanation
1039
+
1040
+ Imagine two kids want to bake a cake.
1041
+
1042
+ 1. One kid knows the recipe
1043
+ 2. One kid knows what is inside the kitchen
1044
+
1045
+ The recipe kid says, “We need chocolate.”
1046
+
1047
+ The kitchen kid says, “We do not have chocolate, but we have cocoa.”
1048
+
1049
+ Then they talk until they find the best cake they can make.
1050
+
1051
+ If the cake still tastes good, uses what the kitchen has, and finishes on time, they get a star.
1052
+
1053
+ ReplicaLab is that, but for science experiments.
1054
+
1055
+ ---
1056
+
1057
+ ## 30. Final recommended positioning
1058
+
1059
+ ### 30.1 Best main pitch
1060
+
1061
+ **ReplicaLab is an OpenEnv scientific negotiation environment where a Scientist agent and a Lab Manager agent collaborate to design valid experiment replications under real world lab constraints. We train the Scientist with RL so it learns to ask better questions, make better tradeoffs, and reach better replication plans over time.**
1062
+
1063
+ ### 30.2 Best track framing
1064
+
1065
+ **Primary:** Multi Agent Interactions and World Modeling
1066
+
1067
+ **Supporting:** Long Horizon Planning and Self Improvement
1068
+
1069
+ ### 30.3 Best sponsor framing
1070
+
1071
+ **Primary sponsor fit:** Halluminate and Snorkel AI
1072
+
1073
+ **Optional supporting narrative:** Fleet AI through the Judge as an oversight layer
1074
+
1075
+ ### 30.4 Best MVP framing
1076
+
1077
+ 1. Train only the Scientist
1078
+ 2. Keep the Lab Manager rule based
1079
+ 3. Keep the Judge rubric based
1080
+ 4. Ship 3 scenario families
1081
+ 5. Show one strong before versus after training demo
1082
+
1083
+ ---
1084
+
1085
+ ## 31. Final “done” definition
1086
+
1087
+ ReplicaLab is done for the hackathon when we have:
1088
+
1089
+ 1. A working OpenEnv environment
1090
+ 2. A deployed HF Space on port 7860
1091
+ 3. A public GitHub repo
1092
+ 4. A Colab notebook with visible reward improvement
1093
+ 5. A one minute YouTube demo
1094
+ 6. A clear README
1095
+ 7. A clean story that judges understand in under one minute
1096
+
1097
+ That is the real finish line.