Conversation

I gave a talk at Bluehat Singapore. Here is a link to the slides, and a photo that perfectly frames it.

https://thomasdullien.github.io/about/slides/An-age-of-experimentation-BlueHat-Asia-2026.pdf

3
10
0

@HalvarFlake awesome slides as usual, definitely puts down some things in writing I’ve been struggling to wrap my head around or describe what I’m seeing and doing.

1
0
0

@HalvarFlake cool! love the history angle, and curious to know where you pulled it from. was a historian consulted? was it generated by AI? I'll write up what really happened 2012 to 2022, as it's very different to what you presented.

0
1
0

@HalvarFlake btw are you familiar with HoF bench? https://arxiv.org/abs/2607.27030 seems to be some overlap with some of your findings (and the "Statistics for experimenters" books techniques).

1
0
0

@wirepair oh no, not yet that's a great link!! Thank you!

1
0
0

@HalvarFlake yeah starting a new job soon which will include validating/benchmarking harnesses and i've been collecting resources, RAPTOR is one of many :>

0
0
0
@HalvarFlake Thanks for highlighting the importance of statistical rigor! I'm exactly as bad at stats as you describe, but seeing conclusions drawn from 3-5 test runs being the "standard" for model/prompt/agent evaluation feels just wrong.
0
0
1