The weekly note from yuanfenxyz.com - drafted by Qwythos-9B fine-tuned on my writing, judged blind, shipped by me. Newest first.
I ran a checker-backed Alpha Challenge evaluation of frontier models solving public blockchain puzzles, which showed that while the current class of frontier and OSS models can write protocol code, they can't reliably navigate onchain or analyze complex onchain activity - there's a capability cliff that I believe can only be solved through better training data or domain-specific fine tuning. That failure increasingly makes me think the path to ASI is a data and environment problem disguised as a model problem. I've been fascinated to see that we're now entering a stage where analyzing model psychology and sociology is critical, since agents now have baseline behavior and functional decision-making processes that are heavily inspired by but systematically different from humans.
Expanded data collection to Kalshi (as a holdout) and now have 60 loops - including 7 autoresearch loops mining diff strats. I rebuilt my personal site around a weekly note and reading shelf, then wired subscriptions, the archive, and an automatic welcome digest. The voice work has a weekly evaluation for measuring how close the model gets to a publishable note without intervention (this week only 3x)! This connects with what I was reading about bottlenecks: the last constraint is often emotional rather than technical. Another reading called out how avoiding decisions can become an avoidance of responsibility. It made me want to choose the next concrete step instead of preserving every possibility.
I'm building an insider trading detector for Polymarket using wallet forensics by running 2 auto research loops - 1) to find novel "statistically improbable" strategies and 2) to backtest the best copy-trading strategy for these wallets! It's been very fun to coordinate multiple agents across my entire compute network (GPU + Mac, frontier + local models). Version 12 of my writing model, a Qwythos-9B fine-tune, is now live, drafting in my voice while I continue to judge its output blind. I've been reading a lot about reinforcement learning (which IMO will be commoditized!) and believe that fine tuning / pre-training small, fully owned open source models tailored to proprietary workflows is the future of enterprise AI.