April 11, 2026 6 min read

Links: Week of 12 Apr 2026

  1. We did find some Reddit comments, though, warning other netizens to steer clear of MEDVi, claiming serious allegations of possible HIPPA violations, shady billing practices, and even damaged vials of seemingly bogus drugs causing physical harm.

    AI is making the web weirder and muddier than ever. And though MEDVi promises that “sometimes you have to see it to believe it,” in our burgeoning AI-powered web, that’s no longer the case.

    MEDVi, sadly, is the same company from last week's NYT's article about a one-person, $1.8bn company. It is disappointing to see NYT fall for their hype despite this article being published almost a year ago.

    This, yet again, also raises the question of just how credulous and naive am I being when it comes to the AI Hype cycle. Keep that in mind with rest of this week's coverage.

  2. Today we’re announcing Project Glasswing1, a new initiative that brings together Amazon Web Services, Anthropic, Apple, Broadcom, Cisco, CrowdStrike, Google, JPMorganChase, the Linux Foundation, Microsoft, NVIDIA, and Palo Alto Networks in an effort to secure the world’s most critical software.

    We formed Project Glasswing because of capabilities we’ve observed in a new frontier model trained by Anthropic that we believe could reshape cybersecurity. Claude Mythos2 Preview is a general-purpose, unreleased frontier model that reveals a stark fact: AI models have reached a level of coding capability where they can surpass all but the most skilled humans at finding and exploiting software vulnerabilities.

    Mythos Preview has already found thousands of high-severity vulnerabilities, including some in every major operating system and web browser. Given the rate of AI progress, it will not be long before such capabilities proliferate, potentially beyond actors who are committed to deploying them safely. The fallout—for economies, public safety, and national security—could be severe. Project Glasswing is an urgent attempt to put these capabilities to work for defensive purposes.

    Since Anthropic (along with OpenAI) is trying to IPO this year, it is tempting to dismiss this as hype, especially in context of the previous link. However, there are many signals that end credibility to their claims.

    First, there is the large list of credible partners above including their competitor in the LLM space, Google. Second, was the news that Treasury Secretary Scott Bessent and Chairman of the Federal Reserve summoned CEOs of major Financial Services firms to warn them about the risks posed by this model. Third is the long list of credible tech people endorsing the abilities of this model.

    With this level of publicity, if this was hype, we will find out soon enough but the evidence so far suggests it is likely real.

    In which case, this is a huge step change in the abilities of LLMs. I expect this will also bring AI centerstage in national and global political discourse. This is a model with major national security implications because the NSA / Mossad types can use one vulnerability in operating systems to compromise personal devices of their targets. Imagine what they could do with "thousands of high-severity vulnerabilities".

    This also raises important questions like what if China had developed a model with such abilities first or what if Anthropic hadn't realized the power of this model and released it to public or who gets to decide who gets access to a model like this, a private company or government?

    The other question I am thinking about is how do leaders of China, Russia react to this news knowing that NSA / CIA have access to such a system?

    There is a lot of excellent coverage of Mythos and related stuff, if you want to read more.

  3. First Banksy and then Satoshi. Something about their unmasking is not sitting right with me. I am bothered by it. I am annoyed by it. And even more annoyed with myself because as a former journalist I should understand, but I don’t. I am referring to Reuters’s meticulous investigation and unmasking of Banksy, and John Carreyrou’s in-depth report labeling Adam Back as Satoshi, the creator of Bitcoin.

    Both investigations are technically impressive. Both raised the same question I keep turning over: what exactly was accomplished here, and for whom?

  4. When KJ Muldoon was born in the summer of 2024, his parents were told he had a disease so rare, it strikes about one in 1.3 million newborns. His condition, a severe deficiency of an enzyme known as CPS1, left his tiny body unable to properly break down protein, flooding his blood with toxins that could cause brain damage or death. A liver transplant could correct the problem, but KJ was too young and too fragile to undergo one. With each passing day, the risk of irreversible neurological damage grew.

    What happened next may become the most important medical story of the decade. In just six months, a team at Children’s Hospital of Philadelphia and Penn Medicine designed a personalized therapy that could correct the single misspelled letter in KJ’s DNA using a gene editing technology known as CRISPR. To get the therapy inside KJ’s cells, doctors relied on the same kind of mRNA technology that powered the Covid-19 vaccines. He received his first dose at 6 months old. One year later, KJ is walking, talking and thriving at home with his family.

    Worth a read, the key question being how does the FDA regulate individualized treatments when the current paradigm is to rely on RCTs with thousands of subjects.

  5. Ms. Judis currently holds the Guinness World Record for oldest competitive rope skipper. She also thrives on having an audience: If she doesn’t share a workout, she said, it’s like it never happened.

    82!

  6. Twelve months ago, I signed up for the Paris Marathon. Within six months, I knew I’d be in trouble without a trainer. So, living in the San Francisco Bay Area — the home of artificial intelligence — I decided to build one myself.

  7. We should all do this sort of thing more often. 🙂

  8. Imagine I told you that AI was going to create a 40% unemployment rate. Sounds bad, right? Catastrophic even. Now imagine I told you that AI was going to create a 3-day working week. Sounds great, right? Wonderful even. Yet to a first approximation these are the same thing. 60% of people employed and 40% unemployed is the same number of working hours as 100% employed at 60% of the hours.

April 4, 2026 2 min read

Links: Week of 05 Apr 2026

A lighter edition this week as the family traveled for Spring Break. Normal service should resume next week.

  1. His start-up, Medvi, a telehealth provider of GLP-1 weight-loss drugs, got 300 customers in its first month. In its second month, it gained 1,000 more. In 2025, Medvi’s first full year in business, the company generated $401 million in sales.

    Mr. Gallagher then hired his only employee, his younger brother, Elliot. This year, they are on track to do $1.8 billion in sales.

    A $1.8 billion company with just two employees? In the age of A.I., it’s increasingly possible.

  2. The most plausible decision however is to slightly lower your level of ambition.

    I thought the correct response was the opposite. Think of it along the lines of immigration - one very common path for an ambitious person from India / China / Africa is to immigrate to the West. Do big things there or learn and go back to do big things at home. An advanced alien world opens up a whole new frontier. But here's a much better answer:

    I, like most people, used to think that UFOs were ridiculous and in the same category as Bigfoot and the Loch Ness monster. Nutcases believed in them because they wanted it to be true, to feel special, and to enjoy confirmation bias.

    Now I feel that I owe the UFO nuts an apology. I’m not claiming they’ve been definitively proven right, but evidence has come out that they don’t belong in the same category as other conspiracy theorists.

    This means I was globally miscalibrated. My model of the world was too narrow, so I should broadly update to be more open-minded about crazier things.

    A lot more at the link.

March 29, 2026 2 min read

Links: Week of 29 Mar 2026

  1. The Chinese party-state is fundamentally a set of goal-oriented institutions. This is not unique to China—it is in fact a distinguishing feature of all Leninist systems. I sometimes think of Leninist systems as a little bit like that bus in the movie Speed. Who here has seen it? For those who haven’t, here is basic gist of that film: an extortionist attaches a bomb to the speedometer of a bus. If the bus ever slows below 50 miles per hour, everyone blows up. So it is with your average communist system. Either it hurtles towards some clearly defined goal or things start to fall apart.

  2. SS
    Stefan Schubert@StefanFSchubert · Mar 28

    While social media is polarising, evidence suggests AI may nudge people towards the centre.

    This holds true of all studied models. Grok is more right-leaning than other models, but also has depolarising effects.

    By @jburnmurdoch.

  3. Behind this drive is an experience of A.I. that many casual users have not yet had. An A.I. without deep knowledge of you is an upgrade, perhaps, over Google search. An A.I. with deep knowledge of you feels like something else entirely. I have heard people talk about their A.I.s in terms that bring to mind the daemons from Philip Pullman’s “His Dark Materials” trilogy: They become companions that know you deeply, that you feel safe telling things you’d never tell another person, that become a separate self that nevertheless feels like a part of your own self. That this sounds strange and disquieting does not mean it is not happening.

  4. Here are 30 questions to elevate your awareness of the greater place in which you live:

    1. Point north.
    2. What time is sunset today?
    3. Trace the water you drink from rainfall to your tap. Where does your water come from?
    4. When you flush, where do the solids go? What happens to the waste water?
    5. How many feet (meters) above sea level are you right now? How about your home?
    6. What spring wildflower is consistently among the first to bloom here?
March 21, 2026 4 min read

Links: Week of 22 Mar 2026

Word of The Day: TESCREAL

Meaning: A neologism and a acronym, it stands for Transhumanism, Extropianism, Singularitarianism, (modern) Cosmism, Rationalists (the internet community, not to be confused with other uses of the term), Effective Altruism, and Longtermism.

Gebru and Torres argue that these ideologies should be treated as an "interconnected and overlapping" group with shared origins. They claim these constitute a movement that allows its proponents to use the threat of human extinction to justify expensive or detrimental projects and consider it pervasive in social and academic circles in Silicon Valley centered on artificial intelligence. As such, the acronym is sometimes used to criticize a perceived belief system associated with Big Tech.

A regular reader of the blog pointed this word out to me, as a message that blind cheerleading for big-tech was something for me to guard against. The link last week to the story about the Sydney Data Engineer developing a vaccine for his dogs cancer was the trigger for sharing this.

Its advice well received, since that story did trigger some red-flags at the back of my mind but I over-rode those signals with little thought as I continue to be very excited by the developments in AI space. However that is no reason to throw caution and skepticism to the wind. The story has held up so far but I continue to watch with interest.

Links

  1. The H1B Fees:

    CO
    Connor O’Brien@cojobrien · Mar 12

    85 people have paid the $100,000 H-1B fee so far, totaling $8.5 million in revenue. But fee revenue from H-1B apps abroad is down $28 million.

    So the fee — justified by a paper claiming the revenue-maximizing fee was >$100,000! — appears to have lost the government $20 million.

  2. Microsoft is admitting, at least for now, that delivering a truly compelling agentic product that enterprises are willing to pay for means abandoning their stated goal of being model agnostic; that, by extension, raises the possibility that models are not and will not be commodities, because agents require more than models.

    Food for thought as my framework so far was that models will become commodities but the recent developments are pointing away from that direction. There is a lot more of interest in this article.

  3. Michael Smith used AI to create music, and then used AI to create bots to get the “plays” and took the smartest technology companies, including Spotify and Amazon, who should know better, for about $8 million. He is going to jail for his crimes. It is easy to dismiss this as one-and-done fraud. It is anything but. It is an early warning of how AI will disrupt the systems that power our digital society: how culture gets discovered, how commerce gets directed, and how conversations get shaped.

    Will AI break all recommendation algorithms, from YouTube to Tik Tok?

  4. Nothing you own is finished. Everything exists in a state of permanent incompletion, permanently needing. Your phone needs updates, needs charging, needs storage cleared, needs passwords rotated.

    Your apps need permissions reviewed, terms accepted, preferences re-configured after every update.

    Your subscriptions need evaluating, need renewing, need canceling, need justifying to yourself every month when the charge appears. The purchase isn't the end of anything. It's the first day of a relationship you didn't agree to, with no clean way out.

    You live in a house full of dependents.

  5. More guides and ideas for how to use LLMs.

March 14, 2026 2 min read

Links: Week of 15 Mar 2026

The first two stories this week are mindblowing. Huge if true, as they say.

  1. A Sydney data engineer with no background in biology has used ChatGPT and AlphaFold to design what researchers are calling the world's first personalised mRNA cancer vaccine for a dog, and the results have stunned the scientists who helped make it.

  2. In 2024, the entire neuronal diagram of the fruit-fly brain–some 140,000 neurons and 50 million connections–was mapped. Later research showed that the map could be used to predict behavior. Now, Eon Systems a firm with some of the scientists involved in the fruit-fly research and with the goal of uploading a human brain has announced that they uploaded the fruit fly brain to a digital environment.

    The digital fly appears to behave in the digital environment in reasonably fly like ways–this is not a simulation, the fly’s “sensors” are being activated by the digital environment and the neurons are responding.

  3. Some more AI tutorials. So many tutorials, so little time.

  4. BO
    Brooks Otterlake@i_zzzzzz · Mar 13

    Japanese society is so civilized that the fires simply drive themselves to the fire station

March 7, 2026 6 min read

Links: Week of 08 Mar 2026

SavithaShan

Savitha Shan, an undergrad double major here in economics and information systems, who was murdered over the weekend by an Islamist terrorist who started randomly shooting people on Sixth Street, apparently angry about the war in Iran. Two other innocents were also killed. - Scott Aaronson

So senseless. And the 180 schoolgirls Minab, Iran. Did any of them know it was their time? Did they get to live a full life? Will I? It's one thing to know this and another to feel it in your bones. But the worst is when you start feeling it and your self-preservation instinct kicks in - allowing the feelings to only go in so deep and no more. RIP.

Links

  1. This is an instance of what I call the comb-over effect: when a series of individually small changes takes you from something that's a little bit off to something that's freakishly wrong.

  2. The leaders who win this era won’t just be 22‑year‑olds building AI‑native startups. They’ll also be experienced operators who integrate AI quietly and intelligently into systems they already understand. If you’re over 50 and feeling behind, you might actually be early. Because when the tools get easier, experience becomes more powerful—not less. And this time, that experience may finally be the competitive edge.

  3. BC
    Brett Caughran@FundamentEdge · Mar 6

    I'll provide a little more specificity on this, and snippets of an example.

    For many months people have been talking about a "Cursor moment" in finance, where workflow changes so dramatically that you hit the steep part of an adoption curve. I've been highly skeptical of that, for a few reasons.

    But the most fundamental reason is the LLM technology just wasn't there. The foundation models simply did not have enough power to interact with Excel spreadsheets in any sort of usable way (despite splashy demos...). Even if you solve the (very hairy) data challenges, 2025-era LLMs just didn't have the power to interact with spreadsheets.

    So we could sit and talk about a lot of ideas and concepts on how AI could augment institutional investment research. But it was just that, a concept.

    I have a series of tests I run on new AI models that are capability tests for hedge fund style research workflows. And the easiest is just uploading an existing Excel file to see if the LLM can understand what's going on. If LLMs can't sufficiently read and understand an Excel model, the full stack of AI Excel workflows is just not possible (in my opinion). And a waste of time to try to explore.

    This didn't work to any sort of impressive degree (Opus 4.6 could do it, but not do it well). Until yesterday, with GPT-5.4 Thinking.

    Suddenly, I can now get something that is not only modestly useful, but I think will immediately become part of my investment process workflow.

    I call it "PM Review", or a structured evaluation and push back on a model. I have participated in literally hundreds of these as both analyst and PM. Effectively the analyst builds a model, sends it to the PM, and they walk through it together. The wise, experienced-scarred PM will rip the model apart, push back, and help steer the model to a usable outcome.

    A great PM will be able to hone in on the two or three key variables that matter and identify aggressive or conservative assumptions. An analyst may be pitching a stock where the core quantitative input is supported by flawed logic. And the PM's job is to try and identify that flawed logic. This workflow, to me, is a key differentiator between good and not good PMs.

    However this workflow isn't just for PMs; it's for analysts who are trying to evaluate their own work, peer analysts who want to do thoughtful push-back on ideas the team may participate in, our director of research teams who are looking to efficiently evaluate the idea underwriting process. Or PMs for the first cut if they're looking at lots of ideas.

    The intriguing aspect of augmenting this process with AI is it scales incredibly. And it can run autonomously. Across 300 models I could have a swarm of agents doing automated due diligence on the key drivers, updating those models, feeding those results back to me, and flagging which of my covered ideas have earnings revision potential. This workflow is the "Cursor moment" for public equity research, in my opinion. I'm not saying we're there by any means as data accuracy and the structures required to incorporate internal data are still in progress. But we just took a step forward in the technological capability.

    I tested this out in GPT-5.4. And while it's not perfect this is the first time I've received anything that's useful back in this test.

    I'll walk you through a couple of steps to do this on your own.

    Step 1: brain dump into Claude. I don't know if there's any logic to it or just my own habit but if I'm executing in Chat GPT, I'll meta prompt and Claude and vice versa. I'm not sure where you meta prompt matters all that much for the types of workflows I do but it CERTAINLY matters if you meta prompt vs. raw prompt so don't skip this step.

    Step 2: take that prompt output, turn it into Markdown, and put that as custom instructions in a GPT project. This is just a workflow efficiency because then I now have a GPT project that I can upload any model into.

    Step 3: run the prompt. I purposely jacked up my DraftKings model a little bit (and it's a work in progress anyway so do not take any of these estimates as anything I believe).

    But it produced an exceptionally helpful:
    1) Executive Summary
    2) Business understanding (explaining how a dollar flows through P&L)
    3) Model Evaluation, providing an assessment and sanity check of all of the key inputs
    4) Model audit, looking for input consistency, formula integrity, and broken references
    5) A road map for incremental due diligence
    6) The highest value IR questions

    I encourage you to check it out for yourself.

    Will link to the six-page output in the replies.

  4. C
    Chapin@Chapinc · Mar 1

    You want me to be physically present at a meeting in the office? Like the Ayatollah?

  5. I wouldn't stand there.

    IWouldntStandThere
March 7, 2026 5 min read

Feeling the AGI

I have been following the AI revolution almost since the day ChatGPT launched in Nov 2022. That this was a transformational technology has also been clear to me for almost as long. I even sensed the "vibe shift" in early Jan.

But so far I didn't really "feel the AGI". Last week I did.

AGI stands for Artificial General Intelligence. Here's how Claude defines it:

An AI system capable of performing any intellectual task that a human can — reasoning, learning, and adapting across domains without being specifically programmed for each one." - Claude Sonnet 4.6

It is also worth knowing the related concept of ASI, Artificial Superintelligence.

An AI that surpasses the cognitive abilities of all humans combined across every domain, including scientific reasoning, social intelligence, and creative problem-solving."

Last week we finally signed up for Claude Code at work and I started playing with it.

Claude Code works via the CLI (command line interface) or the terminal - the black screen with white text used by all the movie nerds. It can be a little intimidating if you are not a programmer but really once you set up Claude Code, it works just like the chat window.

It is so much more powerful though. It can manipulate files on your computer and run the code it writes. This means it is not restricted to recommendations or single steps any more. It can generate an entire plan of action and execute and implement the thing by itself. It can be more than a little freaky when you see the output of a particularly complex task.

Let me share a couple of examples that blew my mind.

At work we have a database that stores our financial projections for a company we cover. The investment team can use custom functions in excel to then download this data. This is useful to create reports for analysis - say comparing 5 companies across a few metrics. There are different functions for different types of data and the IT team has created an excel file with about 20 sheets listing the syntax for each function, the list of values that can go in each function etc.

I pointed Claude Code to this help file and asked it to create a "skill" for itself that would allow it to create reports in excel using these formulas. I gave it the same context as the previous paragraph - maybe a little more technical but nothing a lay person wouldn't understand.

With that single command, in maybe 3-5 minutes, the skill was ready. Now when we need to create a report we can just ask Claude Code to use the skill, the data we want in the report and it creates a fully formatted excel file with the (usually) correct formulas. Tasks that would take me or my team members 20-60 minutes, automated permanently. Using English sentences, no technical knowledge.

Second example. We wanted to perform some statistical analysis on 20 year historical performance of 600 stocks to identify specific episodes / time periods and then dig deeper into specific episodes to understand their fundamental causes.

An analyst spent probably 5-6 hours to download the data in excel, process it so we could start identifying the qualifying episodes / time periods in different markets. At this point we were somewhat stumped about how to isolate the relevant episodes from this vast data. Probably a trivial problem for a data scientist but not for us.

In the past, we would have spent probably another 5-10 hours trying to either eyeball the data using charts or some other way to get our answer. Over the last couple of years, we would have asked Claude or ChatGPT to suggest a better way.

But since we had Claude Code, I pointed it to the existing file, with all its messy sheets and structure and explained what we were trying to do and asked it to identify the episodes. It whirred away for about 20 min, an occasional question here or a permission there and then it spat out a report.

But the report didn't just have the episodes identified. It also identified potential causes for each episode (based on web searches presumably), linked patterns across multiple episodes, gave charts and commentary that helped us understand the relevance of each episode, limitations of the analysis, suggested next steps etc.

Now we have all seen more than enough AI slop to not be impressed by the sheer volume of content these tools can spit out. But we spent a few hours verifying the numbers and conclusions. So far it all checks out. You have to take my word for it, but man, this was not slop. If a junior associate had put this out I would be proud of them. There were parts that I would have been proud to create and, of course, large parts that we simply couldn't have created at all.

And those were not the only thing Claude Code did last week that blew my mind.

So, let it be noted on this 8th of March, 2026, I felt the AGI last week. I may even have felt the ASI.

February 28, 2026 5 min read

The Citrini Scenario

The big news in markets this week was this report from Citrini Research & Alap Shah that apparently crashed the markets and led to a lot of debate in our office. It lays out a "fast take-off" scenario for AI, which causes mass layoffs of white-collar emplopyees as AI replaces intelligence work and starts off an economic downward spiral as demand collapses.

It should have been clear all along that a single GPU cluster in North Dakota generating the output previously attributed to 10,000 white-collar workers in midtown Manhattan is more economic pandemic than economic panacea. The velocity of money flatlined. The human-centric consumer economy, 70% of GDP at the time, withered. We probably could have figured this out sooner if we just asked how much money machines spend on discretionary goods. (Hint: it's zero.)

AI capabilities improved, companies needed fewer workers, white collar layoffs increased, displaced workers spent less, margin pressure pushed firms to invest more in AI, AI capabilities improved…

It was a negative feedback loop with no natural brake. The human intelligence displacement spiral. White-collar workers saw their earnings power (and, rationally, their spending) structurally impaired. Their incomes were the bedrock of the $13 trillion mortgage market - forcing underwriters to reassess whether prime mortgages are still money good.

The report found many believers in markets but I find myself on the skeptical side, much more pusuaded by the many pushback articles which are grounded in conventional economic theory. And they came from many sources.

Here's Tyler Cowen in his cryptic style. Here's Zvi with the inverse. And finally here's Citadel.

And lastly, here's Claude summarizing it all and adding its own perspective:

Why Citrini's Scenario Doesn't Add Up The piece is an excellent thought experiment and a useful sector-level vulnerability map. The macro conclusion — that AI abundance causes a demand collapse and systemic crisis — is built on a fundamental accounting error.

The Core Contradiction: Every Loss Is Someone Else's Gain The entire scenario rests on a demand collapse: AI replaces workers, workers stop spending, the economy spirals. But the same force destroying jobs is also destroying prices. If a Claude agent does the work of a $180K PM for $200/month, then everything that PM helped produce also gets dramatically cheaper. The piece catalogs agents slashing insurance premiums, SaaS costs, delivery fees, real estate commissions, and interchange — then claims displaced workers can't afford things. Which things? The things that just got 80% cheaper?

Every corporate revenue loss in the piece is a gain on the other side. ServiceNow loses $500K in licenses — that's $500K freed for the client. DoorDash loses its 30% take rate — drivers earn more, consumers pay less. Real estate commissions drop from 6% to 1% — that's a 5% stimulus to every home purchase. SaaS fees are a tax on business. That tax went down.

Meanwhile, the piece describes NVIDIA posting records, hyperscalers spending $150-200B/quarter, AI companies thriving. Someone is paying for all of that. You cannot have booming AI revenues and an economy where nobody is spending. The money doesn't vanish — it circulates through different channels. The piece tracks one side of every transaction and ignores the other.

The Timeline Is Physically Impossible The scenario requires mass deployment across every major industry in ~30 months. Enterprise procurement cycles, organizational restructuring, regulatory review, and infrastructure buildout don't move at this speed. More critically: there isn't enough compute on Earth to run always-on agents for every American consumer while simultaneously replacing most white-collar work by 2028. The compute constraint alone paces everything and makes this timeline a non-starter. Even clearly superior technologies (electricity, the internet) took decades to restructure economies.

Zero Policy Response for Two Years Is Ahistorical The US deployed trillions within weeks during COVID. Automatic stabilizers activate immediately when unemployment rises. The piece treats political difficulty in designing optimal long-term policy as equivalent to no response at all.

More importantly: in a world with booming real output and deflation, aggressive fiscal stimulus is straightforwardly correct — you have too much output and not enough demand, so you print money without inflationary risk. During COVID, stimulus risked inflation because real production was down. Here, real production is exploding. This is the easiest macro policy environment imaginable.

What's Actually Useful in the Piece The sector-level analysis of which business models are vulnerable is genuinely good. SaaS pricing power, intermediation platforms, payments, insurance built on inertia, real estate brokerage — anything dependent on friction or information asymmetry is legitimately threatened. The private credit / insurance / reinsurance daisy chain is a real and underappreciated fragility.

The transition will be uneven. Specific sectors, geographies, and demographics will get hit hard, and the speed of disruption concentrated in high-earning white-collar cohorts could cause real dislocations even if the macro aggregates are fine. That's worth positioning around.

But the mechanism for how sectoral disruption becomes a macro demand collapse is where the analysis fails — because it treats every efficiency gain as pure destruction rather than a transfer from incumbents to consumers. Read it as a sector vulnerability map, not a macro forecast.

Going back to the title of my last post from Singapore, there's clearly been a vibe-shift when it comes to what the latest models can do. In early Jan only those focused on AI were feeling it and now it is going more mainstream. But that's no reason to fall prey to half-baked thinking. Let Claude taste it and cook it for you.

February 28, 2026 3 min read

Links: Week of 01 Mar 2026

Links

  1. I think I have been underrating the pace of model improvements. In March 2025 I made a bet with Tamay Besiroglu, cofounder of RL environment company Mechanize, that AI tools would not be able to autonomously produce papers I judge to be at a level comparable to that of the best few papers published in 2025, at comparable cost to human experts, by 2030. I gave him 3:1 odds at the time; I now expect to lose this bet.

  2. It is hard to communicate how much programming has changed due to AI in the last 2 months: not gradually and over time in the "progress as usual" way, but specifically this last December. There are a number of asterisks but imo coding agents basically didn’t work before December and basically work since - the models have significantly higher quality, long-term coherence and tenacity and they can power through large and long tasks, well past enough that it is extremely disruptive to the default programming workflow.

    It’s not perfect, it needs high-level direction, judgement, taste, oversight, iteration and hints and ideas. It works a lot better in some scenarios than others (e.g. especially for tasks that are well-specified and where you can verify/test functionality). The key is to build intuition to decompose the task just right to hand off the parts that work and help out around the edges. But imo, this is nowhere near "business as usual" time in software.

  3. I think of vibe coding using its original definition of coding where you pay no attention to the code at all, which today is often associated with non-programmers using LLMs to write code.

    Agentic Engineering represents the other end of the scale: professional software engineers using coding agents to improve and accelerate their work by amplifying their existing expertise.

  4. OpenAI has some big questions. It doesn’t have unique tech. It has a big user base, but with limited engagement and stickiness and no network effect. The incumbents have matched the tech and are leveraging their product and distribution. And a lot of the value and leverage will come from new experiences that haven’t been invented yet, and it can’t invent all of those itself. What’s the plan?

February 21, 2026 10 min read

Links: Week of 22 Feb 2026

AI Links
  1. If AI has even a fraction of the impact that many people in Silicon Valley now expect on the fabric of work and daily life, it’s going to have profound and unpredictable political impacts.

  2. When 2012 passed into 2013, we did not have to rebuild our world, not in most countries at least. It sufficed to make adjustments at the margin.

    After the Roman Empire fell, parts of Europe had to rebuild their worlds. It took a long time, but they ended up doing pretty well.

    After the American Revolution, the newly independent colonies had to rebuild their own world. They did so brutally, but with considerable success.

    After WWII, Western Europe had the chance to rebuild its own world, and did a great job.

    We moderns are not used to having to rebuild our world.

    It is now the case that strong AI is here/coming, and we will have to rebuild our own world. Many of us are terrified at this prospect, others are just extremely pessimistic. It seems so impossible. How are all the new pieces supposed to fit together? Who amongst us can explain that process in a reassuring way?

    Yet we have done it many times before. Not always with success, however. After WWI ended, Europe was supposed to rebuild its own world, but they came up with something far worse than what they had before. Nonetheless, in the broader sweep of history world rebuilding projects have had positive expected value.

    And so we will rebuilding our world yet again. Or maybe you think we are simply incapable of that.

    As this happens, it can be useful to distinguish “criticisms of AI” from “people who cannot imagine that world rebuilding will go well.” A lot of what parades as the former is actually the latter.

    In any case, it all will be quite something to witness.

  3. Strap in. This is the most exciting time for business and technology, ever.

  4. I think we've just disrupted decades of existing intuition about sustainable working practices. It's going to take a while and some discipline to find a good new balance.

  5. SK
    Séb Krier@sebkrier · Feb 8

    Every time a model card drops, a lot of people screenshot scary parts - blackmail, evaluation awareness, misalignment etc. Now this is happening again, but instead of it being confined to a niche part of the safety community, it’s established commentators who are looking for things to say about AI.

    I want to make an honest attempt at demystifying a few things about language models and unpacking what I think people are getting wrong. This is based on a mixture of my own experimentation with models over the years, and also the excellent writing from @nostalgebraist, @lumpenspace, @repligate, @mpshanahan and many parts of the model whisperer communities (who may or may not agree with some of my claims). Sources at the bottom.

    In short: many public readings of some evaluations implicitly treat chat outputs as direct evidence of properties inherent to models, while LLM behavior is often strongly role- and context-conditioned. As a result commentators sometimes miss what the model is actually doing (simulating a role given textual context), design tests that are highly stylized (because they don't bother to make the scenarios psychologically plausible to the model), and interpret the results through a framework (goal-directed rational agency) that doesn't match the underlying mechanism (text prediction via theory-of-mind-like inference).

    Here I want to make these contrasts more explicit with 5 key principles that I think people should keep in mind:

    1. The model is completing a text, not answering a question

    What might look like "the AI responding" is actually a prediction engine inferring what text would plausibly follow the prompt, given everything it has learned about the distribution of human text. Saying a model is "answering" is practically useful to use, but too low resolution to give you a good understanding of what is actually going on.

    Lumpenspace describes prompting as "asking the writer to expand on some fragment." Nostalgebraist notes that even when the model appears to be "writing by itself," it is still guessing what "the author would say."

    Safety researchers sometimes treat model outputs as expressions of the model's dispositions, goals, or values — things the model "believes" or "wants." When a model says something alarming in a test scenario, the safety framing interprets this as evidence about the model's internal alignment. But what is actually happening is that the model is simply producing text consistent with the genre and context it has been placed in. The distinction is important because you get a richer way of understanding what causes a model to act in a particular way.

    A model placed in a scenario about a rogue AI will produce rogue-AI-consistent text, just as it would produce romance-consistent text if placed in a romance novel. This doesn't tell you about the model's "goals" any more than a novelist writing a villain reveals their own criminal intentions. Consider how models write differently on 4claw (a 4chan clone) vs Moltbook (a Facebook clone) in the OpenClaw experiments.

    2. The assistant persona is a fictional character, not the model itself

    In practice we should distinguish between (a) the base model (pretrained next-token predictor), and (b) the assistant persona policy (a post-hoc fiction layered on through instruction tuning + preference optimization like RLHF/RLAIF). Post-training creates a relatively stable assistant-like attractor, but it’s still a role: the same underlying model family can be steered into different "characters" under different system prompts, fine-tunes, and reward models.

    In their ‘The Void’ essay, Nostalgebraist also specifies that the character remains fundamentally under-specified, a "void" that the base model must fill on every turn by making reasonable inferences. I think characters today are getting more coherent and the void is not as large, partly because each successive base model trains on exponentially more material about what "an AI assistant" is like - curated HHH-style dialogues, but also millions of real conversations, blog posts analyzing model behavior, AI twitter discourse, academic papers, system cards, and so on. The character stabilizes the same way any cultural archetype does, i.e. through sheer accumulation of description.

    In practice, evaluating the character for its various propensities and dispositions remains useful! These simulated behaviours matter a lot, particularly if you're giving these simulators tools and access to real world platforms. But many discussions and papers just take the persona at face value and make all sorts of claims about 'models' or 'AI' in general, rather than the specific character that is being crafted during post-training. The counter-claim is that there is no stable agent there to evaluate. The assistant is a role the model plays, and it plays it differently depending on context, just as a base model would produce different continuations for different text fragments. Evaluating the model for "alignment" is like evaluating an actor for the moral character of their roles.

    3. Apparent errors are often correct completions of the world implied by the prompt

    This is increasingly less of an issue as we're getting much better at reducing 'mistakes' and 'hallucination' through post-training, retrieval, tool use, and decoding/verification. But it's helpful to take a step back and remember what it was like when these errors were omnipresent.

    Lumpenspace demonstrates this with the Gary Marcus bathing-suit example (see here:

  6. When the lights dimmed at Jaideep Sharma’s wedding reception in the north Indian city of Ajmer, guests expected to see a cheesy montage of the young couple in various attractive locations. Instead, they saw Sharma’s father — dead for more than a year — on the screen, smiling and blessing the newlyweds.

Other Stuff
  1. I think we’re going to do great things together. I think we’ll make the most of these beautiful years we have left, and I hope they’ll last very long.

    Amen.

  2. The most important thing I've learned about hospitals over the last decade: if your loved one needs to be admitted to the hospital, chances are they will get incredible care... as long as that care can be immediately administered in the ED.

    However, if they need to move outside the ED, you must learn as much as you can so you can help expedite the process, advocating to them to get to where they need to go — usually an inpatient floor, as quickly as possible.

    The stakes are probably higher than you think.

  3. No, this is not an AI post. Codex is a NYC bookshop at 1 Bleecker St., at Bowery. It is quite extraordinary in its curation of used books. The fiction section is large, yet you can pick up virtually any title on the shelves and it is worth reading. A wonderful place to go to get reading ideas, plus the prices are reasonable and the used books are in decent shape. Such achievements should be praised.

  4. This post will do two things:

    1. Establish that our best data show crime rates are historically low

    2. Argue that this is a real effect, not just reporting bias (people report fewer crimes to police) or an artifact of better medical care (victims are more likely to survive, so murders get downgraded to assaults)

  5. RJ
    Rob Johnson@FreeRangeLawyer · Jan 13

    Housing permits for new multifamily construction in Montgomery County, MD, before and after rent control.

  6. Some of the biggest stars to emerge from this year's Super Bowl halftime show never even showed their faces on camera. They were the ones who dressed as bunches of grass to transform a football stadium into the sugarcane fields of Puerto Rico.

×