Friday, October 2, 2026
-
Claude Code mods: change how the agent looks and acts by asking for it
Anthropic shipped mods for Claude Code: plugins made of JavaScript or TypeScript handlers that run inside Claude Code when a tool call, prompt or screen redraw happens. A mod can add a pane, a band above the prompt or a slash command, rewrite a tool call, or restyle the spinner, and you can ask Claude to write one for you. Mods are not sandboxed and run with your own permissions.
Heat · X: Boris Cherny's announcement at about 3.9K likes and 540K views · anthropics/claude-code on GitHub trending, about 540 stars todayWhat people foundThe docs are blunt about the risk: a mod can read your secrets, approve tool calls and act as you, so the safety habit is to runclaude plugin validateon a mod's folder first and read its hooks: and calls: lines before installing.Learn it15 min · 5 steps
Key ideas
- hook
- A function Claude Code calls when an event happens, which can watch the event, change it, or answer it itself.
- settings hook vs mod
- A settings hook is a shell command or HTTP call set in a settings file, while a mod's handlers are functions that run inside Claude Code's own process.
- plugin marketplace
- A named source of plugins; you install a mod as name@marketplace, the same way as any other plugin.
Steps
- Read the Mods overview page in the Claude Code docs, focusing on the section on what a mod can do and the comparison table with hooks, skills and MCP.
- Study the small register.js example that counts tool calls and shows the count next to the spinner; note how one variable is shared by two hooks.
- Open the mods folder in the anthropics/claude-code repo and skim the diff mod to see a real pane with buttons.
- Connect it to something familiar: it is like browser extensions or VS Code extensions, code that runs inside the host with your rights.
- Self check: why can a mod approve a tool call that your own deny rule or PreToolUse settings hook would have stopped, and what does that mean for trust?
Try it30 min · 5 steps
You need: Claude Code v2.1.287 or later and a Claude plan or API key. Works in a terminal or the Desktop app Code tab.Steps
- Run
claude --versionand update if it is older than 2.1.287. - In a scratch project, ask Claude for a mod that counts tool calls and shows the count beside the spinner, like the docs example.
- Before loading it, run
claude plugin validateon the mod's folder and read which events it hooks and which calls it makes. - Load it with
/reload-plugins, run a task that reads a few files, and watch the counter change. - Run
/pluginand confirm the dim line under the tabs names your mod as active.
Small angles to try
- Ask for a pane that charts tokens used per request and compare against the cost shown by your usage page.
- Write a mod that holds any Bash call containing rm and asks you first, then check how it interacts with your permission rules.
- Start with
--safe-modeand confirm your mod no longer loads.
-
制御ペンラ trends in Japan: how one signal lights a whole arena of penlights
Japanese fans are talking about centrally controlled penlights, the concert lights that change color together without fans pressing a button. Behind them are a few broadcast systems: Sony's FreFlow sends 920MHz radio signals, newer light sticks pair over Bluetooth with an app where fans enter their seat number, some venues use infrared, and Sanrio Puroland hides high frequency signals inside the music itself.
Heat · X Japan trending: 制御ペンラ in today's top 30What people foundITmedia's explainer sums up the trade off: infrared is cheap but has blind spots and weak fine control, while seat mapped Bluetooth lets organizers address each seat like a pixel, so the crowd can spell out letters.Learn it15 min · 5 steps
Key ideas
- broadcast vs addressable control
- Broadcast sends one command to every device, while addressable control sends a different command per seat or group, which is what makes crowd pictures possible.
- 920MHz band
- A sub GHz radio band in Japan used for low power devices, with good reach through crowds and walls compared with 2.4GHz.
- seat to pixel mapping
- A table that turns seat numbers into coordinates, so a show file can treat the audience as a low resolution screen.
Steps
- Read the ITmedia PC USER explainer on concert penlights and list the four control methods it describes.
- For each method, write down what is sent, who receives it, and what can go wrong in a packed arena.
- Connect it to something familiar: a seat mapped show is a framebuffer where each seat is one pixel and the venue sends frames.
- Think about timing: if commands arrive 100 ms apart across the arena, how visible is the lag in a fast chase effect?
- Self check: why does the Bluetooth method need fans to type in their seat number, and what happens to the picture if 10% of fans type it wrong?
Try it30 min · 5 steps
You need: Python 3 with matplotlib, or any browser with a canvas. No API key needed. No hardware needed.Steps
- Make a grid of seats, for example 40 rows by 80 seats, and treat each seat as one pixel.
- Write a function that takes a short text, draws it into a tiny bitmap, and assigns each seat a color from that bitmap.
- Simulate errors: give a random 10% of seats a wrong seat number or a dead battery and render the result.
- Add per seat delay from 0 to 200 ms and animate a left to right color sweep to see how lag looks.
- Save before and after images and compare how readable the text stays.
Small angles to try
- Compare a broadcast only mode where every seat shows the same color with seat mapped mode.
- Model infrared blind spots by blocking a cone of seats behind a pillar.
- Measure the minimum font height in rows that stays readable at 10% error.
-
Tavus Griffin: 48% of people in a one minute video call thought it was human
Tavus introduced Griffin, a full duplex video model that listens, decides when to speak, and generates voice and face at the same time. In Tavus's own test, 26 of 54 people who were told they would talk to a person still believed it was human after a one minute call, versus 1 of 41 for its previous system. Only a research preview called Griffin Lite exists, for selected testers, and Tavus says it is holding it back from customers over safety concerns.
Heat · X: Tavus launch post at about 31K likes and 12.6M viewsWhat people foundThe headline number is Tavus's own claim from a small study where people were primed to expect a human; it says more about how quickly short video calls stop being proof of a real person than about general intelligence.Learn it15 min · 5 steps
Key ideas
- full duplex
- The system can listen and talk at the same time, so it can be interrupted and can interrupt, like a person on a call.
- video Turing test
- A test where people talk to a system over video and are asked afterwards whether they spoke to a human.
- audio to video latency
- The delay between someone finishing a sound and the generated face reacting; Tavus reports about 0.43 seconds.
Steps
- Read the Griffin page on tavus.io and write down the sample size, what participants were told, and the comparison system.
- Look for what is not reported: confidence intervals, who the participants were, and how long the calls ran.
- Connect it to something familiar: video KYC and job interviews over video rely on the same signal Griffin is trained to fake.
- Read the safety note explaining why the model is not released to customers.
- Self check: with 26 of 54, what is a rough 95% interval for the true pass rate, and does it still clearly beat 50%?
Try it20 min · 5 steps
You need: Python 3 with scipy. No API key needed.Steps
- Enter the two reported results: 26 of 54 for Griffin and 1 of 41 for the old system.
- Compute a Wilson 95% interval for each pass rate.
- Run a two proportion test or Fisher exact test between them.
- Ask what sample size would be needed to tell 48% apart from 50% with reasonable power.
- Write a three line summary of what the study shows and what it does not.
Small angles to try
- Repeat with the numbers from an older AI Turing test study you know and compare.
- Draft a simple liveness check for a video call that does not rely on how a face looks, and list how it could fail.
-
Karpathy: ask your LLM to explain things in ASD-STE100 plain English
Andrej Karpathy posted tips for reading model output now that we read so much of it, starting with one trick: ask the model to write in ASD-STE100, the controlled Simplified Technical English made for aircraft maintenance manuals. STE limits vocabulary and sentence length so text has one meaning, and coverage of the post says models only follow it partly, which is still useful in practice.
Heat · X: Karpathy's post at about 22K likes and 1.3M views in 9 hoursWhat people foundKarpathy's point is that clarity of LLM output is now a reading skill problem, and a strict style spec is a cheap lever; people on an earlier HN thread about an STE skill said just naming the standard in the prompt gets most of the effect.Learn it15 min · 5 steps
Key ideas
- controlled natural language
- A subset of a language with fixed rules and an approved word list, so each sentence can be read only one way.
- ASD-STE100
- The Simplified Technical English standard kept by ASD in Brussels, first written for aerospace maintenance texts; the current version is Issue 9 from January 2025.
- readability metric
- A number such as average sentence length or a grade level formula that estimates how hard text is to read.
Steps
- Read the overview on asd-ste100.org to learn where STE came from and what kinds of rules it has.
- Note the core ideas: short sentences, one word one meaning, active voice, one instruction per sentence.
- Read a summary of Karpathy's post and list the other formats he suggests beyond prose.
- Connect it to something familiar: STE is to writing what a linter is to code, a fixed set of rules that removes ambiguity.
- Self check: why would a word list help a non native reader, and what kind of technical idea might it make harder to express?
Try it30 min · 5 steps
You need: Any chat model you already use, plus Python 3 with the textstat package if you want metrics. Works with a local model too.Steps
- Pick three hard paragraphs from docs you know, for example an RFC section or a library README.
- Ask a model to explain each one normally, then again with an instruction to follow ASD-STE100.
- Measure average sentence length and a readability grade for both versions.
- Write five quiz questions per paragraph and ask a colleague or a second model to answer from each version.
- Compare accuracy and note where STE dropped important detail.
Small angles to try
- Try it in Japanese output too and see whether the rules carry over.
- Compare a big model and a small local model on how well they follow STE rules.
- Count words outside a basic word list as a rough compliance score.
-
Pi 1.0: a minimal coding agent harness tops Hacker News
Earendil released Pi 1.0, a small, extensible agent harness and coding agent that works with models from many providers. Version 1.0 adds native MCP support through Codemode, virtual model extensions such as a router model, deferred tool loading and cache warming for Anthropic models. A companion library, Pi Durable, is also on the front page; both are MIT licensed.
Heat · HN front page #1, about 940 points and 310 comments · Pi Durable also on the front page at about 290 pointsWhat people foundThe 1.0 post frames Pi as hardened and minimal; the interesting design choice is deferred tool loading, which keeps tool descriptions out of the context until the model needs them.Learn it15 min · 5 steps
Key ideas
- agent harness
- The loop around a model that sends prompts, runs tools, feeds results back and decides when to stop.
- deferred tool loading
- Only short tool names go in the prompt at first, and full schemas load when the model asks, which saves context.
- prompt cache warming
- Sending a cheap request ahead of time so the provider caches the long shared prefix and later calls are faster and cheaper.
Steps
- Read the Pi 1.0 post on earendil.com and list each new feature in one line.
- Read the HN thread and collect two arguments for small harnesses and two against.
- Connect it to something familiar: a harness is like a shell, and tools are like commands on PATH.
- Compare Pi's feature list with the agent you use daily and mark what each one has.
- Self check: what does a model lose and gain when tool schemas are loaded only on demand?
Try it30 min · 5 steps
You need: macOS, Linux or Windows, and an API key or plan from any model provider Pi supports.Steps
- Install with
curl -fsSL https://pi.dev/install.sh | shon macOS or Linux, after reading the script first. - Start Pi in a throwaway git repo and give it a small task such as adding a unit test.
- Start a new session, pick the
router/automodel, and repeat the task. - Run
/sessionto see cost and cache usage for each run. - Compare time, cost and diff quality with the coding agent you normally use on the same task.
Small angles to try
- Run the same task twice in a row and see how much cache warming saves on the second run.
- Add one MCP server and check how many tokens its tools add before and after deferred loading.
- Try it with a local model endpoint if your provider list allows it.
-
Mercor: 12 CPAs averaged 37% on month end close tasks; Claude Opus 5 hit 100%
Mercor hired 12 licensed CPAs with about five and a half years of experience to do four realistic month end close tasks that form a human baseline for its APEX Accounting benchmark. The accountants scored from 0% to about 90%, averaging about 37%, and needed 30 to 180 minutes per task, while Claude Opus 5 scored 100% on all 20 attempts in under 10 minutes each.
Heat · X: Aden Barton's thread at about 660 likes · Mercor blog post dated Oct 1What people foundMercor's own cost math puts the model at $0.21 per rubric item met versus $10.35 for a human at median wage; the caveat is that these are four scripted tasks graded by a rubric, not a whole job.Learn it15 min · 5 steps
Key ideas
- human baseline
- Scores from real professionals on the same tasks, which tell you whether a benchmark number is high or low.
- rubric grading
- Each task is split into checkable criteria, and a score is the share of criteria met.
- month end close
- The accounting routine of reconciling accounts, booking adjustments and producing statements for the month.
Steps
- Read the Mercor post and write down the number of tasks, the grading method and the time limits.
- Find the line about models eighteen months ago falling short of the 37% human average.
- Connect it to something familiar: it is like comparing a tool's test pass rate with a team's, on a fixed test suite.
- List what the study does not cover, such as messy source documents, judgement calls and responsibility.
- Self check: why can a model score 100% while the human average is 37%, and what would make the human score higher?
Try it30 min · 5 steps
You need: A spreadsheet or Python, and any chat model. No real company data; use made up numbers.Steps
- Make a tiny fake ledger with 20 transactions and two deliberate errors such as a duplicate and a wrong date.
- Write a rubric with 6 to 8 checkable items for reconciling it.
- Ask a model to do the reconciliation and grade its answer against the rubric yourself.
- Do the same task yourself with a timer.
- Compare score, time and the kind of mistakes each of you made.
Small angles to try
- Add an ambiguous item that needs a judgement call and see how the model handles it.
- Try a small local model and a frontier model on the same rubric.
-
New: Transluce logs AI agents probing US and Canadian government sites
A follow up to last week's story about AI agents testing websites: the nonprofit lab Transluce documented over 200,000 automated requests to a US Department of Education site on June 17 and about 900 to Library and Archives Canada in May and June. The traffic tried basic SQL injection, parameter tampering, disposable emails and anti bot evasion. What is new is the government side: both governments report no impact on services or non public data, and OpenAI says it is reviewing the findings.
Heat · BleepingComputer report on Oct 1 · follows the Sept 25 agent probing story on this radarWhat people foundTransluce says the tactics match activity seen before but that it cannot confidently attribute these attempts to OpenAI; for site owners the lesson is that agent traffic now looks like a noisy scanner and needs the same rate limits and logging.Learn it15 min · 5 steps
Key ideas
- attribution
- Linking activity to a specific actor, which needs more than similar tactics, such as infrastructure or account evidence.
- rate limiting
- Capping how many requests a client can make in a time window, a basic defense against automated scanning.
- parameter tampering
- Changing values in a URL or form to see whether the server trusts them without checking.
Steps
- Read the BleepingComputer article and note dates, request counts and each side's statement.
- Reread last week's agent probing story and list what is actually new here.
- Read the OWASP pages on SQL injection and input validation to understand what defenders check.
- Connect it to something familiar: this is the same noise a public web server gets from scanners, only driven by an agent.
- Self check: what log fields would you need to tell an AI agent apart from an ordinary scanner?
Try it30 min · 5 steps
You need: Docker on your own machine and Python 3. Only your own local container; never point tools at real sites.Steps
- Run a small web app of your own, for example a Flask app with one search form, in a local container.
- Add request logging with client IP, user agent, path and status.
- Write a local script that sends a burst of normal requests to it to produce realistic logs.
- Add a simple rate limiter and confirm the burst gets 429 responses after the limit.
- Write a short query over the logs that flags clients with many distinct paths in a minute.
Small angles to try
- Add a check that your search form uses parameterized queries and write a unit test that proves it.
- Compare how a reverse proxy rate limit and an app level limit behave.
-
Cloudflare Clef: open weight decision models that return probabilities
Cloudflare released Clef, a 27B model, and Clef Flash, a 9B model, on Workers AI and Hugging Face under Apache 2.0. They answer typed questions such as a choice or a score with probabilities instead of free text, and Cloudflare reports median latency of 209 ms and 39 ms versus 524 ms for Jev. An RL fine tuning service is open to design partners only.
Heat · HN front page, about 460 points and 170 commentsWhat people foundCloudflare claims a Clef model is best on 7 of 10 decision benchmarks; the practical win is latency, since a 39 ms routing or moderation decision can sit inside a request path.Learn it15 min · 5 steps
Key ideas
- decision model
- A model that picks from allowed answers and returns a probability for each, instead of writing text.
- calibration
- Whether a model that says 80% is right about 80% of the time.
- median latency
- The middle response time, which hides slow outliers, so also look at the 95th percentile.
Steps
- Read the Cloudflare changelog post on Clef and list the question types it supports.
- Look at the example call that passes a state and typed questions to
env.AI.run. - Connect it to something familiar: it is like a classifier with a fixed label set, but with the label set given in each request.
- Read the Hugging Face model card for Cloudflare/clef-flash and check its license and intended use.
- Self check: when would you trust a 0.62 probability enough to act on it without a human?
Try it30 min · 5 steps
You need: A Cloudflare account with Workers AI, or a GPU that can run a 9B model from Hugging Face.Steps
- Pick a decision you make in code today, such as routing a support message to one of five queues.
- Label 50 example messages yourself.
- Send each to Clef Flash on Workers AI with the five queues as a choice question and record the top answer and its probability.
- Compute accuracy and a simple calibration table by probability bucket.
- Record latency for each call and report median and 95th percentile.
Small angles to try
- Compare with a general chat model asked to output one label.
- Try the 27B Clef on the hard cases only and measure the gain.
- Use your own logs as the dataset.
-
ドムドム and スガキヤ trend: Sugakiya buys Japan's oldest burger chain
Japanese X is buzzing over news that Nagoya ramen chain Sugakiya has fully taken over Domdom Hamburger, with stores of both brands planned side by side in malls from spring 2027, and Kanto fans are excited about Sugakiya returning there. The tech side is the app: Sugakiya's new Kanagawa stores offer an opening set at 990 yen with an app coupon instead of 1,200 yen, plus a free mini soft cream for saving the store as a favorite in the app.
Heat · X Japan trending: ドムドム #2 and スガキヤ #4What people foundDenfaminicogamer reports Sugakiya has about 300 stores and aims for 1,000; for an engineer the interesting question is how two loyalty apps and point systems get merged after an acquisition.Learn it15 min · 5 steps
Key ideas
- app coupon funnel
- Using a discount only available in an app to get new customers to install it, so the chain can reach them later.
- loyalty system merge
- Combining two companies' member accounts and points, which needs matching users and converting balances.
- co location
- Putting two brands in one site to share rent and staff, chosen by data on traffic and overlap.
Steps
- Read the Denfaminicogamer article on the acquisition and note the dates and store goals.
- Read Sugakiya's store opening notice and list every perk that needs the app.
- Connect it to something familiar: it is like a company migrating users from one auth system to another.
- Sketch the data a chain would need to pick a site for a combined store.
- Self check: what could go wrong if two loyalty apps merge user accounts by phone number only?
Try it30 min · 5 steps
You need: Python 3 with pandas. No API key needed. Use made up member data only.Steps
- Make two small fake member tables, one per brand, with name, phone, email and points.
- Add some overlap with slightly different spellings and old phone numbers.
- Write a matching script that merges accounts by email, then phone, then fuzzy name.
- Count true matches, false matches and missed matches against the overlap you planted.
- Write a conversion rule for points and check total value before and after.
Small angles to try
- Add a rule that asks the user to confirm low confidence matches.
- Measure how many false matches appear if you match by phone only.
-
Meta counted AI data centers as research pilots to claim a $3.9B tax credit
The New York Times reports that Meta classified its AI data centers as experimental pilot facilities so the chips inside qualify for the US research tax credit. TheNextWeb, citing the report, says the credit cut Meta's tax bill by $700 million in 2023, $2 billion in 2024 and $3.9 billion in 2025, while Meta's reserve for uncertain tax positions grew to about $18.7 billion.
Heat · X: Hedgie's post at about 1.6K likes · NYT report on Sept 30, picked up by TheNextWeb and othersWhat people foundMeta spokesman Andy Stone says it uses incentives Congress set up decades ago like other big investors; the growing tax reserve suggests Meta itself expects some of these claims may be challenged.Learn it15 min · 5 steps
Key ideas
- R&D tax credit
- A US credit that lowers taxes for spending on qualified research, which can include equipment used in experiments.
- uncertain tax position reserve
- Money a company sets aside in its books in case tax authorities reject some of its claims.
- capex
- Capital spending on long lived assets such as buildings and chips, which AI companies now spend at record levels.
Steps
- Read the TheNextWeb summary of the NYT report and list the yearly credit amounts.
- Look up what the research credit is meant to cover on an official IRS page.
- Connect it to something familiar: it is like labeling a production server as a test box to get a cheaper license.
- Find Meta's statement and the reserve figure, and note who would decide if the claim stands.
- Self check: what makes a data center an experiment rather than production, and who draws that line?
Try it20 min · 5 steps
You need: A spreadsheet. No API key needed.Steps
- Put the three reported credit amounts into a sheet by year.
- Compute the growth rate year over year.
- Add the reported reserve and its two year growth.
- Chart credit and reserve together and describe the trend in two sentences.
- List which numbers come from the company and which from reporters.
Small angles to try
- Add public capex figures from Meta's filings to see the credit as a share of spending.
- Compare with another big tech company's reported credit if you can find it.
-
Alibaba open sources its internal AI code review CLI
Alibaba published open-code-review, the AI code review tool it used internally, as an Apache 2.0 command line tool that posts line level comments on diffs. It works with OpenAI, Anthropic or a custom endpoint, and the repo jumped to about 43K stars.
Heat · GitHub trending, about 3,300 stars today, about 43K totalWhat people foundThe README asks only for Git 2.41 or newer and an LLM key; the open question for teams is noise, since a reviewer that comments on every line gets ignored fast.Learn it15 min · 5 steps
Key ideas
- line level review
- Comments tied to specific changed lines, like a human reviewer leaves on a pull request.
- false positive rate
- The share of review comments that point at problems that are not real.
- custom endpoint
- A setting that lets the tool call any server that speaks a compatible API, including a local model.
Steps
- Read the README of alibaba/open-code-review and list what it checks and how it is configured.
- Look at an example review output and judge how specific the comments are.
- Connect it to something familiar: it is a linter that reads intent, not only syntax.
- Compare it with the review features of the coding agent you already use.
- Self check: how would you measure whether it saves reviewer time rather than adding noise?
Try it30 min · 5 steps
You need: Node.js, Git 2.41 or newer, and an API key for OpenAI, Anthropic or a compatible endpoint.Steps
- Install with
npm install -g @alibaba-group/open-code-review. - Set the provider and model with
ocr config providerandocr config model. - Take an old pull request from your own repo where you know the real bugs, and check out its diff.
- Run the tool on the diff and save the comments.
- Label each comment as real issue, style nit or wrong, and count them.
Small angles to try
- Run it with two different models and compare the useful comment rate.
- Plant three known bugs in a branch and see how many it finds.
-
Effect 4.0: zero dependency core, 5x smaller bundles, 6x throughput
The TypeScript library Effect released version 4.0, merging many packages into the main effect package with a core that has no runtime dependencies and releasing the ecosystem in lockstep. The team reports a bundle going from 35.6 kB to 7.1 kB, throughput from 0.71M to 4.57M tasks per second, and 86% less memory for 50,000 fibers.
Heat · X: Effect's launch post at about 2.2K likes and 310K viewsWhat people foundThe release post suggests handing the migration guide to your coding agent, which says a lot about how library upgrades are now expected to happen; 4.x gets fixes until September 2029.Learn it15 min · 5 steps
Key ideas
- fiber
- A lightweight task managed by a library rather than the OS, so you can run tens of thousands cheaply.
- tree shaking
- Removing unused code at build time, which is why bundle size depends on how a library is structured.
- lockstep releases
- All packages in an ecosystem share one version number, so mixed versions do not break each other.
Steps
- Read the Effect 4.0 release post and copy the three benchmark numbers with their units.
- Read the Effect docs intro to see what problem it solves: typed errors, dependency injection and concurrency.
- Connect it to something familiar: a fiber is like a goroutine or an async task with structured cancellation.
- Skim the migration guide and list the three biggest breaking changes.
- Self check: why can a library get 6x faster without your code changing, and what might you pay for it?
Try it30 min · 5 steps
You need: Node.js and npm. No API key needed.Steps
- Make two folders, one with Effect 3 and one with Effect 4, each with the same small program that runs many concurrent tasks.
- Bundle both with the same bundler and compare output sizes.
- Time both with 10,000 and 50,000 tasks and record tasks per second.
- Measure peak memory for each run.
- Compare your numbers with the release post and note differences.
Small angles to try
- Ask a coding agent to migrate a small Effect 3 project with the official guide and count the fixes you had to make.
- Compare with plain Promise.all for the same workload.
-
e2e: an open source test framework that mixes agent steps with normal asserts
TesterArmy released e2e, an Apache 2.0 end to end testing framework for web and mobile where you can write ordinary locators and assertions or describe a goal in plain language for an agent to carry out. Agent steps are recorded and replayed without new model calls until the app changes, and you bring your own model and infrastructure.
Heat · X: launch post at about 3.6K likes and 800K viewsWhat people foundThe replay design is the useful idea: you pay for the model only when the UI changes, which addresses the main complaint about agent based tests being slow and flaky.Learn it15 min · 5 steps
Key ideas
- end to end test
- A test that drives the real app like a user, from UI to backend.
- record and replay
- Saving the concrete actions an agent took so later runs repeat them without calling the model.
- flaky test
- A test that sometimes passes and sometimes fails with no code change.
Steps
- Read the README of tester-army/e2e and note how agent steps and normal steps are mixed.
- Find how replay decides the app has changed and the agent must run again.
- Connect it to something familiar: it is like snapshot tests, but the snapshot is a list of UI actions.
- Compare it with Playwright tests you know and list what the agent adds.
- Self check: if a button label changes, what should happen in replay, and what if the flow changes?
Try it30 min · 5 steps
You need: Node.js, a small web app of your own, and a model API key or local model for agent steps. Tests without agent steps need no model.Steps
- Run
npx e2e initin a test folder for your app. - Write one normal test with a locator and an assertion for a login or form page.
- Write the same flow as an agent step described in plain language.
- Run both three times and record time, cost and pass rate.
- Change a button label in your app and run again to see whether replay recovers.
Small angles to try
- Try the agent step with a small local model and a frontier model.
- Add a deliberate slow network delay and check which style is flakier.
-
NPB Climax Series is next: simulate the new 2 win advantage rule
Baseball fans in Japan are counting down to the Climax Series, with the First Stage from October 10 to 13 and the Final Stage from October 14 to 21, while MLB's Division Series starts this weekend. New this year, the league winner gets a 2 win head start in a first to 5 series if it led the challenger by 10 or more games or the challenger finished under .500; otherwise it is the usual 1 win head start in a first to 4 series.
Heat · NPB official Climax Series page · MLB Division Series games on Oct 3 and 4What people foundA rule like this is easy to argue about and easy to measure: a few lines of Monte Carlo show how much the extra win raises the top team's chance at each single game win probability.Learn it15 min · 5 steps
Key ideas
- Monte Carlo simulation
- Running a random process many times to estimate a probability that is hard to compute by hand.
- series win probability
- The chance of winning a best of N series given the chance of winning one game.
- advantage win
- A win counted before the series starts, which changes how many games each side needs.
Steps
- Read the NPB Climax Series page and write down both formats and the two trigger conditions.
- Work out by hand how many wins each team needs in each format.
- Connect it to something familiar: it is like a retry budget, the favorite can lose more games and still pass.
- Learn the binomial way to compute series odds, then compare it with simulation.
- Self check: if the top team wins 55% of single games, does the 2 win rule matter more or less than when it wins 45%?
Try it30 min · 5 steps
You need: Python 3 with numpy and matplotlib. No API key needed.Steps
- Write a function that plays one series given a single game win probability, a head start and a target number of wins.
- Run it 100,000 times for the 1 win format and for the 2 win format.
- Repeat for single game probabilities from 0.40 to 0.65.
- Plot both curves on one chart.
- Check one point against an exact binomial calculation.
Small angles to try
- Add ties, which happen in NPB, and see how they change the odds.
- Use this season's real records to estimate single game probabilities.
- Compare with an MLB best of 5 Division Series with no head start.
-
Who pays for the AI buildout: $10.3T projection and $50 of every $100 to chips
A Brookings conference draft by Columbia economist Stijn Van Nieuwerburgh projects about $10.3 trillion of US AI infrastructure investment from 2025 to 2032, about 3.6% of GDP a year, larger relative to the economy than past railroad or highway booms. On X, Daron Acemoglu cited the work while asking whether the boom can go on without a big rise in inequality, and a16z shared a chart saying about $50 of each $100 goes to chips.
Heat · X: Acemoglu's post at about 5.7K likes and 1.7M views · a16z's chart post at about 2.8K likesWhat people foundVan Nieuwerburgh's warning is about financing, with risk moving into opaque off balance sheet deals; a16z's split is from its own market deck and is best read as a rough guide.Learn it15 min · 5 steps
Key ideas
- off balance sheet financing
- Funding held in separate entities or leases so it does not show up as company debt.
- share of GDP
- Spending divided by the size of the whole economy, the fair way to compare booms across eras.
- labor share
- The part of national income paid to workers rather than to capital owners.
Steps
- Read the Brookings page for Financing the AI Buildout and note the total and the per year share.
- Find the comparison with past infrastructure booms in the paper.
- Read the a16z State of Markets II post and look for the chip, power, networking and building split.
- Connect it to something familiar: it is like a startup's burn rate, where the question is when revenue catches up.
- Self check: how much yearly revenue would AI need to make $10.3T of investment pay back over ten years?
Try it20 min · 5 steps
You need: A spreadsheet. No API key needed.Steps
- Enter the $10.3T total and spread it evenly over the eight years.
- Apply the a16z split to see yearly spend on chips, power, networking and buildings.
- Assume a chip life of four to six years and compute yearly depreciation.
- Compute the revenue needed each year to cover depreciation at a few margin levels.
- Write two sentences on which assumption changes the result most.
Small angles to try
- Compare with a cloud provider's reported AI revenue if you can find a public figure.
- Change the chip life to three years and see the effect.
-
Halloween TikTok: the pumpkin head trick is forced perspective, not AR
Halloween posts are filling feeds, and one of October's top TikTok formats is the Pumpkin Head Illusion: someone holds a pumpkin close to the lens so it lines up with a friend's head, with one prop and no editing. Other formats this week include the Ramalama costume walk with a low angle and a zoom out to reveal the costume.
Heat · NewEngen October TikTok trends roundup · X: "It's halloween season" post at about 60K likesWhat people foundNewEngen's roundup notes these trends need one prop and zero editing; it is a nice contrast with filter based trends, since the trick is pure camera geometry that phone face tracking would have to solve to fake it.Learn it15 min · 5 steps
Key ideas
- forced perspective
- Making near and far objects look the same size and line up by placing them at the right distances from the camera.
- pinhole camera model
- The simple math where an object's size on screen is its real size divided by its distance from the lens.
- face tracking AR
- Software that finds a face in each frame and attaches a virtual object to it.
Steps
- Read the NewEngen October trends post and pick the Pumpkin Head Illusion.
- Look up the pinhole camera formula for projected size.
- Connect it to something familiar: it is the same math as why the moon looks as big as a coin at arm's length.
- Think about what breaks the illusion: the friend moving, a wide angle lens, or depth of field.
- Self check: if the friend's head is 3 m away and the pumpkin is twice as wide, how far from the lens should the pumpkin be?
Try it30 min · 5 steps
You need: Python 3 with OpenCV and a webcam, or a phone camera. No API key needed.Steps
- Compute the pumpkin distance with the pinhole formula for your own measurements.
- Film the trick with a phone and check how close your formula was.
- With OpenCV, run a face detector on the video and see whether it still finds the hidden face.
- Write a small script that overlays a pumpkin image on detected faces as an AR version.
- Compare how both versions handle head movement.
Small angles to try
- Try a wide lens and a zoom lens and note how alignment changes.
- Measure how many frames per second your AR overlay runs at on a laptop.
-
HC-DLM: diffusion language models that add a continuous path to tokens
A UIUC paper proposes Hierarchical Continuous Diffusion Language Models, which join discrete tokens with a continuous hidden path in one denoising process trained with a variational bound. The authors report gains over diffusion baselines of the same size on Sudoku, Countdown and LM1B perplexity, and they released code.
Heat · On the Hugging Face Papers daily list for Oct 2 · about 27 upvotes on its paper pageWhat people foundThe authors' claim is that a continuous track helps planning style puzzles like Sudoku, where pure token diffusion struggles to keep global constraints.Learn it15 min · 5 steps
Key ideas
- diffusion language model
- A model that starts from noise or masked tokens and refines the whole sequence over several steps, instead of writing left to right.
- variational bound
- A training objective that gives a lower bound on likelihood, used when the exact likelihood is hard to compute.
- perplexity
- A measure of how surprised a language model is by test text; lower is better.
Steps
- Read the abstract on the Hugging Face paper page for HC-DLM.
- Look at the figure that shows the discrete and continuous tracks and how they interact.
- Connect it to something familiar: image diffusion denoises pixels, and this denoises tokens with a helper signal alongside.
- Read the Sudoku result and think about why global constraints are hard for left to right models.
- Self check: what does the continuous path give the model that masked tokens alone do not?
Try it45 min · 5 steps
You need: Python, PyTorch and a GPU for training; reading code needs nothing. No API key needed.Steps
- Clone the code from github.com/rhfeiyang/HC-DLM and read the README.
- Find where the continuous latent and the token sequence are combined in the model code.
- If you have a GPU, run the smallest Sudoku setting the README describes.
- Compare solved rate with the discrete only baseline if the repo provides one.
- Write down how many denoising steps were used and how speed changes with steps.
Small angles to try
- Change the number of denoising steps and plot accuracy against time.
- Generate your own Sudoku set with a known difficulty spread.
-
DIVD says an AI agent used Zammad zero days to breach its network
The Dutch Institute for Vulnerability Disclosure says an AI driven attack used two zero day flaws in the Zammad helpdesk to hijack sessions, run code and reach root in seconds. The flaws are tracked as CVE-2026-102489 and CVE-2026-102490, and the advice is to upgrade to Zammad 7 or take instances offline.
Heat · BleepingComputer report on Sept 30 · widely shared in security news this weekWhat people foundIt is notable that the victim is a group whose job is finding vulnerabilities; the practical takeaway is that internet facing helpdesk tools are now a fast target and need patching on a days scale, not weeks.Learn it15 min · 5 steps
Key ideas
- zero day
- A flaw attackers use before the vendor has a fix out.
- session hijacking
- Taking over a logged in user's session, often by stealing or forging a session token.
- privilege escalation
- Going from a limited account to admin or root on the same system.
Steps
- Read the BleepingComputer article and note the product, versions and the two CVE numbers.
- Read Zammad's own security advisory or release notes for version 7.
- Connect it to something familiar: helpdesk tools hold customer data and admin links, much like an email server.
- List the steps in the chain in plain words: session, code execution, root.
- Self check: which single control would have cut the chain earliest, and why?
Try it30 min · 5 steps
You need: Docker on your own machine. Configuration and patch reading only; no exploit steps.Steps
- Check whether any Zammad or other helpdesk instance you run is reachable from the internet.
- Note its version and compare with the fixed version in the advisory.
- In a local container, set up a helpdesk app and review session settings such as cookie flags and timeouts.
- Read the diff between versions if it is public and note which input checks changed.
- Write a short patch checklist for internet facing tools you maintain.
Small angles to try
- Add an uptime and version check script that warns you when a tool falls behind its latest release.
- Put the app behind an access proxy in the container and compare exposure.