Code415 Radar

by Ziyu Guo · @Code415zg

What people are paying attention to in AI, LLMs and CS today. Every item comes with a short path to learn it and a small test to try it yourself. Updated every evening, Japan time.

Monday, October 5, 2026

16 items · 焼肉きんぐ app leak hits 10.79M; Strata runs a 125B MoE on one GPU
  1. #1Trends × tech

    焼肉きんぐ app breach: 10.79M members leaked, and what a leak without passwords enables

    Monogatari Corporation, which runs the 焼肉きんぐ restaurant chain, said on October 5 that an outsider got into its official app system and exposed 10,788,963 of its 10,808,784 registered members, about 99.8%. Member ID, name, email and phone number were taken, while passwords, birthdays, point history and card data were not. Access was detected on October 2, the cause is still under investigation, and the company has reported to the privacy commission and the police.

    Heat · X Japan trending: 焼肉きんぐ #1, アプリ登録者 #11, 情報流出 #15 · Yahoo News Japan and ASCII coverage
    What people foundBecause almost every member lost the same three fields, this looks like a bulk read of one table rather than password guessing, but the company has not named a cause yet, so treat any attack story as speculation. The practical risk now is follow up phishing by SMS and email that pretends to be the restaurant; the company says it never asks for passwords or card numbers.
    Learn it15 min · 5 steps

    Key ideas

    Data minimization
    Store only the fields a feature truly needs, so a breach can only expose those fields.
    Smishing
    Phishing sent by SMS, which works well when attackers already know your name and phone number.
    Breach notification
    The legal duty in Japan to report a leak to the Personal Information Protection Commission and tell affected people.

    Steps

    1. Read the company's official notice first and list exactly which fields leaked and which did not.
    2. Read the ASCII article to see how Japanese tech press framed the incident and the timeline from October 2 to 5.
    3. Compare it with a breach you remember where passwords leaked: what changes for users when only contact data is exposed?
    4. Look at why card data was safe: the app never stored it, which is data minimization in action.
    5. Self check: if you were a member, which three messages in the next month would you treat as suspicious, and why?
    Try it30 min · 5 steps
    You need: Python 3 with the built in sqlite3 module; your own test data only. No API key needed.

    Steps

    1. Create a small SQLite table that mirrors a typical loyalty app: id, name, email, phone, password hash, birthday, points.
    2. Fill it with 1,000 fake rows generated by a short script.
    3. Write two queries that simulate what an attacker gets from a full read of the table versus a read of only the contact fields.
    4. For each column, write one line on whether the app truly needs it and what the user loses if it leaks.
    5. Drop or move the columns you marked unnecessary and rerun the leak simulation to compare the exposure.

    Small angles to try

    • Add a separate encrypted column for phone numbers and see what a raw table dump still reveals.
    • Write a simple rule that flags SMS text claiming to be from a restaurant and asking for a login or card number.
    • Measure how big the leaked file would be for 10.8 million rows of contact data.
  2. #2Tools & open source

    Strata runs the 125B Qwen3.8-Flash-Next MoE on one 12GB gaming GPU

    Strata is an MIT licensed engine that runs the 125 billion parameter Qwen3.8-Flash-Next mixture of experts model on home PCs with NVIDIA or AMD cards. It keeps the most used experts on the GPU, holds all experts in system RAM, lets the CPU handle the rest, and serves an OpenAI compatible API plus a browser chat. The README reports 53 to 94 tokens per second of generation on an RTX 5070 with 12GB, with at least 32GB RAM and about 80GB of disk.

    Heat · HN front page #2, about 660 points and 300 comments · GitHub about 12,000 stars
    What people foundHN readers pushed back on the headline: quietFalcon noted that generation speed is the easy half of MoE offloading and asked about prompt processing at long context, and others asked how much quality the 2 and 3 bit quantizations cost. Note that the HN title mentions a 4090, while the README's measured numbers are on a 5070.
    Learn it15 min · 5 steps

    Key ideas

    Mixture of experts
    A model where each token only uses a few expert sub networks, so most weights sit idle at any moment.
    Expert offloading
    Keeping hot experts in fast GPU memory and cold ones in RAM or on disk, trading memory for bandwidth.
    Quantization
    Storing weights in fewer bits, such as 2 or 3 bits, to shrink memory at some cost in accuracy.

    Steps

    1. Read the Strata README section on how the model is split across GPU, RAM and CPU.
    2. Skim the HN thread for the measured numbers people posted from their own cards.
    3. Read the Edge0 paper abstract, which studies the same memory wall problem by predicting expert routing one token ahead.
    4. Connect it to CPU caches: hot experts on the GPU behave like a cache, and a cache miss costs a trip to slower memory.
    5. Self check: why does a MoE model offload much better than a dense model of the same size?
    Try it45 min · 5 steps
    You need: Windows 10 or 11 or Linux, an NVIDIA RTX 20 series or newer or AMD RX 7900 or 9000 series card with at least 12GB VRAM, 32GB RAM, about 80GB free disk. No API key needed.

    Steps

    1. Clone or download the Strata repository from GitHub.
    2. On Windows double click START-HERE.bat; on Linux run ./setup.sh in the Strata folder, and let it pick a model size for your hardware.
    3. Wait for the roughly 70GB download, then open the local chat at http://127.0.0.1:8080.
    4. Send the same three prompts, one short, one with 4,000 tokens of context, one coding task, and write down prompt speed and generation speed for each.
    5. Point any OpenAI compatible client at the local API and run a small set of your own coding questions to judge quality.

    Small angles to try

    • Compare two quantization levels on the same prompts and note quality versus speed.
    • Watch GPU memory and system RAM use during a long context run.
    • Run the same questions against a hosted model and compare answer quality per minute of waiting.
  3. #3Engineering & CS

    Google pauses its open source bug bounty after a flood of invalid AI reports

    Google paused its Open Source Software Vulnerability Rewards Program, which since 2022 paid researchers for bugs in projects like Go, Angular and Bazel. Its stated reason is a significant rise in automated submissions, most of them not valid, and it promises an update in the first quarter of 2027. Google's other reward programs, including Cloud and patch rewards, stay open.

    Heat · TechCrunch, October 4 · BleepingComputer lead story, October 5
    What people foundThis is the bounty side of the same problem Linux maintainers raised last week: AI tools make reports cheap to file but not cheap to check, so triage cost moves onto maintainers. Google gave no volume numbers, so claims about how many reports were AI written are not public.
    Learn it15 min · 5 steps

    Key ideas

    Bug bounty
    A program that pays outside researchers for valid security bugs.
    Triage
    The work of reading a report, reproducing it and deciding if it is real and how severe it is.
    Signal to noise
    The share of reports that are valid; when it drops, reviewers spend most time on junk.

    Steps

    1. Read the TechCrunch article for Google's exact statement and what stays open.
    2. Read the BleepingComputer write up for the program's history and payout range.
    3. Compare with the Linux kernel CVE flood covered on October 3: same cause, different outcome.
    4. Think about what a maintainer would need from a report to reject it in under five minutes.
    5. Self check: what would make an AI generated report cheap to verify instead of cheap to write?
    Try it30 min · 5 steps
    You need: Any LLM you already use, a small open source project of your own, and its source code. No third party systems.

    Steps

    1. Pick a small project you own with a few hundred lines of code.
    2. Ask an LLM to list possible security bugs in it, and save the raw list.
    3. For each claimed bug, try to reproduce it with a unit test in your own project and time how long each check takes.
    4. Count how many claims were real, duplicates, or invented.
    5. Write a short report template that would have let you reject the invalid ones faster, for example requiring a failing test.

    Small angles to try

    • Repeat with a second model and compare valid rates.
    • Ask the model to include a reproduction test with each claim and see if the valid rate changes.
    • Measure minutes of review per valid bug for each model.
  4. #4Trends × tech

    光遺伝学 trends after the Nobel: switching neurons on and off with light

    The 2026 Nobel Prize in Physiology or Medicine went to Karl Deisseroth, Peter Hegemann and Georg Nagel for discoveries about light gated ion channels and optogenetics, and 光遺伝学 shot into Japan's X trends. The core idea is to put an algae protein, channelrhodopsin, into chosen neurons so that a pulse of light opens a channel and makes the cell fire. That gives researchers a switch with millisecond timing for specific cell types, which is now a standard tool in neuroscience.

    Heat · X Japan trending: 光遺伝学 #22 · CNN, Scientific American and Nobel press release coverage
    What people foundFor engineers the interesting part is the interface: light pulses act like a digital control signal for chosen neurons, so experiments look like timing and control problems, including closed loop setups where a recording decides when to fire the light.
    Learn it15 min · 5 steps

    Key ideas

    Channelrhodopsin
    A light sensitive ion channel from green algae that lets ions flow into a cell when hit by blue light.
    Optogenetics
    Using genes to make specific cells respond to light, then controlling them with light pulses.
    Closed loop control
    A setup where measurements from the brain decide in real time when to apply the light.

    Steps

    1. Read the Nobel press release for the official citation and a short history.
    2. Read the CNN article for a plain summary of who did what.
    3. Connect it to something familiar: a GPIO pin that you toggle with precise timing, but the pin is a neuron type.
    4. Look up one closed loop optogenetics example and note what signal triggers the light.
    5. Self check: why does targeting a cell type matter more than just stimulating a brain region with electricity?
    Try it30 min · 5 steps
    You need: Python 3 with numpy and matplotlib. No API key needed.

    Steps

    1. Write a simple leaky integrate and fire neuron in Python: a voltage that leaks toward rest and spikes when it crosses a threshold.
    2. Model a light pulse as an input current that is on for a set number of milliseconds.
    3. Drive the neuron with pulse trains at 5, 20 and 50 Hz and record when it spikes.
    4. Plot input pulses and spikes on one chart and note where the neuron stops following the light.
    5. Change the pulse width and find the shortest pulse that still causes a spike.

    Small angles to try

    • Add a second neuron that only receives light half the time and compare firing.
    • Add noise to the input and measure how reliable the spikes stay.
    • Build a tiny closed loop rule that only fires the light when the voltage is below a level.
  5. #5AI in the world

    Bad PDF redaction exposes Google data center water and power use in Nebraska

    Nebraska data centers had to file water and electricity reports and claimed the figures were trade secrets, but a reporter found the black boxes in the PDFs could be highlighted and copied. Google's Lincoln site showed 52.65 MW of peak demand and about 13.3 million gallons of water a year, and its Papillion site about 548 million gallons. Six reporting facilities together used about 765 million gallons a year.

    Heat · HN front page #10, about 310 points and 420 comments
    What people foundThe HN thread split: tptacek argued the water volume is small next to farm and city use, while aliasxneo asked why companies hide numbers that are on their side. The engineering lesson is older than AI: drawing a black box over text does not remove the text.
    Learn it15 min · 5 steps

    Key ideas

    Visual redaction
    Covering text with a shape, which leaves the underlying text in the file.
    True redaction
    Removing the text from the document itself before export.
    Peak demand
    The highest power a site draws at once, which drives grid planning more than average use.

    Steps

    1. Read the 1011 Now article for the numbers and how the redaction failed.
    2. Skim the HN thread for the debate on whether the water numbers are large.
    3. Compare 52.65 MW with the power use of a city you know to get a sense of scale.
    4. Look up how your PDF tool's real redaction feature differs from drawing a box.
    5. Self check: name two ways to verify that a redacted PDF really has no hidden text.
    Try it20 min · 5 steps
    You need: A PDF editor and a text extraction tool on your own machine, with a document you create yourself. No API key needed.

    Steps

    1. Make a test PDF with a fake secret number on it.
    2. Hide the number by drawing a black rectangle over it and export.
    3. Extract all text from the exported file with any text extraction tool and check whether the number appears.
    4. Redo it with your editor's real redaction feature and extract the text again.
    5. Write a short checklist for anyone on your team who shares redacted documents.

    Small angles to try

    • Check the PDF metadata and earlier versions inside the file too.
    • Try a scanned image PDF and see what OCR recovers.
    • Write a small script that fails if any listed secret string is found in an exported PDF.
  6. #6Tools & open source

    pstack's /correct turns repeated agent corrections into repo rules

    Lauren Tan, @poteto, added a /correct skill to pstack, her Cursor plugin of agent workflow skills, and says it ships in pstack 0.15.9. The skill mines past corrections for classes of mistakes and fixes each at the highest level that works: architecture first, then types and lint, then tests, with docs last. It also keeps a table that pairs each rule with what enforces it.

    Heat · X: 4,046 likes and 188K views on the launch post, plus a 2,032 like follow up · PR merged October 3
    What people foundHer point, in her words, is not to micromanage agents but to correct the environment that shapes their behavior. It fits last week's 'documentation, not memory' debate, but pushes further: prefer a type or lint rule over a prose note, because the agent cannot ignore a failing check.
    Learn it15 min · 5 steps

    Key ideas

    Agent skill
    A packaged instruction file an agent loads to do one kind of task the same way every time.
    Enforcement ladder
    Fixing a mistake with the strongest tool available: architecture, then types, then lint, then tests, then docs.
    Mistake class
    A group of similar errors that share one root cause and can be prevented by one rule.

    Steps

    1. Read the description of pull request 494 in cursor/plugins to see what /correct is meant to do.
    2. Read Flavio Copes' pstack overview for how the plugin is installed and the other skills it includes.
    3. Connect it to code review: a reviewer who keeps leaving the same comment should add a lint rule instead.
    4. Read the follow up X post about constraints in the codebase being freeing for humans and agents.
    5. Self check: for one mistake your agent repeats, which rung of the ladder could block it?
    Try it30 min · 5 steps
    You need: Cursor with plugins enabled and a repo of your own. No extra API key beyond your Cursor setup.

    Steps

    1. Install pstack in Cursor by following the install steps in Flavio Copes' guide or the plugin listing.
    2. Pick a repo where you have corrected your agent several times for the same thing.
    3. Run the /correct skill and read the mistake classes it finds.
    4. Review each proposed fix and note which rung it chose: types, lint, tests or docs.
    5. Run the agent on a fresh task that used to trigger the mistake and see if the new check catches it.

    Small angles to try

    • Try the same idea by hand in another agent tool by writing the lint rule yourself.
    • Count corrections per hour before and after for one week.
    • Compare a docs only fix with a lint fix for the same mistake.
  7. #7Trends × tech

    トリプル台風 trends: three typhoons at once, and AI versus physics track models

    Typhoons No. 27, 28 and 29 are active at the same time and トリプル台風 is trending in Japan. No. 27 was a very strong 925 hPa storm near Ogasawara on October 5, expected to pass east of Japan on October 6, while No. 28 crossed the date line from the eastern Pacific. tenki.jp notes 2026 is only the third year since 1951 with a typhoon in every month from January to October.

    Heat · X Japan trending: トリプル台風 #10 · Google Trends Japan: 今日の天気 10,000+ searches
    What people foundWeathernews says the AI models its typhoon team uses had track errors about 40% smaller than physics models, but physics models still do better on rain, so it uses both. Storms close together can steer each other, which makes days like this a hard test for any model.
    Learn it15 min · 5 steps

    Key ideas

    Track forecast
    A prediction of where a storm's center will be over the next days, usually shown as a cone.
    AI weather model
    A neural network trained on decades of past weather that predicts the next state much faster than a physics simulation.
    Fujiwhara effect
    Two nearby storms rotating around each other and changing each other's paths.

    Steps

    1. Read the tenki.jp forecaster post for the current three storms and the record facts.
    2. Read the Weathernews article on how its team mixes AI and physics models.
    3. Open the JMA typhoon position table to see what official track data looks like.
    4. Connect it to ensembles in ML: many runs with small changes give a spread, which is the forecast cone.
    5. Self check: why could an AI model beat physics on track but lose on rainfall?
    Try it30 min · 5 steps
    You need: Python 3 with pandas and matplotlib, and public JMA data. No API key needed.

    Steps

    1. Pick one finished 2026 typhoon from the JMA position table.
    2. Copy its positions into a CSV with time, latitude and longitude.
    3. Plot the track on a simple latitude and longitude chart.
    4. Save today's forecast positions for typhoon No. 27 from a news page or JMA, then compare them with the real positions once the storm ends.
    5. Compute the distance error in kilometers for each forecast time using the haversine formula.

    Small angles to try

    • Compare forecasts from JMA and another agency for the same times.
    • Plot all three current storms together and watch for paths bending toward each other.
    • Measure how many days the official position table lags behind the storm.
  8. #8Tools & open source

    ds4: antirez's small C engine for running big MoE models locally

    Salvatore Sanfilippo, antirez, the creator of Redis, published DwarfStar 4, a narrow C inference engine for high memory Macs, CUDA and ROCm machines under the MIT license. It targets DeepSeek V4 Flash first and also supports GLM 5.x and Qwen 3.8 Flash Next, with a CLI, an HTTP server and a small coding agent. The README says it was built with strong help from AI coding agents.

    Heat · HN about 360 points and 100 comments · GitHub about 23,500 stars
    What people foundOne HN commenter, neomantra, argued that hosted models on bigger hardware are still smarter and faster and did not recommend the local path yet. The fun part for engineers is the scope: one person, a narrow C codebase, a few model families, and agents helping write it.
    Learn it15 min · 5 steps

    Key ideas

    Inference engine
    The program that loads model weights and runs the math to produce tokens.
    Prefill versus decode
    Reading the prompt in parallel versus generating one token at a time, which stress different hardware.
    Narrow engine
    An engine that supports a few models very well instead of every model.

    Steps

    1. Read the ds4 README sections on supported models and hardware.
    2. Compare its scope with llama.cpp, which aims to support many models.
    3. Look at the reported prefill and generation speeds and note which is limited by memory bandwidth.
    4. Read the HN thread for real user numbers and objections.
    5. Self check: why can a narrow engine be faster than a general one on the same hardware?
    Try it45 min · 5 steps
    You need: A Mac with Apple silicon and lots of unified memory, or a supported CUDA or ROCm machine, plus disk for model weights. No API key needed.

    Steps

    1. Clone the repo with git clone https://github.com/antirez/ds4.git and enter the folder.
    2. Build with make on a Mac, or the make target the README lists for your GPU, such as make cuda-generic.
    3. Download a supported model following the README.
    4. Run ./ds4 -p "your prompt" with a short and a long prompt and note prefill and generation speed.
    5. Start ./ds4-server --ctx 32768 and send the same prompts from any HTTP client.

    Small angles to try

    • Compare speed and answers with llama.cpp on the same model and quantization.
    • Try the built in coding agent on a small task.
    • Test how speed drops as context grows.
  9. #9Engineering & CS

    大和証券 trends: a contractor's inquiry system leaks about 110,000 customers

    Daiwa Securities said on October 5 that a server run by an outside contractor for its customer inquiry system was accessed without permission. NHK reports about 110,000 customers affected, with names, email addresses and account numbers. A security news site reports the access ran overnight from October 2 to 3 and that Daiwa's own trading systems were not hit.

    Heat · X Japan trending: 大和証券 #12 · NHK coverage
    What people foundThis is a supply chain lesson: a support ticket tool held personal data and account numbers, so its security mattered as much as the core system. It was announced the same day as the 焼肉きんぐ leak; there is no public sign the two are related.
    Learn it15 min · 5 steps

    Key ideas

    Third party risk
    Security risk that comes from vendors who hold or process your data.
    Support ticket data
    Customer messages that often contain more personal detail than the main database expects.
    People versus records
    One person can have many records, so a leak count of records is usually larger than the count of people.

    Steps

    1. Read the NHK article for the official facts.
    2. Read the security site's write up for the contractor name and timeline, noting it is a secondary source.
    3. Think about what customers typically paste into a support form.
    4. Compare with item 1: here the weak point was a vendor, not the company's own app.
    5. Self check: which fields in a support system should be masked before staff or vendors see them?
    Try it30 min · 5 steps
    You need: Python 3 and your own sample text. No API key needed.

    Steps

    1. Write 50 fake support messages that sometimes include an email, phone number or account number.
    2. Write a small script with regular expressions that finds and masks those fields.
    3. Count how many sensitive values your masking catches and misses.
    4. Add a rule that refuses to store a message until masking runs.
    5. Write down which vendor systems in your own work would hold data like this.

    Small angles to try

    • Try an LLM for detection and compare it with the regex approach.
    • Measure masking speed on 100,000 messages.
    • Add Japanese format phone numbers and addresses to the test set.
  10. #10Trends × tech

    A free manga panel layout generator goes viral: procedural layouts in the browser

    A Japanese indie developer released a free web tool that generates many manga page layouts at once and lets you tweak them without starting over. You choose a style, panel count and page size, set chances for angled panels, bleed and splits, and export SVG or PNG for drawing apps. It runs entirely in the browser, and a seed number makes any layout reproducible.

    Heat · X: 20,346 likes and 25,493 bookmarks on the launch post
    What people foundThe bookmarks outnumber the likes, a sign people saved it to use. Under the hood it looks like recursive rectangle splitting with weighted randomness, and the mutation mode slightly changes ratios and angles of an existing layout, similar to a genetic algorithm step.
    Learn it15 min · 5 steps

    Key ideas

    Procedural generation
    Making content with rules and randomness instead of by hand.
    Seeded random generator
    A random number generator that gives the same sequence for the same seed, so results can be repeated.
    Recursive subdivision
    Splitting a rectangle, then splitting the pieces, until you reach the panel count.

    Steps

    1. Open the tool and generate a batch, then change the seed and see what changes.
    2. Try mutation mode on one layout and note which properties move.
    3. Read about binary space partitioning, the same idea used for game level layouts.
    4. Connect it to SVG: each panel is just a polygon you can export.
    5. Self check: how would you guarantee no panel gets too thin to draw in?
    Try it45 min · 5 steps
    You need: A browser and a text editor; plain JavaScript or Python. No API key needed.

    Steps

    1. Write a function that splits a page rectangle horizontally or vertically at a random ratio.
    2. Call it recursively until you have a target number of panels.
    3. Use a seeded random generator so the same seed gives the same page.
    4. Draw the result as SVG with a gutter between panels.
    5. Add a mutation function that nudges each split ratio a little and compare the results.

    Small angles to try

    • Add angled cuts by splitting along a tilted line.
    • Score layouts by panel size variety and keep the best of 50.
    • Export a page your favorite drawing app can open.
  11. #11Engineering & CS

    Citrix NetScaler SAML zero day exploited; researchers suspect more than DoS

    Citrix patched CVE-2026-88779, rated 8.7, in NetScaler ADC and Gateway when they are set up as a SAML service provider or identity provider, and confirmed targeted attacks on unpatched devices. Citrix calls it a denial of service flaw. BleepingComputer reports attack traffic since late September carrying commands that download further code.

    Heat · BleepingComputer, October 4
    What people foundKevin Beaumont reported honeypots running the downloaded malware, which suggests remote code execution beyond the official denial of service rating. If you run NetScaler with SAML, treat the patch as urgent rather than waiting for the rating to change.
    Learn it15 min · 5 steps

    Key ideas

    SAML
    An XML based standard for single sign on between an identity provider and a service.
    Edge device
    A gateway or VPN box at the network edge, a frequent target because it faces the internet.
    Severity rating
    The vendor's estimate of impact, which can lag behind what attackers actually achieve.

    Steps

    1. Read the BleepingComputer article for affected versions and fixed builds.
    2. Read Citrix's security bulletin linked from the article.
    3. Look up how a SAML login flow works so you know where the crafted request goes.
    4. Compare with earlier NetScaler incidents where ratings were later raised.
    5. Self check: which logs would show that your gateway had received these requests?
    Try it20 min · 5 steps
    You need: Read only work: the vendor bulletin and, if you run NetScaler, your own device's version and configuration. No exploit steps.

    Steps

    1. List every NetScaler appliance you are responsible for and its build number.
    2. Check whether each one is configured as a SAML service provider or identity provider.
    3. Compare the builds with the fixed versions in the bulletin.
    4. Write down the patch plan and who approves it.
    5. Review logs for unusual authentication requests since late September, following the vendor's guidance.

    Small angles to try

    • Write a small script that compares an inventory list with fixed build numbers.
    • Draft an internal note explaining why a DoS rating may understate risk.
    • Set an alert for future bulletins on your edge devices.
  12. #12Research

    On or off policy distillation? KL direction matters more, says a Cambridge study

    Researchers at Cambridge ran a systematic study of how small models learn from bigger ones during distillation. They found that the direction of the token level KL loss matters more than whether training rollouts come from the student or the teacher: forward KL is stable across rollout sources, while reverse KL is sensitive. Learning rate drives forgetting, and gains from on policy data do not always survive further training.

    Heat · Hugging Face Paper of the Day, October 2, about 175 upvotes
    What people foundFor anyone distilling a small model, the practical message is to choose the loss direction before worrying about expensive on policy rollouts, and to watch the learning rate for forgetting. No code release is listed.
    Learn it15 min · 5 steps

    Key ideas

    Distillation
    Training a small student model to copy the output distribution of a larger teacher.
    Forward versus reverse KL
    Two ways of measuring the gap between distributions; forward covers all of the teacher's options, reverse focuses on the teacher's main mode.
    On policy rollouts
    Training on text the student itself generates rather than text from the teacher or a dataset.

    Steps

    1. Read the abstract and the main findings on the Hugging Face paper page.
    2. Review forward and reverse KL with a two peak example on paper.
    3. Find the figure comparing loss directions across rollout sources.
    4. Connect it to RLHF, where on policy data is also expensive to produce.
    5. Self check: why would reverse KL make a student ignore rare but correct answers?
    Try it45 min · 5 steps
    You need: Python 3 with PyTorch on a CPU or small GPU. No API key needed.

    Steps

    1. Build a tiny teacher: a fixed probability distribution over 10 tokens with two peaks.
    2. Create a student with learnable logits for the same 10 tokens.
    3. Train one student with forward KL and one with reverse KL and record the final distributions.
    4. Plot both students against the teacher.
    5. Repeat with three learning rates and note which setting forgets earlier training fastest.

    Small angles to try

    • Make the teacher have three peaks and compare again.
    • Train the student on samples it draws itself versus samples from the teacher.
    • Measure how many steps each loss needs to match the main peak.
  13. #13Models & launches

    Opus 5.5 versus GPT-6.1 Sol: one prompt WebGPU shootouts spread on X

    X users are pitting Claude Opus 5.5 against OpenAI's GPT-6.1 Sol and ChatGPT-6 Astra with the same single prompt, then posting the results side by side. Examples include a soft body plush octopus with fur running on WebGPU in the browser and a 3D parallax landing page. Another post asking Claude to draw the inside of its own mind drew more than 4,700 likes.

    Heat · X: 4,741 likes and 247K views on @trikcode's Opus 5.5 build · 783 likes on the WebGPU octopus comparison
    What people foundThese comparisons are fun but are single samples with no fixed seed, so treat them as taste tests, not benchmarks. The poster of the octopus demo says both versions are pure code with no ready made models or textures; that is the poster's claim.
    Learn it15 min · 5 steps

    Key ideas

    One prompt test
    Giving two models the same single instruction and comparing the first output.
    Sampling variance
    The same model can give quite different outputs to the same prompt.
    WebGPU
    A browser API for running GPU graphics and compute from JavaScript.

    Steps

    1. Watch the octopus comparison post and note what exactly differs.
    2. Read a structured comparison of the two models from an evaluation site you trust.
    3. Learn the basics of WebGPU from MDN to understand what the demo needs.
    4. Think about how many runs you would need before trusting a difference.
    5. Self check: what would make a one prompt test fair?
    Try it45 min · 5 steps
    You need: Access to both models through their apps or APIs, and a browser with WebGPU support. API keys needed only if you use the APIs.

    Steps

    1. Write one prompt for a small WebGPU or canvas scene with clear requirements.
    2. Run it three times on each model and save every output.
    3. Open each result in the browser and note whether it runs, frame rate and visual quality.
    4. Score each run against your requirements on a simple 1 to 5 scale.
    5. Post the full table including failures, not just the best run.

    Small angles to try

    • Use a non visual task, such as a parser, with automatic tests.
    • Count tokens and time per run.
    • Ask each model to fix its own failed run once and score again.
  14. #14Tools & open source

    RemoveMacAI frees about 12GB by turning off Apple Intelligence on macOS 27

    A small open source command line tool disables Apple Intelligence features on macOS 27, such as Writing Tools, Genmoji and summaries, and removes the downloaded on device models. Its README says this frees about 12GB, uses configuration profiles instead of editing system files, and can be fully undone.

    Heat · HN front page #9, about 460 points and 280 comments
    What people foundThe HN thread focused on trust: several people objected to installing with a pipe to bash, and one asked how anyone can know it is safe. The tool offers a dry run and a status command, which is the right place to start.
    Learn it15 min · 5 steps

    Key ideas

    Configuration profile
    A macOS mechanism for setting managed policies, often used by companies for their Macs.
    On device model
    A model stored and run locally rather than in the cloud.
    Pipe to bash install
    Running a script straight from the internet, which skips any review unless you read it first.

    Steps

    1. Read the README sections on what is removed and how revert works.
    2. Read the HN thread for the safety concerns.
    3. Look at your own Mac's storage view to see how much space system data uses.
    4. Read the install script itself before running anything.
    5. Self check: what would you check in a script before giving it admin rights?
    Try it20 min · 5 steps
    You need: A Mac on macOS 27 with Apple silicon that you own, ideally after a backup. No API key needed.

    Steps

    1. Read the install script in the repo before installing.
    2. Install with Homebrew using brew install omlahore/tap/removemacai.
    3. Run removemacai status to see the current state.
    4. Run removemacai off --dry-run and read what it would change.
    5. If you go ahead, measure free disk space before and after, and know that removemacai revert restores everything.

    Small angles to try

    • Use the keep option to keep one feature and compare space saved.
    • Check which processes stop using memory afterwards.
    • Write a short review of the script's safety for others.
  15. #15Research

    OneStreamer: a 4B model that remembers earlier video after frames scroll away

    Researchers at Nanjing University released OneStreamer, a 4 billion parameter model for live video that keeps a layered memory of captions, so evidence survives after frames leave the context window. They report the best results on all 8 streaming benchmarks they tested and release a training set of about 1 million samples. The model is on Hugging Face as MCG-NJU/OneStreamer-4B.

    Heat · Hugging Face daily papers, October 2, about 160 to 220 upvotes
    What people foundThe idea is simple and reusable: instead of keeping raw frames, compress the past into text summaries at several levels and query those. The claims are the authors' own and have not been reproduced independently yet.
    Learn it15 min · 5 steps

    Key ideas

    Streaming video model
    A model that answers questions while video keeps arriving, instead of after the whole clip.
    Hierarchical memory
    Summaries at several levels, such as per minute and per scene, that replace raw frames.
    Context window
    The amount of input a model can attend to at once.

    Steps

    1. Read the abstract and method figure on the Hugging Face paper page.
    2. Look at the model card for MCG-NJU/OneStreamer-4B to see how it is meant to be used.
    3. Connect it to log rotation: keep summaries of old logs instead of every line.
    4. Read which benchmarks it was tested on and what they measure.
    5. Self check: what kind of question would a caption memory fail to answer?
    Try it40 min · 5 steps
    You need: Python 3, a GPU with enough memory for a 4B vision model, and your own short video. No API key needed.

    Steps

    1. Open the model card and follow its usage instructions to load the model.
    2. Record a 5 minute video of your desk where an object moves early on.
    3. Ask the model at the end where the object was at the start.
    4. Compare with a plain image model given only the last frames.
    5. Note memory use and response time.

    Small angles to try

    • Try a longer 20 minute video.
    • Write your own caption memory with any vision model and compare.
    • Ask questions about text that appears briefly on screen.
  16. #16AI in the world

    The AI buildout needs about 9% of US GDP in AI revenue, says a Brookings paper

    Columbia economist Stijn Van Nieuwerburgh's Brookings paper estimates about $10.3 trillion of US AI investment from 2025 to 2032. New this week is the attention on the other side of the ledger: to earn a 10% return, AI services would need about $3.7 trillion a year in revenue by 2032, around 9.2% of projected GDP. An X post comparing that share to what Americans spend on food spread widely.

    Heat · X: 1,592 likes and 135K views on the 9% of GDP post
    What people foundVan Nieuwerburgh compares the opaque financing behind data centers, special purpose vehicles, private credit and leases, to patterns seen before the 2008 mortgage crisis, and says better measurement and transparency are the first policy step. The food comparison is the X poster's framing, not the paper's.
    Learn it15 min · 5 steps

    Key ideas

    Required revenue
    The income an investment must earn to repay its cost plus a target return.
    Special purpose vehicle
    A separate company created to hold one project's assets and debt.
    Share of GDP
    A way to compare spending with the size of the whole economy.

    Steps

    1. Read the Brookings summary page for the main numbers.
    2. Read the roic.ai write up for the 9% revenue figure.
    3. Compare 9% of GDP with categories you know, such as health care or food.
    4. Think about which AI products would need to grow to reach that revenue.
    5. Self check: what growth rate per year turns about $100B today into $3.7T by 2032?
    Try it20 min · 5 steps
    You need: A spreadsheet or Python. No API key needed.

    Steps

    1. Put $100B of yearly AI revenue in 2025 in a sheet.
    2. Compute the yearly growth needed to reach $3.7T by 2032.
    3. Try three scenarios with lower targets and see how the needed growth changes.
    4. Compare with the fastest growing companies you know.
    5. Write one sentence on which assumption you find least realistic.

    Small angles to try

    • Add a cost line for depreciation of chips over 5 years.
    • Model what happens if hardware costs fall 30% a year.
    • Compare with the dot com era telecom buildout.

Get the Radar by email

One email each evening, Japan time, with that day's items. Opening soon: leave your email and we'll send a confirm link when it starts.