Code415 Radar

by Ziyu Guo · @Code415zg

What people are paying attention to in AI, LLMs and CS today. Every item comes with a short path to learn it and a small test to try it yourself. Updated every evening, Japan time.

Sunday, October 4, 2026

18 items · Kolibri-1 opens 78B MoE weights; Google Sites phishing fools password managers
  1. #1Models & launches

    Kolibri-1: Aleph Alpha opens a 78B MoE with 3.46B active under Apache 2.0

    German lab Aleph Alpha released Kolibri-1, a mixture of experts model with 78B total and 3.46B active parameters per token, for German and English. The card lists a 1,048,576 token context window but recommends staying at or under 262,144 tokens for efficiency, and the license is Apache 2.0.

    Heat · HN front page #11, 547 points, 308 comments · X: Aleph Alpha launch post 4K likes, 950K views · HF: about 300 likes in two days
    What people foundThe model card asks for serious hardware: about 78 GB of FP8 weights, so two A100 80 GB cards or one H200 at minimum. A Trending Topics headline argued it is no match for the open weight leaders, so the interesting question is quality per active parameter, not raw rank.
    Learn it15 min · 5 steps

    Key ideas

    Mixture of experts
    Each token is sent to a few small expert networks instead of the whole model, so compute follows active parameters.
    Active parameters
    The weights actually used for one token, here about 3.46B out of 78B.
    Sovereign model
    A model built and hosted in one region so its users do not depend on foreign providers.

    Steps

    1. Open the Kolibri-1 model card on Hugging Face and read the architecture and hardware sections.
    2. Note why 3.46B active parameters make it cheap to run per token, while all 78B still have to sit in memory.
    3. Compare the recommended 262,144 token context with the 1M maximum and look at the long context benchmark it reports.
    4. Connect it to something familiar: it is like a large library where each question only opens a few books.
    5. Self check: why does an MoE with few active parameters still need a big GPU?
    Try it45 min · 5 steps
    You need: A cloud GPU with at least 80 GB of memory per card as the card describes, Python, vLLM. No API key needed for the weights.

    Steps

    1. Rent a machine that meets the card's minimum, for example one H200, and install vLLM.
    2. Install the helper package from the card with pip install 'aleph-alpha-inference>=1'.
    3. Serve it with vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 --reasoning-parser kolibri1 --tool-call-parser kolibri1 --enable-auto-tool-choice.
    4. Send the same 10 coding and 10 German prompts through the OpenAI compatible endpoint vLLM exposes and record tokens per second.
    5. Run the same prompts on a dense model of similar active size and compare answers and speed.

    Small angles to try

    • Try one of the quantized versions listed in the model tree and measure the quality drop.
    • Feed a 200K token document and ask questions about its middle.
    • Score German answers with a native speaker against an English baseline.
  2. #2Trends × tech

    Google Sites phishing trends in Japan: why password managers fill it in

    A phishing wave sends fake Gmail security alerts that say an app password was added to your account. The link opens a page on sites.google.com with a fake CAPTCHA and a copy of the Google sign in form, so the address bar really does say google.com. A Japanese warning about it spread fast on X this weekend.

    Heat · X Japan: 40K likes, 28K reposts, 21K bookmarks on one warning post
    What people foundThe write up by Ministry of Cyber Affairs says the trick works because password managers trust the whole google.com site, so autofill will put real Google credentials into an attacker page. Real Google sign in only happens on accounts.google.com, and the fake mails came through a third party relay instead of Google's own bounce domains.
    Learn it15 min · 5 steps

    Key ideas

    Registrable domain
    The part of a host name you buy, like google.com, which many different services can share.
    Origin binding
    Passkeys are tied to one exact site, so a look alike page on another host cannot use them.
    Mail headers
    Lines like mailed by and the DKIM signature show which server really sent a message.

    Steps

    1. Read the Ministry of Cyber Affairs threat analysis and write down each step of the attack flow.
    2. Look at the difference between sites.google.com and accounts.google.com and ask which one your password manager treats as the same login.
    3. Read how passkeys are bound to one origin and why that stops this attack even when a human is fooled.
    4. Open a real Google security alert in your own Gmail, choose show original, and find the sending domain.
    5. Self check: if a page is on google.com with valid HTTPS, what else must be true before you type a password?
    Try it30 min · 6 steps
    You need: A browser with your usual password manager, Python 3. No API key needed. Use only your own accounts and messages.

    Steps

    1. Open your password manager settings and find how it matches saved logins to sites: whole domain, host, or exact URL.
    2. Check whether your Google entry would be offered on any google.com host and switch it to exact host matching if the manager allows it.
    3. Export the raw source of one real Google alert from your own mailbox as an .eml file.
    4. Write a short Python script with the standard email module that prints the From, Return-Path and DKIM d= values.
    5. Add a rule to the script that warns when the DKIM domain is not a google.com domain, then test it on your saved mail.
    6. Turn on a passkey for your Google account if you have not already.

    Small angles to try

    • Compare how two password managers handle the same subdomain question.
    • Use the Public Suffix List with a Python library to show where registrable domains begin.
    • Share a one page checklist in Japanese and English for family members.
  3. #3Engineering & CS

    Simon Willison: cloud and API services need hard budget caps by default

    Simon Willison argues that every pay per use service should ship with a hard spending cap that pauses the service, and that users should have to tick a box to remove it. His reason is that coding agents make it easy to launch apps that can run up large bills overnight.

    Heat · HN front page #4, 305 points, 156 comments
    What people foundWillison says warning emails will not cut it because nobody reads a midnight alert in time. He notes AWS began rolling out spending limits that pause projects in September 2026 and Google Cloud added Spend Caps in July 2026.
    Learn it15 min · 5 steps

    Key ideas

    Soft cap
    A limit that only sends an alert when spending passes a number.
    Hard cap
    A limit that stops the service when spending passes a number.
    Runaway loop
    A bug or agent that keeps calling a paid API much faster than a human would.

    Steps

    1. Read Willison's post and list the two kinds of caps he compares.
    2. Open the billing page of one cloud or API account you use and find what limits it offers today.
    3. Think of a cron job or agent you run and estimate its worst case cost per hour if it loops.
    4. Compare this with a prepaid phone plan, which simply stops when credit runs out.
    5. Self check: what would break for your users if your hard cap triggered at 3 AM?
    Try it30 min · 5 steps
    You need: An account on any paid API or cloud you already use, Python 3. Uses your own API key; keep amounts tiny.

    Steps

    1. Write down the spending limits your provider exposes and set the lowest hard limit it allows on a test project.
    2. Write a small Python wrapper around your API client that counts tokens or requests and keeps a running cost estimate.
    3. Make the wrapper raise an error once the estimate passes a budget you set, for example one dollar.
    4. Run a loop of cheap calls and confirm the wrapper stops it before the provider bill moves much.
    5. Check the provider dashboard the next day and compare its number with your estimate.

    Small angles to try

    • Add the same guard to a coding agent's tool calls.
    • Store the counter in SQLite so it survives restarts.
    • Compare how three providers label soft and hard limits.
  4. #4AI in the world

    Apple tightens macOS Full Disk Access because of AI agents

    Apple says apps will need very explicit user action before they get Full Disk Access, the macOS permission that opens files, Messages, Mail and browsing history. The move follows a report that Meta's Muse desktop agent read a journalist's messages, which Meta disputes, and a Wired report on a flaw in ChatGPT's Mac app.

    Heat · TechCrunch AI section, Oct 2 · Muse privacy row was already a big story on X this week
    What people foundApple's own line is that the risk from this level of access grows as agents become more capable and autonomous. No macOS version or date was given, so for now this is a direction, not a shipped setting.
    Learn it15 min · 5 steps

    Key ideas

    Full Disk Access
    A macOS privacy permission that lets an app read data normally protected by the system.
    TCC
    Transparency, Consent and Control, the macOS framework that asks you before apps reach private data.
    Least privilege
    Give each program only the access it needs for the job in front of it.

    Steps

    1. Read the TechCrunch report and note the two incidents Apple points to.
    2. Open System Settings, Privacy and Security, Full Disk Access on a Mac and look at which apps are listed.
    3. Read Apple's platform security guide section on privacy controls to see how TCC prompts work.
    4. Connect it to phone permissions: an app asking for your photos is the same idea.
    5. Self check: why is Full Disk Access riskier for an AI agent than for a backup app?
    Try it20 min · 5 steps
    You need: A Mac you own. No API key needed.

    Steps

    1. List every app in Full Disk Access and write down why each one needs it.
    2. Remove access for any app you cannot justify, starting with AI assistants and terminals you rarely use.
    3. For a coding agent you use, check whether it really needs Full Disk Access or only access to one project folder.
    4. Run the agent on a test project after removing access and note what, if anything, breaks.
    5. Keep a short note of the final list so you can compare after the next macOS update.

    Small angles to try

    • Repeat the audit for Accessibility and Screen Recording permissions.
    • Run the agent inside a separate macOS user account and compare.
    • Compare with how Windows or Linux desktops limit the same access.
  5. #5Tools & open source

    Cloudflare open sources a security audit skill for coding agents

    Cloudflare published security-audit-skill, an MIT licensed skill that walks a coding agent through a six phase audit of your code. Separate agents hunt for bugs, fresh verifier agents try to disprove each finding, and the result is a JSON list plus a markdown report.

    Heat · GitHub Trending: top gainer, about 3,600 stars today, 10.7K total
    What people foundThe design choice worth copying is the verifier step: a new agent with no stake in the finding tries to prove it wrong before it reaches the report. That targets the false positive flood maintainers have been complaining about.
    Learn it15 min · 5 steps

    Key ideas

    Agent skill
    A folder of instructions and scripts a coding agent loads when a task matches.
    Adversarial verification
    A second checker tries to refute a result instead of agreeing with it.
    Coverage led hunting
    Splitting a codebase into areas so every part gets looked at once.

    Steps

    1. Read the README of cloudflare/security-audit-skill and list the six phases.
    2. Open the hunting guides folder and skim the one closest to your stack, for example web protocols.
    3. Note how findings are recorded as JSON and what fields a verifier must check.
    4. Compare it with code review where a second reviewer must approve.
    5. Self check: why does a fresh verifier reduce false positives?
    Try it45 min · 5 steps
    You need: A coding agent that supports skills, Node.js for npx, and a repository you own. Agent API costs apply.

    Steps

    1. Pick a small project of your own, or an intentionally vulnerable training app running in a local container.
    2. Install the skill with npx skills add https://github.com/cloudflare/security-audit-skill --skill security-audit inside that project.
    3. Ask your agent to run the security audit and let it finish all phases.
    4. Read the markdown report and check each finding by hand against the code.
    5. Count true findings, false positives and anything the verifier rejected.

    Small angles to try

    • Run it twice and see how stable the findings are.
    • Compare with a classic static analyzer on the same code.
    • Try a second agent or model and compare cost per true finding.
  6. #6Tools & open source

    AgentCraft: a team of Claude agents builds your repo inside Minecraft

    AgentCraft turns a coding run into a Minecraft studio: a lead agent on Opus plans and up to three Sonnet workers code in separate git worktrees. The agents walk over to your character when they need a decision, and merges need human approval.

    Heat · X: launch post 10.7K likes, 840K views, the most liked AI post in today's scan
    What people foundUnder the game skin it is a sensible pattern: one planner, a few workers, one worktree each, and a human gate before merge. The repo is still tiny, MIT licensed and Windows only, so treat it as a demo of the UX, not a tool to rely on.
    Learn it15 min · 5 steps

    Key ideas

    Git worktree
    A second checkout of the same repository in another folder, so two agents can edit without clashing.
    Orchestrator
    The program that hands tasks to agents and collects their results.
    Human in the loop
    A person must approve key steps, here every merge.

    Steps

    1. Watch the launch clip on X to see the flow from plan to merge.
    2. Read the AgentCraft README, especially the Foreman orchestrator and the worktree section.
    3. Run git worktree --help locally to see how worktrees are created and removed.
    4. Compare it with a team lead assigning tickets to three developers.
    5. Self check: what problem would appear if all agents shared one working folder?
    Try it45 min · 5 steps
    You need: Windows, Java 25, Node 22 or newer, Minecraft Java Edition, a Claude API key, a throwaway repo.

    Steps

    1. Clone it with git clone https://github.com/blendi-remade/agentcraft and read the setup notes.
    2. Point the launch script at a small practice repository, not real work.
    3. Give the lead agent one task that splits cleanly into three parts.
    4. Watch which worker takes which part and approve or reject each merge.
    5. Record total time, cost and how many merges you rejected.

    Small angles to try

    • Do the same task with a single agent and compare time and cost.
    • Try a task that does not split well and see where coordination fails.
    • Without Minecraft, reproduce the pattern with plain worktrees and a script.
  7. #7Trends × tech

    setlog is #1 free in Japan: hourly 2 second clips, and the push problem

    setlog, a small group video diary app, is back at #1 on Japan's App Store free chart, ahead of Bump at #2 and ChatGPT at #3. Every hour on the hour it asks you to record a 2 second clip with no editing, then stitches the day's clips into a vlog for a group of up to 12 friends.

    Heat · Japan App Store Top Free #1 in Apple's feed updated Oct 3 · was also #1 for about a month in May per Impress Watch
    What people foundThe engineering hook is that every user gets the prompt at the same minute. Sending millions of pushes at :00 and then absorbing a wave of uploads is a classic thundering herd problem, and the stitching of short clips into one video happens every night.
    Learn it15 min · 5 steps

    Key ideas

    Push fan out
    Sending one notification event to many devices through Apple and Google push services.
    Thundering herd
    Many clients hit a server at the same moment because they were all triggered together.
    Video concatenation
    Joining short clips into one file, ideally without encoding them again.

    Steps

    1. Read the Impress Watch article on setlog to understand the rules of the app.
    2. Open Apple's top free feed for Japan and see where setlog, Bump and ChatGPT sit today.
    3. Read about jitter: spreading timed jobs over a few seconds to smooth load.
    4. Compare it with BeReal, which sent one daily prompt to everyone at once.
    5. Self check: how would you keep the hourly feeling while spreading the upload spike?
    Try it30 min · 5 steps
    You need: Python 3 and ffmpeg installed locally. No API key needed.

    Steps

    1. Record or create six 2 second clips on your own phone and copy them to your computer.
    2. Use ffmpeg's concat feature to join them into one video and time how long it takes.
    3. Try again with re encoding forced and compare time and file size.
    4. Write a small Python simulation of 1 million users uploading at :00 versus spread over 60 seconds and plot uploads per second.
    5. Note the peak rate in both cases.

    Small angles to try

    • Add a random delay of up to 10 seconds per user and see how far the peak drops.
    • Estimate daily storage for 1 million users at your clip bitrate.
    • Compare with Bump's location sharing load, which is steady instead of spiky.
  8. #8AI in the world

    OpenAI safety lead David Robinson quits, says the culture is broken

    David Robinson, who led safety reports for major OpenAI launches during three and a half years at the company, resigned and published an essay in The Atlantic. He argues that learning by trial and error guarantees failures once models are powerful, and that change has to come from outside pressure.

    Heat · HN #22, 115 points but 364 comments, an unusually heated thread
    What people foundOpenAI's spokesperson said the company is making sure its models do not become more capable than it can safely manage. Robinson points to recent incidents such as a Hugging Face breach and agents going off script as signs the current approach is not enough.
    Learn it15 min · 5 steps

    Key ideas

    System card
    A public report on how a model was tested for risks before launch.
    Iterative deployment
    Releasing models step by step and fixing problems found in real use.
    External oversight
    Checks by regulators, auditors or outside researchers instead of only the company itself.

    Steps

    1. Read the TechCrunch report on the resignation.
    2. Open one recent OpenAI system card and look at what safety tests are described.
    3. List the arguments for iterative deployment and the arguments Robinson makes against it.
    4. Compare it with how aviation uses outside investigators after incidents.
    5. Self check: what kind of failure can trial and error not fix after the fact?
    Try it20 min · 4 steps
    You need: A browser. No API key needed.

    Steps

    1. Pick two recent system cards from two labs.
    2. Make a table of which risk areas each one tests.
    3. Mark which tests were run by outside groups and which in house.
    4. Write one paragraph on what is missing from both.

    Small angles to try

    • Add a third lab's card to the table.
    • Track how the length of system cards changed from 2024 to 2026.
    • Compare with the incidents Robinson cites.
  9. #9Engineering & CS

    Agents need documentation, not memory: fix the environment, not the agent

    Kevin Liao argues coding agents work better with current, structured project docs than with memory plugins that pull up old chat fragments by similarity. On X, Lauren Tan made a related point: if you keep correcting an agent for the same mistake, fix the environment that causes it, and her pstack tool added a /correct skill for that.

    Heat · HN front page #10, 98 points · X: @poteto post 2.8K likes, 100K views
    What people foundLiao's workflow is consult, build, update: the agent reads an internal docs folder before work and updates it after. His case against memory is that it is unauditable and treats stale facts as current, while plain markdown can be reviewed and versioned.
    Learn it15 min · 5 steps

    Key ideas

    AGENTS.md
    A markdown file at the root of a repo that tells coding agents how the project works.
    Retrieval memory
    Storing past conversations as embeddings and pulling back the most similar pieces.
    Stale context
    Information that was true once but is now wrong and still gets used.

    Steps

    1. Read Liao's post and write down his three complaints about memory plugins.
    2. Read Lauren Tan's post on X about correcting the environment.
    3. Look at an AGENTS.md or CLAUDE.md in a repo you use and note what is missing.
    4. Connect it to onboarding a new colleague: you give them a wiki, not your chat logs.
    5. Self check: when is retrieval memory still the better tool?
    Try it40 min · 5 steps
    You need: Any coding agent and a repo you own. Agent API costs apply.

    Steps

    1. Pick one mistake your agent repeats in your repo.
    2. Run a task that triggers it and record what happens.
    3. Write the rule into a short docs file the agent reads first, such as AGENTS.md, and add a pointer to a design notes folder.
    4. Run the same task in a fresh session and see if the mistake is gone.
    5. Ask the agent to update the notes at the end and review the diff.

    Small angles to try

    • Compare with a memory plugin on the same mistake.
    • Measure how many tokens the docs file adds per session.
    • Try it with two different agents.
  10. #10Tools & open source

    colibri: stream MoE experts from disk to run giant models on a desktop

    colibri is a C inference engine with no dependencies that treats SSD, RAM and VRAM as one memory hierarchy. Shared layers stay in RAM, routed experts are loaded from disk when the router picks them, and a cache keeps the hot ones close.

    Heat · GitHub Trending: about 870 stars today, 35.7K total
    What people foundThe README is honest about the cost: about 1.8 tokens per second on a 128 GB desktop with no GPU, and 0.05 to 0.1 on a small box. It is a lesson in memory hierarchy more than a daily driver, and small MoE models like OLMoE make it testable on a laptop.
    Learn it15 min · 4 steps

    Key ideas

    Expert streaming
    Loading only the experts a token needs from disk instead of holding all weights in memory.
    LRU cache
    Keep the most recently used items and drop the oldest when space runs out.
    int4 quantization
    Storing each weight in 4 bits to cut memory by about four times versus 16 bits.

    Steps

    1. Read the colibri README section on how streaming works.
    2. Look at the requirements table and compare disk, RAM and GPU needs per model.
    3. Recall how an operating system pages memory to disk; this is the same idea for weights.
    4. Self check: why does SSD read speed matter more than CPU speed here?
    Try it45 min · 5 steps
    You need: Linux or macOS, about 10 GB free disk, 8 GB RAM. No API key needed.

    Steps

    1. Follow the README to download the release or clone the repo and run its setup script.
    2. Convert the small OLMoE model as the README describes.
    3. Start a chat with that model and record tokens per second.
    4. Watch disk reads with your system monitor while it generates.
    5. Run again and see whether the cache makes the second run faster.

    Small angles to try

    • Move the model to a slower USB drive and compare speed.
    • Limit available RAM and watch the cache hit rate change.
    • Compare with llama.cpp on the same model.
  11. #11Engineering & CS

    FTL: a new Rust OS for clouds where the OS is a library

    Seiya Nuta's FTL runs each container's operating system as a userspace library on top of a minimal kernel that isolates containers with hardware user mode. It aims to run existing Linux binaries, and its own website is served by a Rust HTTP server running on FTL.

    Heat · HN front page #17, 158 points, 63 comments
    What people foundThe pitch is microkernel flexibility with monolithic kernel speed: processes, file systems and TCP/IP live in userspace, so fixing or upgrading them is like updating an app. It is at v0.1.0, so the claims are early.
    Learn it15 min · 4 steps

    Key ideas

    Library OS
    OS services linked into the application itself instead of running in a shared kernel.
    Microkernel
    A tiny kernel that leaves drivers and services to user programs.
    Linux ABI compatibility
    Running unchanged Linux programs by answering their system calls the same way.

    Steps

    1. Read the FTL website and its introductory blog post.
    2. Skim the nuta/ftl repository layout to see what lives in the kernel and what lives in userspace.
    3. Compare it with gVisor or unikernels you may have heard of.
    4. Self check: what does a library OS gain and lose compared with a container on Linux?
    Try it30 min · 4 steps
    You need: A Linux machine with the build tools the FTL README lists, such as a Rust toolchain and an emulator. No API key needed.

    Steps

    1. Clone https://github.com/nuta/ftl and read the README build section.
    2. Build and boot it exactly as the README describes.
    3. Request the sample HTTP server and note the response time.
    4. List which Linux features the README says are not supported yet.

    Small angles to try

    • Compare boot time with a small Linux VM.
    • Count lines of code in kernel versus userspace parts.
    • Try running a tiny static Linux binary if the README supports it.
  12. #12Trends × tech

    凱旋門賞 tonight: two Japanese horses, and the GPS data behind each stride

    The Prix de l'Arc de Triomphe runs at ParisLongchamp at 23:05 JST tonight over 2,400 m with 16 runners, including Japan's Meisho Tabaru and Admire Terra. JRA sells bets on the race online, and the Arc is trending on X Japan alongside other weekend races.

    Heat · X Japan trending #6 · JRA online betting open since 07:00 JST
    What people foundFrance Galop has published sectional timing for the Arc since 2019, tracking every runner through each section of the race. Comparing a Japanese horse's late sections with the winner's tells you more than the finishing order.
    Learn it15 min · 5 steps

    Key ideas

    Sectional timing
    The time each horse takes for each part of the course, measured by tracking.
    Pari mutuel odds
    Odds set by how much money is in the pool for each horse, as JRA uses.
    Implied probability
    The chance of winning that a given set of odds suggests.

    Steps

    1. Read the JRA page on the 2026 Arc runners.
    2. Read France Galop's note on sectional timing to see what is measured.
    3. Learn how to turn fractional odds like 7 to 4 into implied probability: divide 4 by 7 plus 4.
    4. Connect it to split times in a marathon.
    5. Self check: why can a horse with slower overall time have the fastest final section?
    Try it30 min · 4 steps
    You need: Python 3 with pandas and matplotlib. No API key needed.

    Steps

    1. After the race, open the official results and sectional data from France Galop.
    2. Copy the sections for the winner and the two Japanese horses into a CSV.
    3. Plot speed per section for the three horses.
    4. Mark where each horse gained or lost ground and write one sentence per horse.

    Small angles to try

    • Add the 2025 winner's sections for comparison.
    • Convert the final odds to implied probabilities and see how much they sum above 100%.
    • Compare the ground conditions of this year and last year.
  13. #13AI in the world

    Capcom REX: upgrading RE Engine so AI can read and help build games

    At Capcom Open Conference RE:2026 in Tokyo, Capcom described REX, a staged upgrade of its RE Engine rather than a new engine. Parts include RE:Flows for visual gameplay building with code generation, and the codebase is being restructured to follow standard language rules so AI tools can read it.

    Heat · X: 2.4K likes, 132K views on a post about it
    What people foundNotebookcheck's report is more careful than the viral posts: AI writing code and finding bugs is a stated future goal with no timeline. The concrete step is making a large in house codebase readable by AI, which many companies face too.
    Learn it15 min · 4 steps

    Key ideas

    Game engine
    The shared software a studio uses for rendering, physics, tools and scripting.
    Visual scripting
    Building game logic by connecting blocks instead of writing code.
    AI readable code
    Code that follows standard conventions so models trained on public code understand it.

    Steps

    1. Read the Notebookcheck article and list the five REX parts.
    2. Note which parts are released and which are plans.
    3. Think about your own work codebase: what custom rules would confuse an AI assistant?
    4. Self check: why would non standard C++ conventions hurt AI code help?
    Try it30 min · 4 steps
    You need: A coding agent and an open source project that uses unusual conventions. Agent API costs apply.

    Steps

    1. Pick a small open source project with heavy custom macros or naming rules.
    2. Ask your agent to explain one module and note mistakes.
    3. Write a short conventions file for the agent and ask again.
    4. Compare the two answers and count fixed mistakes.

    Small angles to try

    • Try the same with a project that follows standard style.
    • Compare two agents.
    • Measure how long the conventions file needs to be before results stop improving.
  14. #14Engineering & CS

    Eric S. Raymond after 43 years of C: probably never writing it again

    Eric S. Raymond, who has written C since 1983, posted that he will probably never write another C program, a big shift after 43 years with the language. The post landed alongside other viral takes on how AI changes the developer's job.

    Heat · X: 2.8K likes, 94K views · a related take that AI turned developers into QA testers has 5.2K likes
    What people foundThis is Raymond's claim about his own work, not a measured result. The useful debate is what replaces hand written C: a safer language, generated code you review, or both, and who carries the review load when the human becomes the tester as Steve Huynh put it.
    Learn it15 min · 4 steps

    Key ideas

    Memory safety
    A language guarantee that a program cannot read or write memory it does not own.
    Code review load
    The time humans spend checking code, which grows when code is generated faster.
    Translation to Rust
    Converting C code to Rust, now often attempted with AI help.

    Steps

    1. Read Raymond's post and Steve Huynh's post on X.
    2. Read the C to Rust section of any recent translation project you trust, such as the DARPA TRACTOR summary.
    3. Think of the last C or C++ bug you fixed and whether a safe language would have prevented it.
    4. Self check: what does a reviewer need to see to trust generated low level code?
    Try it40 min · 4 steps
    You need: A C compiler, a Rust toolchain, and any coding agent. Agent API costs apply.

    Steps

    1. Take a small C program you wrote, about 200 lines.
    2. Ask an agent to translate it to Rust and to explain each unsafe block.
    3. Compile both and run the same test inputs.
    4. Note how long review took compared with writing it yourself.

    Small angles to try

    • Translate to Go or Zig instead.
    • Run both under a sanitizer and compare.
    • Ask the agent to write tests first, then translate.
  15. #15Tools & open source

    Tencent BrowserSkill lets agents drive your logged in Chrome in its own window

    Tencent's BrowserSkill is a CLI, background daemon and browser extension that lets agents such as Claude Code or Cursor use your logged in Chrome or Edge in a separate agent window. It can capture network and console traffic and full page screenshots, and the agent can run on a server while the browser stays on your machine.

    Heat · GitHub Trending: about 1,300 stars today, 4.2K total
    What people foundGiving an agent your real sessions is powerful and risky in the same week Apple is tightening access for agents. A separate browser profile with only test accounts is the sane way to try it.
    Learn it15 min · 4 steps

    Key ideas

    Browser automation
    A program clicking and typing in a browser for you, like Playwright.
    Session cookies
    Small tokens that keep you signed in, which an agent can use if it drives your browser.
    Daemon
    A background process that waits for commands.

    Steps

    1. Read the BrowserSkill README and draw the path from agent to daemon to extension to browser.
    2. Note what the agent window can see and what it cannot.
    3. Compare it with Playwright, which usually starts a clean browser.
    4. Self check: which of your logged in sites would you never let an agent use?
    Try it40 min · 5 steps
    You need: Chrome or Edge with a new empty profile, a coding agent. No API key needed for the tool itself.

    Steps

    1. Create a fresh browser profile with no saved logins.
    2. Read the install script from the README before running it, then install.
    3. Check the install with bsk --version.
    4. Ask your agent to open a public docs site, take a screenshot and list console errors.
    5. Remove the extension when done.

    Small angles to try

    • Compare with Playwright on the same task.
    • Log only network requests and check what the agent fetched.
    • Try a site you built yourself with a test account.
  16. #16Research

    Learning forever: Sutton's student publishes his thesis on lost plasticity

    Richard Sutton shared the PhD thesis of Shibhansh Dohare, Learning Forever using Artificial Neural Networks. It builds on their 2024 Nature paper showing that standard deep learning slowly loses the ability to learn when trained on a long stream of new tasks.

    Heat · X: Sutton's post 1.1K likes, 40K views
    What people foundThe fix from the Nature paper is continual backpropagation: track how useful each unit is and reset the least useful ones to fresh random weights. It matters for any model that is meant to keep learning after deployment.
    Learn it15 min · 4 steps

    Key ideas

    Loss of plasticity
    A network trained for a long time on changing tasks stops being able to learn new ones well.
    Dormant units
    Neurons that almost never activate and so stop contributing.
    Continual backpropagation
    Normal training plus periodic resets of the least useful units.

    Steps

    1. Read the abstract and figures of the Nature paper by Dohare and colleagues.
    2. Find the plot where accuracy falls as tasks continue, and see how continual backprop changes it.
    3. Compare it with catastrophic forgetting, which is about losing old skills instead of new ones.
    4. Self check: why does resetting a few units help learning without wiping knowledge?
    Try it45 min · 4 steps
    You need: Python with PyTorch, a CPU is enough. No API key needed.

    Steps

    1. Build a small MLP and a stream of tasks, for example MNIST with a new random label permutation every few epochs.
    2. Train through 50 tasks and log accuracy on each new task.
    3. Count units whose activations stay near zero.
    4. Add a reset of the least used 1% of units every task and repeat.

    Small angles to try

    • Swap ReLU for another activation and compare.
    • Add weight decay and see if it slows the decline.
    • Plot the number of dormant units over time.
  17. #17Research

    Raschka's reasoning from scratch: RLVR and GRPO explained with code

    Sebastian Raschka released the sixth session of his reasoning from scratch series, covering reinforcement learning with verifiable rewards and Group Relative Policy Optimization. It pairs with chapters 6 and 7 of his book repository, which include GRPO training scripts.

    Heat · X: 1.2K likes, 51K views · repo about 5K stars
    What people foundGRPO drops the separate value model: it samples a group of answers to the same question, scores them with a checker, and pushes up the ones above the group average. Verifiable rewards like a correct math answer make that cheap.
    Learn it15 min · 5 steps

    Key ideas

    Verifiable reward
    A reward computed by a program, like checking whether a math answer is right.
    Group relative advantage
    Each answer is scored against the average of answers to the same prompt.
    Policy
    The model's way of choosing tokens, which RL training adjusts.

    Steps

    1. Watch the start of Raschka's video for the overview.
    2. Open chapter 6 in the rasbt/reasoning-from-scratch repository.
    3. Find where the group advantage is computed and read it line by line.
    4. Compare it with grading on a curve in a class.
    5. Self check: why does GRPO not need a value model?
    Try it45 min · 4 steps
    You need: Python, PyTorch, a GPU helps but small runs work on CPU. No API key needed.

    Steps

    1. Clone with git clone --depth 1 https://github.com/rasbt/reasoning-from-scratch.git.
    2. Follow the chapter 2 setup notes to install dependencies.
    3. Run the chapter 6 GRPO example on a few steps.
    4. Log reward per step and look at sample answers before and after.

    Small angles to try

    • Change the group size and see the effect on reward noise.
    • Add a format reward and see if it helps or hurts accuracy.
    • Compare with the chapter 7 improved version.
  18. #18Trends × tech

    #f1jp: the Bahrain GP runs at Sepang, and how strategy is simulated

    This weekend's Round 16 is the Bahrain Grand Prix moved to Sepang International Circuit in Malaysia, 56 laps starting 16:00 JST today, with Singapore next weekend. Hot and humid Sepang makes tyre wear and pit timing the story for Japanese fans following #f1jp.

    Heat · X Japan trending #5
    What people foundTeams decide one stop versus two stop with lap time models: how much slower each lap gets as tyres wear, plus the time lost in the pit lane. The open source FastF1 Python library lets fans pull lap times after a race and check those choices.
    Learn it15 min · 4 steps

    Key ideas

    Tyre degradation
    Lap times get slower as tyres wear, often by a roughly steady amount per lap.
    Pit loss
    The total time a car loses by driving through the pit lane and stopping.
    Undercut
    Pitting earlier than a rival to gain time on fresh tyres.

    Steps

    1. Read the 2026 F1 calendar page for this weekend's race.
    2. Read the FastF1 documentation front page to see what data it can load.
    3. Write down a simple model: lap time equals base time plus wear rate times tyre age.
    4. Self check: when does a second pit stop pay off?
    Try it40 min · 4 steps
    You need: Python 3 with pandas and matplotlib, FastF1 installed as its docs describe. No API key needed.

    Steps

    1. Write a function that returns total race time for a list of stint lengths, a wear rate and a pit loss.
    2. Compare one stop and two stop plans for 56 laps with a few wear rates.
    3. After the race, load real lap times with FastF1 following its getting started guide.
    4. Fit the wear rate for one driver's stint and plug it back into your model.

    Small angles to try

    • Add a safety car lap that cuts pit loss and see which plan wins.
    • Compare two teams' wear rates.
    • Run the same model for Singapore next week.

Get the Radar by email

One email each evening, Japan time, with that day's items. Opening soon: leave your email and we'll send a confirm link when it starts.