Code415 Radar

by Ziyu Guo · @Code415zg

What people are paying attention to in AI, LLMs and CS today. Every item comes with a short path to learn it and a small test to try it yourself. Updated every evening, Japan time.

Saturday, October 3, 2026

18 items · Pope's AI art post goes viral; decision models come to llama.cpp
  1. #1Trends × tech

    Pope Leo XIV on AI art goes viral: what content provenance can prove

    Pope Leo XIV posted on X that it is becoming urgent to tell human art apart from what machines produce, calling the gap ontological, not just aesthetic, and asked for an alliance with artists and cultural institutions. The post became the most liked general post in today's scan. The main technology people point to is C2PA Content Credentials, a signed record attached to a file that says which tool made or edited it.

    Heat · X: about 347K likes, 16.9M views, 30K bookmarks · covered by TechCrunch and Axios
    What people foundC2PA describes Content Credentials as a nutrition label for digital content, but a label only proves where a file came from and how it was edited. A missing label does not mean a picture was made by AI, because screenshots and many uploads strip it.
    Learn it15 min · 5 steps

    Key ideas

    Content Credentials
    A signed manifest stored with an image or video that lists the tools and edits behind it, built on the C2PA specification, now at version 2.3.
    manifest stripping
    Most metadata, including a C2PA manifest, is lost when you take a screenshot or a platform recompresses the file.
    provenance vs detection
    Provenance records a file's history from the maker's side, while AI detectors guess from the pixels and can be wrong both ways.

    Steps

    1. Read the TechCrunch story to see exactly what the Pope said and how people reacted.
    2. Open c2pa.org and note who sits on the steering committee and what a manifest contains.
    3. Open contentauthenticity.org and find the open source c2patool and the Verify website.
    4. Connect it to something familiar: a manifest works like a signed git commit history for a picture, and a screenshot is like copying the files without the .git folder.
    5. Self check: if an image has no Content Credentials, what can you honestly conclude about who or what made it?
    Try it30 min · 5 steps
    You need: A computer with a browser, and optionally c2patool from the Content Authenticity Initiative. No API key needed.

    Steps

    1. Find an image that carries Content Credentials, for example one exported from a tool that supports C2PA or a sample from the C2PA site.
    2. Upload it to the Verify website at contentauthenticity.org and write down what the manifest says: tool, edits, signer.
    3. Take a screenshot of the same image, upload the screenshot, and compare the result.
    4. If you install c2patool, run it on both files and compare the output in your terminal.
    5. Write three lines on what provenance can and cannot prove about this image.

    Small angles to try

    • Try a photo you shot yourself on a phone that supports Content Credentials.
    • Run the image through a social upload and download it again to see if the label survives.
    • Compare with what a free AI image detector says about the same file.
  2. #2Tools & open source

    Decision models come to llama.cpp: score options locally, no text generation

    Hugging Face and the ggml team published a guide to running decision models in llama.cpp through a new /v1/systemone server endpoint. A decision model reads the input once and returns a probability for each option you give it, which suits routing, yes or no checks and ranking. Supported models include Julia-1 at 144M, Laya at 421M, Kev-4B, lev and OpenJev at 27B.

    Heat · X: Georgi Gerganov post about 1.9K likes · Laya and Julia-1 on Hugging Face trending
    What people foundThe support lives in llama.cpp PR #29818, which was still a draft waiting for review from the code owners when checked, and its author notes most of the code is AI written. Treat it as a preview build, not a stable release.
    Learn it15 min · 5 steps

    Key ideas

    decision model
    A model that scores a fixed list of answers instead of writing free text, so the output is a probability per option.
    question types
    The endpoint accepts choice questions with criteria, noul questions for yes or no, and score questions for ordering a list.
    one pass inference
    Because the model reads the input once and does not generate tokens, latency stays close to a single forward pass.

    Steps

    1. Read the Hugging Face blog post on decision models in llama.cpp and copy its routing example.
    2. Look at the model table and note the size, base model and languages of each one.
    3. Open PR #29818 on GitHub and check its current status and review comments.
    4. Connect it to something familiar: this is like a classifier head, but the labels are written in the request, not fixed at training time.
    5. Self check: why can a 421M decision model beat a chat model at routing tickets, and when would it fail?
    Try it30 min · 5 steps
    You need: A llama.cpp build that includes the decision model PR, a machine that can run a 4B GGUF model, curl. No API key needed.

    Steps

    1. Build llama.cpp from the PR branch, since the feature may not be in a release yet.
    2. Start the server with llama serve -hf ggml-org/Kev-4B-GGUF as shown in the blog.
    3. Send a POST to http://localhost:8080/v1/systemone with a state text and one choice question that routes a support message to billing, shipping or technical.
    4. Write 20 support messages of your own and record the top option and its probability for each.
    5. Compare accuracy and latency with asking a small chat model the same question.

    Small angles to try

    • Try a noul question such as whether a message contains personal data.
    • Ask the same question in Japanese and English and compare probabilities.
    • Plot probability against correctness to see if the scores are well calibrated.
  3. #3Engineering & CS

    Linux kernel CVEs jump as AI bug hunters flood maintainers

    The Linux kernel has logged about 7,178 CVEs so far in 2026 against 5,681 in all of 2025, with September alone above 2,000, according to LinuxCVETracker. Greg Kroah-Hartman showed CVE counts per release climbing from about 500 in the 6.x era to over 1,500 in 7.2. Networking maintainer Jakub Kicinski estimated a third to a half of recent net-next patches were low priority AI driven fixes.

    Heat · X: LaurieWired thread about 2.8K likes · HN: Kroah-Hartman talk about 150 points
    What people foundOnly 3 of this year's kernel CVEs are on CISA's Known Exploited Vulnerabilities list, so the count measures how many bugs get fixed and labeled, not how many are being attacked. Kicinski put the maintainer side plainly: they are completely overwhelmed.
    Learn it15 min · 5 steps

    Key ideas

    kernel CNA
    Since 2024 the kernel team assigns CVEs itself and gives one to almost every fix that could matter for security, which raises counts a lot.
    KEV list
    CISA's list of vulnerabilities seen exploited in the wild, a better signal of real risk than raw CVE totals.
    triage load
    Every reported bug costs maintainer time to read and review, even when the fix is small, which is why AI found bugs strain small teams.

    Steps

    1. Read the Tom's Hardware article and note the per release CVE numbers from Kroah-Hartman's slide.
    2. Open LinuxCVETracker for 2026 and look at the monthly chart and the KEV count.
    3. Connect it to something familiar: think of a code review queue at work that suddenly gets ten times more small pull requests.
    4. Watch part of the Kernel Recipes talk on security in the LLM age if you have time.
    5. Self check: why can CVE count go up sharply while real world risk stays roughly flat?
    Try it30 min · 5 steps
    You need: Python 3 with matplotlib, a copy of the kernel CVE data you can download from a public feed. No API key needed.

    Steps

    1. Pick one source of kernel CVE records, such as the NVD data feed or the kernel's own CVE announce archive.
    2. Count CVEs per month for 2024 to 2026 and plot them.
    3. Mark on the plot when the kernel became its own CNA and when AI bug hunting tools became common.
    4. Pull the KEV list from CISA and count how many kernel entries it has per year.
    5. Write two sentences comparing the two curves.

    Small angles to try

    • Split counts by subsystem such as net, fs and drivers.
    • Check how many CVEs your own distro kernel actually patched last month.
    • Compare with CVE growth for another big C project like curl or FFmpeg.
  4. #4AI in the world

    Hinton, Bengio and lab leaders warn automated AI research could speed up fast

    A Cambridge policy paper by 22 authors, including Geoffrey Hinton, Yoshua Bengio, OpenAI chief scientist Jakub Pachocki, Anthropic cofounder Jack Clark and Microsoft's Eric Horvitz, argues AI is on track to automate most AI research within a few years. They say a year of progress could then happen in weeks. They propose reporting on research automation, auditors inside labs, incident reporting and a way to pause datacenter research.

    Heat · X: Hinton post about 3.8K likes, 2.9K bookmarks · Axios coverage
    What people foundThe notable part is who signed it: people running research at OpenAI, Anthropic and Microsoft, writing in a personal capacity. The authors also say the outcome is uncertain and could be stopped by compute limits or diminishing returns.
    Learn it15 min · 5 steps

    Key ideas

    recursive self improvement
    The idea that AI systems help build better AI systems, which then speed up the next round.
    AI R&D automation
    Using AI agents to run experiments, write training code and analyze results that human researchers did before.
    compute bottleneck
    Even fast automated research needs GPUs to run experiments, which can cap how fast progress grows.

    Steps

    1. Read the Axios summary first for the main claims and signers.
    2. Open the CASP report page and read the policy proposals section.
    3. Connect it to something familiar: compare it with a build system where faster compilers make the next compiler faster, and ask what still limits speed.
    4. List which proposals would need new laws and which a lab could adopt alone.
    5. Self check: what single measurement would tell you research automation is really speeding up?
    Try it30 min · 5 steps
    You need: A browser and a spreadsheet. No API key needed.

    Steps

    1. Pick one public AI benchmark that has results over several years, such as a coding benchmark leaderboard.
    2. Write down the best score by quarter for 2024 to 2026.
    3. Plot it and fit both a straight line and a curve that bends upward.
    4. Note which fit matches better and how sensitive that is to the last two points.
    5. Write a short note on what this can and cannot say about the paper's claim.

    Small angles to try

    • Repeat with training compute estimates instead of scores.
    • Add the share of AI written code reported by labs as a second series.
    • Ask a friend to predict the next point before you reveal it.
  5. #5Trends × tech

    A Japanese artist's 'insanely fast' website goes viral: why tiny pages win

    Pixel and 3D artist 空論/秒 posted that they made a website and it loads extremely fast, and millions opened ku-ron.com to check. The page is one main image, four icon links and a note that it is under construction, with no scripts or framework visible in a fetch. It is a good live example of how much speed comes from simply sending less.

    Heat · X Japan: about 44K likes, 7.5M views, 13K bookmarks
    What people foundNothing clever seems to be going on: plain static HTML, no JavaScript and a few images. The remaining weight is mostly the PNG files, which modern formats could shrink further.
    Learn it15 min · 5 steps

    Key ideas

    LCP
    Largest Contentful Paint, the time until the biggest visible element, often the hero image, appears.
    TTFB
    Time to First Byte, how long the server and network take before the first byte of the page arrives.
    render blocking
    Scripts and stylesheets the browser must load before it can draw, which a page with none avoids completely.

    Steps

    1. Open ku-ron.com and view the page source to see how small it is.
    2. Read the web.dev guide on LCP to learn what counts as good.
    3. Connect it to something familiar: compare the source with a typical framework landing page and count script tags.
    4. Open the Network tab and list every request with its size.
    5. Self check: if the page is already fast, which single change would help most, and why?
    Try it30 min · 5 steps
    You need: Chrome with DevTools and Lighthouse built in. No API key needed.

    Steps

    1. Open ku-ron.com in Chrome, open DevTools, and run a Lighthouse performance report on mobile settings.
    2. Write down total bytes, request count, TTFB and LCP.
    3. Run the same report on a framework based landing page of your choice.
    4. Make a table comparing the two pages.
    5. Save the main PNG, convert it to WebP or AVIF with any image tool, and note the size saved.

    Small angles to try

    • Throttle the network to slow 3G and compare again.
    • Test from a phone on mobile data.
    • Build your own one file page with the same content and measure it.
  6. #6Tools & open source

    HyperQwen: Qwen3.8-27B on one RTX 3090, and what 381 tokens per second means

    HyperQwen, an Apache licensed project from syv-ai, runs Qwen3.8-27B on a single 24GB RTX 3090 with patched vLLM, 4 bit weights and speculative decoding. Its README reports about 127 tokens per second for one user and about 1,035 tokens per second across 64 concurrent requests. A viral post quoted 381 tokens per second.

    Heat · X: about 520 likes, 620 bookmarks · GitHub: about 1.8K stars
    What people foundThe 381 figure comes from a fork's test where the model copies a 25K token document already in its prompt, so a drafter that reads from the prompt guesses almost every token. For normal answers, plan around the 127 number.
    Learn it15 min · 5 steps

    Key ideas

    speculative decoding
    A small fast drafter proposes several tokens and the big model checks them in one pass, keeping only the ones it agrees with.
    prompt lookup drafting
    The drafter copies likely next tokens from text already in the prompt, which works very well when the answer quotes the input.
    W4A16
    Weights stored in 4 bits while math runs in 16 bits, cutting memory so a 27B model fits on a 24GB card.

    Steps

    1. Read the HyperQwen README and write down each speed number with its setup.
    2. Read the awoodka fork README to see how the 381 test was built.
    3. Connect it to something familiar: it is like a CPU branch predictor, very fast when the pattern repeats and slow when it does not.
    4. Find what acceptance rate means in speculative decoding and why it changes with the task.
    5. Self check: why is a summarization prompt faster than a creative writing prompt with the same model?
    Try it30 min · 5 steps
    You need: An NVIDIA GPU with 24GB, Docker with GPU support, the HyperQwen repo. No API key needed.

    Steps

    1. Clone the repo with git clone https://github.com/syv-ai/HyperQwen && cd HyperQwen and copy .env.example to .env.
    2. Start the single user profile with docker compose --profile single up -d.
    3. Send three prompts: a free question, a code task, and a request to repeat a long pasted document.
    4. Record tokens per second for each prompt.
    5. Compare the numbers and explain the gap using acceptance rate.

    Small angles to try

    • Try the batch profile with docker compose --profile batch up -d and measure total throughput.
    • Use a Japanese document for the copy test.
    • Run the same prompts in llama.cpp on the same card for a baseline.
  7. #7AI in the world

    Physicist Matthew Schwartz: Claude helped produce 36 papers in three months

    In a guest post on Anthropic's site, Harvard physicist Matthew Schwartz says his Claude setup explored about 400 candidate problems and produced 36 manuscripts across 18 fields with 19 coauthors in three months. He released the toolkit, BootLoops, as open source for any LLM. Examples include 30 elliptic integral results and 4,452 economics papers ported to open source code.

    Heat · X: about 1.2K likes for a summary post · Anthropic research post
    What people foundSchwartz is candid that early results were technically correct, but scientifically unremarkable until domain experts steered the work, and that it needs heavy human checking and a lot of compute. He was a visiting researcher at Anthropic.
    Learn it15 min · 5 steps

    Key ideas

    research loop
    An agent cycle of proposing a problem, attempting it, checking the result and choosing the next step.
    verification bottleneck
    When AI produces results faster than experts can check them, checking becomes the slow step.
    reproduction first
    Starting by reproducing known results lets you trust the pipeline before you trust new ones.

    Steps

    1. Read Schwartz's post on Anthropic's site and list the fields and numbers he gives.
    2. Note his warnings about quality and checking.
    3. Open the BootLoops repo and read how a loop is defined.
    4. Connect it to something familiar: it resembles a CI pipeline where every claim must pass checks before merging.
    5. Self check: what in his setup decides whether a result is worth a paper?
    Try it30 min · 5 steps
    You need: Python, the BootLoops repo, an API key for an LLM of your choice.

    Steps

    1. Clone the BootLoops repo and read its README for setup.
    2. Pick a small problem with a known answer in your own field, such as a closed form for a sum or a known algorithm bound.
    3. Run one loop and save every step the agent takes.
    4. Check the final answer against the known result by hand or with a test.
    5. Count how many steps needed your correction.

    Small angles to try

    • Try a second LLM and compare correction counts.
    • Give it a problem with no known answer and judge the result.
    • Time how long your checking took versus the agent's work.
  8. #8Research

    Neuralink pretrains a brain decoder on 50,000 hours of unlabeled signals

    Neuralink says its trial participants have logged over 50,000 hours with implants, and it used that unlabeled data to pretrain self supervised neural encoders. It reports decoders that kept working for weeks without recalibration and a record of 11.32 bits per second for cursor control. Secondary coverage says weekly recalibration fell from about 55 minutes to about 10.

    Heat · X: Neuralink post about 4.5K likes
    What people foundThis is the same recipe that worked for language: pretrain on lots of unlabeled data, then fine tune with a little labeled data. The numbers are Neuralink's own and are not independently reviewed yet.
    Learn it15 min · 5 steps

    Key ideas

    self supervised pretraining
    Learning useful features by predicting hidden parts of the data itself, with no human labels.
    signal drift
    Neural recordings change day to day as tissue and electrodes shift, which breaks decoders trained on old data.
    bits per second
    A measure of how much information a user can send through the interface per second, used to compare BCIs.

    Steps

    1. Read Neuralink's update on pretraining with 50,000 hours.
    2. Read the TeslaNorth summary for the recalibration numbers and mark which come from Neuralink.
    3. Connect it to something familiar: it is like pretraining a speech model on hours of raw audio before training on a few labeled clips.
    4. Find why drift makes decoders need recalibration.
    5. Self check: what would a shared model across many participants make possible?
    Try it30 min · 5 steps
    You need: Python with numpy and scikit-learn. No API key needed. No hardware needed.

    Steps

    1. Simulate 50 noisy channels driven by a two dimensional cursor velocity.
    2. Add slow drift by changing each channel's tuning a little every simulated day.
    3. Train a linear decoder on day 1 and measure its error on days 2 to 10.
    4. Train an encoder such as PCA on unlabeled data from all days, then a decoder on top, and compare errors.
    5. Plot error over days for both.

    Small angles to try

    • Add channel dropout to mimic failing electrodes.
    • Use an autoencoder instead of PCA.
    • Measure how many labeled minutes each method needs to recover.
  9. #9Trends × tech

    Hanshin at magic number 1: 胴上げ阻止 trends, and how clinch math works

    Hanshin went into Saturday needing one win or draw for its first back to back Central League pennant, and fans also trended 横浜応援 because a Giants loss to DeNA would clinch it too. Hiroshima won the day game 6 to 4, so 胴上げ阻止 trended and the magic number stays at 1. Clinch math is a small, exact algorithm over the standings and the remaining schedule.

    Heat · X Japan trending: こいほー #3, 胴上げ阻止 #8, #阪神タイガース #9 · Yahoo News pickup
    What people foundNPB has draws, so the clinch condition is not just wins, which is why sites like freefielder publish live magic numbers that fans check during games.
    Learn it15 min · 5 steps

    Key ideas

    magic number
    The combined count of your wins and rival losses needed to make first place certain.
    draws in NPB
    Ties are possible and ranking uses winning percentage, which changes the usual MLB style formula.
    elimination check
    For every rival, test whether it can still pass you if it wins every remaining game.

    Steps

    1. Read the Daily Sports story on Hanshin's magic number 1 situation.
    2. Open freefielder's magic number page and see how it lists conditions.
    3. Connect it to something familiar: it is a worst case search, like checking a deadline against the slowest path in a schedule.
    4. Work out by hand why a draw also clinches for Hanshin.
    5. Self check: why does winning percentage with draws make the simple formula wrong?
    Try it30 min · 5 steps
    You need: Python 3. No API key needed.

    Steps

    1. Copy the current Central League standings and each team's remaining games into a small table.
    2. Write a function that returns the best possible final winning percentage for each rival.
    3. Write a function that checks whether Hanshin's guaranteed minimum beats every rival's best.
    4. Enumerate tonight's outcomes for Hanshin and the Giants and print which ones clinch.
    5. Compare your output with freefielder's number.

    Small angles to try

    • Add tiebreak rules for equal winning percentage.
    • Run a Monte Carlo season finish with 50% win odds per game.
    • Do the same for the Pacific League.
  10. #10Research

    Meta: mathematicians plus Muse Spark settle open problems in six papers

    Meta says mathematicians working with Muse Spark 1.1 and 1.2 in thinking mode produced six papers on open problems in probability, PDEs, group theory, optimization and other areas. One disproves a 2024 conjecture with a 384 element group as a counterexample, and another settles a question open since 2015. Meta itself notes that other teams independently announced solutions to some of the same problems.

    Heat · X: AI at Meta post about 880 likes
    What people foundThe format is the useful part: humans chose problems and checked proofs, the model searched and drafted. Counterexamples like the 384 element group are the easiest kind of AI result to verify, because a computer can check them directly.
    Learn it15 min · 5 steps

    Key ideas

    counterexample
    A single concrete case that breaks a general claim, which is fully checkable by computation.
    monomial group
    A group whose irreducible representations all come from one dimensional ones, the property the conjecture was about.
    human in the loop proof
    A workflow where the model proposes steps and a mathematician verifies and directs them.

    Steps

    1. Read Meta's blog post and list the six problems and their areas.
    2. Pick the group theory result and read its abstract.
    3. Connect it to something familiar: finding a counterexample is like finding a failing test case for a function everyone thought was correct.
    4. Note which results are proofs and which are counterexamples.
    5. Self check: why are counterexamples easier to trust than long proofs?
    Try it30 min · 5 steps
    You need: Python, optionally GAP or SymPy for group checks. No API key needed.

    Steps

    1. Choose a simple claim from number theory that is known false for some value, such as a polynomial that gives primes for small inputs.
    2. Write a brute force search that tests the claim for growing inputs.
    3. Record the first counterexample and how long the search took.
    4. Ask any LLM to find a counterexample to the same claim and compare.
    5. Verify the LLM's answer with your script.

    Small angles to try

    • Use a group theory library to check a small group property.
    • Try a claim with no known counterexample and see what the model does.
    • Time model plus check against brute force alone.
  11. #11Research

    Google moves federated learning into TEEs with verifiable privacy

    Google Research described a new federated learning system where aggregation runs inside Trusted Execution Environments on servers, with a key manager using RAFT consensus and public logs of data processing policies. The paper says this gives externally verifiable central differential privacy for the first time. Gboard already uses it in production for English and Japanese next word prediction.

    Heat · X: Google Research post about 1.1K likes
    What people foundThe big shift is trust you can check: instead of trusting Google's servers, anyone can compare the published, reproducibly built binaries against what the TEE attests it is running. Training that took one to two months is also faster.
    Learn it15 min · 5 steps

    Key ideas

    federated learning
    Training a shared model from many devices without uploading their raw data.
    TEE attestation
    A hardware signed proof of exactly which code is running inside a protected enclave.
    central differential privacy
    Adding calibrated noise once at the server so the final model reveals little about any one user.

    Steps

    1. Read Google's blog post on provably private learning from federated data.
    2. Skim the arXiv abstract for the main guarantees.
    3. Connect it to something familiar: attestation is like checking a download's hash, but done by the CPU for running code.
    4. Look at the google-parfait repos to see what is published.
    5. Self check: what does a user have to trust in this design that they did not before, and what no longer?
    Try it30 min · 5 steps
    You need: Python with numpy. No API key needed.

    Steps

    1. Simulate 1,000 clients that each hold a small private dataset of numbers.
    2. Compute a federated average where each client sends a clipped update.
    3. Add Gaussian noise once at the server and measure error against the true mean.
    4. Compare with each client adding its own noise.
    5. Plot error for both approaches as the client count grows.

    Small angles to try

    • Vary the clipping bound.
    • Try 100 versus 10,000 clients.
    • Use a tiny next word model on your own text.
  12. #12Tools & open source

    Muse Gadgets SDK: build your own hardware for Meta's Muse assistant

    Nat Friedman announced Muse Gadgets, an Apache licensed SDK from Meta for building devices that work with the Muse assistant. It has an ESP32 SDK for adding a screen, audio or sensors, and a Linux SDK that turns a Raspberry Pi or spare Linux box into a gadget. You need an SDK token from gadgets.muse.ai and the Muse app in developer mode.

    Heat · X: Nat Friedman post about 4.9K likes · HN front page about 160 points
    What people foundFriedman suggests pointing a coding agent at the repo to build a device, which is a sign of how SDKs are now written for agents as much as for people.
    Learn it15 min · 5 steps

    Key ideas

    ESP32
    A cheap microcontroller with built in WiFi and Bluetooth, common in hobby and smart home devices.
    device token
    A secret that links a device to your account, which must be kept out of public code.
    thin client
    A device that captures input and plays output while the assistant runs in the cloud.

    Steps

    1. Read the muse-gadget-sdk README and note the two SDKs.
    2. Check the requirements: token, terms, app developer mode.
    3. Connect it to something familiar: a gadget is a smart speaker you assemble yourself.
    4. Look at the HN thread for concerns about privacy and lock in.
    5. Self check: what data leaves the device, and when?
    Try it30 min · 5 steps
    You need: A Raspberry Pi or any Linux machine with a microphone, a Muse account, an SDK token, the Muse mobile app.

    Steps

    1. Create an SDK token on gadgets.muse.ai and accept the terms.
    2. Enable developer mode in the Muse app.
    3. Clone the repo and follow the Linux SDK README.
    4. Run the example and ask a question by voice.
    5. Measure the time from end of speech to start of the reply.

    Small angles to try

    • Try the ESP32 SDK on a board you already have.
    • Add a sensor and expose its reading.
    • Ask a coding agent to build a new gadget from the repo and note what it gets wrong.
  13. #13Models & launches

    Cantina apex-flash-1: open weights security model trained on paid bug finds

    Security firm Cantina released apex-flash-1, an MIT licensed model post trained from GLM-5.3-Flash on real vulnerabilities its team found and was paid for. It is meant to work as a focused worker under a larger agent. On Cantina's own held out set of 60 tasks it reports 66.7% pass@1 at about $2.38 total, against 71.7% for Claude Opus 5 High at about $75.

    Heat · X: Cantina post about 1.2K likes, 1.2K bookmarks
    What people foundThe cost gap is the pitch: close to frontier results for about a thirtieth of the price. RuntimeWire notes the benchmark is Cantina's own and not independent, so wait for outside evaluations.
    Learn it15 min · 5 steps

    Key ideas

    post training
    Further training of a base model on a focused dataset to shift what it is good at.
    held out evaluation
    Testing on cases the model never saw in training, here 60 tasks from 20 vulnerability cases.
    pass@1
    The share of tasks solved on the first attempt.

    Steps

    1. Read the model card on Hugging Face.
    2. Read RuntimeWire's article for the training data size and caveats.
    3. Connect it to something familiar: it is a specialist code reviewer hired for one job, guided by a generalist lead.
    4. Compare the cost per task of each model in the table.
    5. Self check: why does a vendor's own benchmark need outside confirmation?
    Try it30 min · 5 steps
    You need: A browser and a spreadsheet. No API key needed.

    Steps

    1. Copy the three rows from Cantina's results table into a spreadsheet.
    2. Compute cost per solved task for each model.
    3. Plot pass@1 against cost on a log scale.
    4. Add any public scores for GLM-5.3-Flash from other benchmarks you can find.
    5. Write what the chart suggests about specialist versus general models.

    Small angles to try

    • Add a column for size and license.
    • Compare to another open security model you know.
    • Check back later for independent evaluations and update the chart.
  14. #14Tools & open source

    ESP32-C3 AdBlock: a $2 board blocks ads for your home network

    An MIT licensed project turns a $2 ESP32-C3 SuperMini into a DNS ad blocker for the whole home network. It stores up to about 537,000 blocked domains as sorted hashes in flash and finds each one with binary search, using about 40 KB of RAM. It has a web dashboard and can update its firmware and blocklist over WiFi.

    Heat · X: about 2.5K likes, 2.7K bookmarks
    What people foundThe neat trick is the data structure, not the hardware: hashing domains and binary searching flash means about 17 flash reads per lookup, so a board with 400 KB of RAM can hold a blocklist that would never fit in memory.
    Learn it15 min · 5 steps

    Key ideas

    DNS sinkhole
    A DNS server that answers 0.0.0.0 for blocked domains so ads never load.
    binary search on flash
    Searching a sorted array stored in flash memory takes about log2 of N reads, around 19 for half a million entries.
    hash collisions
    Storing short hashes saves space but two domains can share one, causing rare false blocks.

    Steps

    1. Read the esp32-c3-adblock README.
    2. Note how the blocklist is built and stored.
    3. Connect it to something familiar: it is Pi-hole with the database replaced by one sorted file.
    4. Work out how many bits of hash keep collisions rare for 537,000 domains.
    5. Self check: why is a sorted array better than a hash table here?
    Try it30 min · 5 steps
    You need: An ESP32-C3 SuperMini with 4MB flash, PlatformIO, Python 3, your own home network. No API key needed.

    Steps

    1. Copy src/secrets.example.h to src/secrets.h and add your WiFi details.
    2. Build the blocklist with python3 tools/build_blocklist.py data/blocklist.bin.
    3. Flash with pio run -t upload and pio run -t uploadfs.
    4. Check with dig @<c3-ip> doubleclick.net, which should return 0.0.0.0.
    5. Point one test device's DNS at the board and time page loads with and without it.

    Small angles to try

    • Without hardware, write the same sorted hash lookup in Python and count reads per query.
    • Measure the false block rate for 32 and 40 bit hashes.
    • Compare latency against your router's normal DNS.
  15. #15Trends × tech

    FamilyMart's quote repost coupon draw: how instant win prizes are paced

    FamilyMart Japan is running a 7 day campaign: follow @famima_now, quote repost the post, tap the image, and learn instantly if you won a free American Dog coupon, with 5,000 winners from October 2 to 8. The post has far more reposts than likes, because the repost is the entry. Behind it, a landing page checks the follow and repost, draws a result and issues a digital coupon.

    Heat · X Japan: about 180K reposts, 36K likes, 4.1M views
    What people foundHow the draw is paced decides fairness: a fixed chance per entry can run out of prizes early, while splitting prizes across time slots keeps late entrants in the game. FamilyMart has not published which method it uses.
    Learn it15 min · 5 steps

    Key ideas

    instant win
    A draw where each entry gets a result right away instead of after the campaign ends.
    prize pacing
    Spreading a fixed number of prizes over time so they do not all go on day one.
    odds per entry
    The chance a single entry wins, which changes over time when prizes are paced.

    Steps

    1. Read the campaign summary to note dates and winner count.
    2. List the steps a user takes and what the system must check at each.
    3. Connect it to something familiar: prize pacing is rate limiting, applied to giveaways.
    4. Sketch both draw methods on paper.
    5. Self check: with 2 million entries mostly on day one, which method is fairer to a day seven entrant?
    Try it30 min · 5 steps
    You need: Python 3 with numpy and matplotlib. No API key needed.

    Steps

    1. Simulate 2 million entries over 7 days, heavy on day one.
    2. Method A: give each entry a fixed win chance so expected winners are 5,000.
    3. Method B: split 5,000 prizes evenly into hourly slots and draw within each slot.
    4. Plot prizes left over time and the win chance by hour for both.
    5. Note when method A runs out and what late entrants see.

    Small angles to try

    • Add a cap of one win per user.
    • Make entries follow a daily wave.
    • Test what happens if the bot share is 10%.
  16. #16Tools & open source

    OpenDots: a self hosted, open source take on always on AI coworkers

    CopilotKit released OpenDots, an MIT licensed template it calls an open source alternative to OpenAI Dots. Each Dot gets its own browser profile, files and terminal, can be reached by web, voice and Slack, and asks for human approval before saving. It works with any OpenAI compatible model.

    Heat · X: Atai Barkai post about 5.3K likes
    What people foundThe README says core features were checked locally in September, while Slack and spoken task handoff still need testing, so start with the web interface.
    Learn it15 min · 5 steps

    Key ideas

    always on agent
    An agent with its own machine and memory that keeps working between your messages.
    approval gate
    A step where a human must confirm before the agent saves or acts.
    OpenAI compatible API
    Any server that speaks the same chat API, so local models can be swapped in.

    Steps

    1. Read the CopilotKit blog post introducing OpenDots.
    2. Read the README sections on setup and limitations.
    3. Connect it to something familiar: a Dot is like a remote dev box with a teammate logged in.
    4. List what each Dot can access on its computer.
    5. Self check: which approval gate would you add before giving it your email?
    Try it30 min · 5 steps
    You need: Node.js 24 and npm, an API key for any OpenAI compatible model or a local server.

    Steps

    1. Clone with git clone https://github.com/CopilotKit/OpenDots.git and enter the folder.
    2. Run npm ci and copy .env.example to .env, then set your model provider.
    3. Start with npm run dev and open http://127.0.0.1:5173.
    4. Give a Dot a small research task in a sandbox folder.
    5. Note every approval it asks for and how long the task took.

    Small angles to try

    • Point it at a local model server.
    • Run two Dots on related tasks.
    • Compare the same task in another agent tool you use.
  17. #17Research

    Ataraxos in Nature: superhuman Stratego for under $8,000 of training

    A team from CMU, NYU, Stanford and MIT published Ataraxos in Nature, a method for imperfect information games. It beat top Stratego player Pim Niemeijer 15 wins, 1 loss and 4 draws, and is state of the art in Hanabi and strong in Dou Dizhu. Training cost under $8,000, about a 500th of the compute of DeepMind's DeepNash.

    Heat · HN: Ars Technica story about 140 points
    What people foundThe cost is the story: superhuman play in a hidden information game used to need a big lab budget, and now fits a small research grant.
    Learn it15 min · 5 steps

    Key ideas

    imperfect information game
    A game where players cannot see everything, like hidden pieces in Stratego or cards in poker.
    regularized equilibrium
    A strategy that is hard to exploit, found by balancing reward against staying close to a reference policy.
    self play
    Training by playing against copies of yourself.

    Steps

    1. Read the Nature abstract and results summary.
    2. Note the match results and compute comparison with DeepNash.
    3. Connect it to something familiar: bluffing in poker is the reason you cannot just search the game tree like chess.
    4. Look at the AtaraxosAI GitHub organization for code.
    5. Self check: why does hidden information break plain minimax search?
    Try it30 min · 5 steps
    You need: Python 3. No API key needed.

    Steps

    1. Implement Kuhn poker, a three card toy game.
    2. Write regret matching to train both players by self play.
    3. Run 100,000 iterations and print the strategy.
    4. Measure exploitability by computing a best response.
    5. Plot exploitability over iterations.

    Small angles to try

    • Swap in Leduc poker.
    • Add a regularization term and compare convergence.
    • Play against the trained strategy yourself.
  18. #18Engineering & CS

    Court blocks Utah's VPN rule: 'geolocation perfection is not possible'

    A federal judge paused the VPN provisions of Utah's SB 73, which required adult sites to block Utah VPN users or verify the age of every visitor. The court found the rule likely violates the Commerce Clause. EFF highlighted the judge's line that geolocation perfection is not presently possible.

    Heat · HN: top story, about 440 points
    What people foundThe ruling leans on an engineering fact: a server cannot reliably tell where a VPN user really is, so the law asked for something that cannot be built.
    Learn it15 min · 5 steps

    Key ideas

    IP geolocation
    Guessing a user's location from their IP address using databases, which is often wrong and easy to hide.
    VPN exit node
    The server whose IP address websites see when you use a VPN.
    preliminary injunction
    A court order that pauses a law while the case continues.

    Steps

    1. Read EFF's post on the ruling.
    2. Note what the law required and why the judge doubted it.
    3. Connect it to something familiar: think of how often a streaming site wrongly guesses your country.
    4. Look up how IP geolocation databases are built.
    5. Self check: what signals besides IP could a site use, and why are they weak too?
    Try it30 min · 5 steps
    You need: Python 3 and a free IP geolocation database file. No API key needed.

    Steps

    1. Download a free IP to country database such as one that offers a local file.
    2. Look up your own public IP with and without a VPN.
    3. Look up the IPs of a few public cloud regions you know.
    4. Record which answers are right and which are wrong.
    5. Write what a site would need to know to enforce the Utah rule.

    Small angles to try

    • Compare two different databases.
    • Check the database's city level accuracy for your IP.
    • Try mobile data versus home internet.

Get the Radar by email

One email each evening, Japan time, with that day's items. Opening soon: leave your email and we'll send a confirm link when it starts.