Tuesday, October 6, 2026
-
Japan's Digital Minister: stop reusing passwords, turn on MFA
On October 6 at the Digital Agency, Digital Minister Toshiharu Furukawa said cyberattacks are not only a problem for companies and asked each citizen to protect their own information. He named three habits: stop reusing passwords, use multi factor authentication, and be careful with suspicious emails and social posts. The remark came after a run of Japanese breaches, and his words 自分の情報は自分で守って spread across X Japan trends.
Heat · X Japan trending: 自分の情報 #2, デジタル庁 #3, 個人情報 #6, 不正アクセス頻発 #8, 情報漏洩 #15, パスワードの使い回し #18, 多要素認証 #26What people foundMuch of the reaction on X read it as the state pushing responsibility onto individuals. Still, the two habits he named are exactly what stops credential stuffing: a password leaked from one site cannot open a second site where it is unique, or an account that also asks for a second factor.Learn it15 min · 4 steps
Key ideas
- Credential stuffing
- Attackers take email and password pairs leaked from one service and try them automatically on many other services.
- TOTP
- A six digit code computed from a shared secret and the current time, defined in RFC 6238 and used by authenticator apps.
- k anonymity range query
- You send only the first five characters of a password hash, so the server never learns which password you checked.
Steps
- Read the IPA 不正ログイン対策特集ページ to see the official Japanese guidance on password reuse and two step login.
- Look at how it separates something you know from something you have, and why it recommends a password manager.
- Connect it to engineering: a unique random password per site works like an API key scoped to one service.
- Self check: if site A leaks your password, which of your accounts are still safe, and why?
Try it30 min · 5 steps
You need: Python 3 and pip, internet access for one HTTPS call. No API key needed.Steps
- Run
pip install pyotp, create a secret withpyotp.random_base32()and printpyotp.TOTP(secret).now()twice, 30 seconds apart, to watch the code rotate. - Turn the secret into a provisioning QR code, scan it with your own authenticator app, and confirm that
totp.verify(code)accepts the code the app shows. - Write a script that takes the SHA-1 of a test password locally and sends only the first 5 hex characters to
https://api.pwnedpasswords.com/range/{prefix}with the headerAdd-Padding: true. - Find your suffix in the response yourself, skip padded rows whose count is zero, and print how many times that password appears in known breaches.
- Compare a weak test password like password123 with one your password manager generated.
Small angles to try
- Count how many of your own accounts still lack a second factor.
- Implement TOTP from scratch with hmac and compare the output with pyotp.
- Measure how much the range query reveals with and without padding.
-
Reflection launches Beam, a 501B open weight MoE model
Reflection AI announced Beam, a sparse mixture of experts model with 501B total and 23B active parameters, pretrained from scratch on 23.8T tokens and released under Apache 2.0. Its blog says RL ran on 10.5K GB300 GPUs for four weeks with over 100 million rollouts, and claims parity with GLM 5.2 on hard reasoning at 3 to 4 times less inference compute. Weights are promised later this month; for now there is an early access signup.
Heat · HN front page #1, 375 points and 112 comments · X: Misha Laskin launch post 1,036 likesWhat people foundReflection's own table shows Beam at 44.4 on DeepSWE v1.1 against 44.0 for GLM 5.2, but behind GLM 5.3 on SWE Bench Pro v2 Hard, 77.2 against 84.3. TechCrunch notes none of the claims have been independently verified yet.Learn it15 min · 4 steps
Key ideas
- Active parameters
- In a mixture of experts model only a few experts run per token, so 23B of the 501B weights do the work for any one token.
- Asynchronous policy gradient
- Rollout workers keep generating with slightly old weights while the trainer updates, so the method has to correct for stale samples.
- Goodput
- The share of cluster time that actually ends up in the final model after failures and restarts, which Reflection puts at 92.3 percent.
Steps
- Read the Reflection blog post Introducing Beam, focusing on the RL section.
- Go through the benchmark table and mark which rows Beam wins and which it loses.
- Compare the 23B active size with other open MoE models you have run to estimate memory and serving cost.
- Self check: why can a 501B MoE model be cheaper per token than a smaller dense model?
Try it30 min · 4 steps
You need: A browser and a Reflection early access account; weights are not downloadable yet. No API key needed to read the specs.Steps
- Sign up for early access on the Reflection platform.
- Pick five coding and agent prompts you already use on GLM 5.2 or Qwen and save them as a fixed eval set.
- When access or weights land, run the same prompts and record answer quality and tokens used.
- Estimate the memory needed for 501B weights at 4 bits and decide which hosting option is realistic for you.
Small angles to try
- Track how fast Beam lands in local runtimes after the weights release.
- Compare tokens per solved task against GLM 5.2 to test the 3 to 4 times compute claim.
-
How did Shazam work without AI? Spectrogram peaks and hashing
A viral post asked how Shazam could identify songs 15 years ago without AI. Avery Wang's 2003 ISMIR paper answers it: the app keeps only the energy peaks of a spectrogram, pairs each peak with nearby later peaks, packs each pair into a 32 bit hash of two frequencies and a time gap, and then looks for many matches that share the same time offset. The paper reports searches of 5 to 500 milliseconds over about 20,000 tracks, even when only 1 to 2 percent of hashes survive the noise.
Heat · X: Nathan Ruff post, 39,818 likes and about 6.6M views, the top AI related post in today's scanWhat people foundWang's paper shows that pairing peaks makes each hash far more specific than a single peak: with a fan out of 10, storage grows 10 times while search gets about 10,000 times faster. For finding an exact recording, signal processing plus a hash table still beats a neural network.Learn it15 min · 4 steps
Key ideas
- Constellation map
- The spectrogram is reduced to a sparse scatter of its loudest time and frequency points, which survive noise and compression.
- Combinatorial hashing
- Each anchor peak is paired with peaks in a target zone after it, and each pair becomes one hash key.
- Offset histogram
- True matches all share the same gap between clip time and song time, so they form one tall bar.
Steps
- Read sections 2.1 to 2.3 of Wang's paper, hosted on a Princeton course page.
- Find the figure where matching hashes line up on a diagonal and become a histogram spike.
- Connect it to search engines: this is an inverted index, the same idea as posting lists.
- Self check: why does a hash of a peak pair survive noise better than a hash of the raw audio?
Try it30 min · 5 steps
You need: Python 3 with numpy and scipy, a few audio files you own. No API key needed.Steps
- Load two or three of your own songs and compute a spectrogram for each with scipy.
- Find local maxima with a maximum filter and a loudness threshold, and plot the constellation.
- Pair each peak with up to 10 later peaks in a small window and store each hash of frequency 1, frequency 2 and time gap in a dict with the song id and anchor time.
- Record 10 seconds of one song playing from your phone speaker, hash the clip the same way, and build a histogram of time offsets per song.
- Check that the right song shows one clear spike, then add noise to the clip and find where recognition fails.
Small angles to try
- Sweep the fan out from 1 to 20 and plot index size against accuracy.
- Try a pitch shifted clip and see why plain hashing breaks.
- Compare with an embedding model on cover songs, which fingerprints cannot handle.
-
Hugging Face turns Claude Code, Codex and 14 more harnesses into RL envs
Hugging Face's OpenEnv now ships a Harbor environment that wraps 16 coding agent harnesses, including Claude Code, Codex, opencode, Pi, Hermes and mini-swe-agent, as RL environments without modifying them. A capture proxy between the harness and your vLLM or SGLang server records token IDs and logprobs, which TRL's async GRPO trainer then learns from. The demo trained LiquidAI LFM2.5 2.6B across four harnesses at once.
Heat · X: Clem Delangue 2,098 likes in today's scanWhat people foundPer the published model card, multi harness RL lifted LFM2.5 2.6B from about 42 to 54 percent pass@1 averaged over four harnesses with 31 percent fewer tool calls, while training on teacher traces plateaued near 47.5 percent. The striking detail: the same weights can score 62 percent in one harness and 33 percent in another.Learn it15 min · 4 steps
Key ideas
- Harness
- The agent program around a model, such as Claude Code or opencode, that decides prompts, tools and the loop.
- Capture proxy
- A pass through server that logs exactly which tokens the model sampled and their probabilities so RL can train on the real agent run.
- GRPO
- A policy gradient method that scores a group of rollouts for the same task and pushes the model toward the better ones.
Steps
- Read the guide The ultimate guide to multi harness RL on the AdithyaSK Space.
- Then read the OpenEnv Harbor docs and look at the harness table and its wire dialects.
- Connect it to evals you know: SWE bench style grading, but inside the agent you actually ship.
- Self check: why would a model trained only inside opencode make invalid tool calls inside Codex?
Try it45 min · 5 steps
You need: Python 3.12 or newer, a GPU running vLLM, and a sandbox backend such as e2b. No paid model API needed if you serve an open model.Steps
- Install with
pip install "openenv[harbor]". - Serve a small model with
vllm serve <model> --return-tokens-as-token-ids --logprobs-mode processed_logprobs. - Check the setup with
openenv harbor info --llm-url $LLM --dataset AdithyaSK/data_agent_rl_environment_eval. - Run one rollout with
openenv harbor rollout --llm-url $LLM --dataset org/tasks --task-index 0 --harness opencode --sandbox e2b --out rollout.json, then repeat with another harness. - Compare reward and turn count between the two harnesses on the same task.
Small angles to try
- Serve
FineEnvs/LFM2.5-2.6B-multiharness-RLand compare it with the base LFM2.5 in your own harness. - Count invalid tool calls a local model makes in each harness.
-
Rakuten Drive breach: one stolen admin login, eight months of access
On October 6, Rakuten Drive said intruders used stolen credentials for a system administrator account between January 29 and September 17, 2026. Stored photos and documents from 15,382 accounts may have been viewed, profile data from 687 accounts was taken, and for 313 of those the attackers also got encrypted passwords and their salt values. The service blocked the access route, paused app downloads and new signups, and says it will notify affected users.
Heat · X Japan trending: 楽天ドライブ #11, alongside 不正アクセス頻発 #8 and 情報漏洩 #15What people foundNo clever exploit was needed: one stolen admin credential gave about eight months of quiet access, which argues for phishing resistant MFA on admin accounts and alerts when an admin reads user content. The notice does not name the password hashing algorithm or its cost setting, and that detail decides how worried the 313 users should be.Learn it15 min · 4 steps
Key ideas
- Salted password hash
- A random salt is mixed into each password before hashing, so two users with the same password get different stored values.
- Key stretching
- Schemes like Argon2 or bcrypt are deliberately slow, which raises the cost of every offline guess.
- Dwell time
- How long an attacker stays inside before being detected, here roughly eight months.
Steps
- Read the 週刊アスキー report and write down the date range and the three account counts.
- Note what the notice leaves out: the hash algorithm, the stretching cost and how the admin credential leaked.
- Think of an internal tool you run where one admin can read user data, and ask what would alert you.
- Self check: why does a salt alone not protect a weak password once the salt is stolen too?
Try it30 min · 5 steps
You need: Python 3 withpip install argon2-cffi, local machine only. No API key needed.Steps
- Hash a test password with
PasswordHasher().hash(...)and confirmverifyaccepts it. - Hash the same password twice and confirm the two stored strings differ because each gets its own salt.
- Time 100 verify calls, then 100 plain SHA-256 hashes of the same string, and compare guesses per second.
- Use a small wordlist of your own to estimate how long checking it against one Argon2 hash takes compared with SHA-256.
- Change the hasher parameters and use
check_needs_rehashto see how an old hash gets flagged for upgrade.
Small angles to try
- Draft a log rule that alerts when an admin account reads user files in bulk.
- Compare bcrypt and Argon2 timings on your laptop.
- Write the breach notice you would want to receive, with the missing technical details filled in.
-
Opus 5.5 agent team proposes two room temperature magnet candidates
Vals AI says a team of Claude Opus 5.5 agents searched in simulation for compensated magnetic semiconductors, materials whose magnetism cancels overall but which still split electrons by spin, useful for spintronics. They proposed a new five element oxide, YBaMnFeO₅, predicted to order up to about 420 K, and flagged a 1999 Prussian blue compound,
KV[Cr(CN)₆], whose original sample stayed magnetic to 376 K. The blog lists caveats: the oxide may be hard to make, and methods disagree on how water in the old sample changes things.Heat · HN front page #4, 263 points and 180 comments · X: Vals AI post 6,031 likesWhat people foundThis is a simulation result, not a lab discovery. The 90+ agents in 3 days figure is Vals AI's claim on X; the blog does not give agent counts. Vals did publish its full DFT inputs, outputs and analysis code, so others can check the work.Learn it15 min · 4 steps
Key ideas
- Compensated magnet
- A material whose magnetic moments cancel overall while its electron bands still split by spin, so it can carry spin currents without stray fields.
- DFT with PBE+U and HSE06
- Density functional theory estimates electronic structure; PBE+U is the fast approximation and HSE06 the slower check used to confirm band gaps.
- Spin window
- The energy range where only one spin direction of carriers exists, which sets how well the material can filter spin.
Steps
- Read the Vals AI blog post, focusing on the candidate table and the caveats section.
- Mark which claims rest on the fast PBE+U pass and which on the HSE06 check, and where the two disagree.
- Compare it with agent coding: fan out search plus verification, like a test suite filtering generated code.
- Self check: why is a predicted ordering temperature in simulation not the same as a measured one?
Try it30 min · 5 steps
You need: A laptop with git and Python for browsing files. No API key needed.Steps
- Clone the published ledger repository named in the Vals AI blog post.
- Browse the folder layout to find the input files and outputs for each candidate.
- Pick YBaMnFeO₅ and trace how the band gap number in the blog maps to an output file.
- Write down each step where an agent chose what to compute next, if the logs show it.
- Note one thing a human materials scientist should verify before trusting the result.
Small angles to try
- Ask a current model to critique the
KV[Cr(CN)₆]water caveat using only the repository files. - Estimate how much of the pipeline is agent judgment versus standard DFT tooling.
-
OpenAI Codex pledges 28 days of daily ships or usage resets
Codex lead Tibo Sottiaux posted that for the next 28 days the team will either ship one clear improvement relevant to most Codex users each day, or grant a full usage reset. A reset refills the weekly quota rather than raising limits. According to an explainx breakdown, the pledge follows a rough month of outages and quality complaints, plus a cut to the Pro plan multiplier from 20x to 10x effective October 30.
Heat · X: Tibo post 25,846 likes and about 7.1M views, the second biggest AI post in today's scanWhat people foundExplainx points out the deal is either or: users are not promised 28 resets, and a reset replaces leftover quota instead of adding to it, so saving usage for later does not pay off.Learn it15 min · 4 steps
Key ideas
- Usage reset
- A one time refill of the weekly Codex quota back to full.
- Plan multiplier
- How many times the base usage allowance a paid tier gets.
Steps
- Read Tibo's post and the explainx breakdown of the 28 day pledge.
- List what changed for Codex in September: outages, model issues and plan changes.
- Compare it with how other coding tools communicate limits and credits.
- Self check: if you have half your quota left on a reset day, what happens to it?
Try it15 min a day · 4 steps
You need: A Codex or ChatGPT plan. No extra API key needed.Steps
- Note your current Codex usage and plan limits today.
- Each day for a week, check the official Codex changelog and Tibo's posts for what shipped.
- Keep a simple log: date, improvement or reset, and whether it changed your workflow.
- At the end of the week, score how many shipped items actually mattered to you.
Small angles to try
- Compare daily Codex changes with Claude Code release notes over the same week.
- Track whether resets cluster around outage days.
-
Dust pretrains small transformers with no backpropagation
Q Labs released Dust, a zeroth order training method that adds noise to each layer's activations at every token and rewards the noise that lowers loss, so each token acts as one member of a search population. They trained models from 2M to 243M parameters on FineWeb, and at 10M to 20M tokens the gap with backprop shrinks as the population grows. Code is MIT licensed.
Heat · HN front page #5, 138 pointsWhat people foundThe authors are candid that Dust still needs orders of magnitude more compute than backprop. The claim worth noticing is that it is roughly 1,000 to 10,000 times more efficient than weight space evolution strategies, and that bigger models were more population efficient than expected.Learn it15 min · 4 steps
Key ideas
- Zeroth order optimization
- Training that only uses loss values from forward passes, never gradients from the chain rule.
- Node perturbation
- Adding noise to a layer's outputs instead of its weights, then turning reward weighted noise into a gradient estimate.
- Virtual population
- Treating every token position as a separate trial so thousands of perturbations are scored in one forward pass.
Steps
- Read the Q Labs Dust research page, starting with the method diagram and the plot against backprop.
- Find the token budgets where Dust loses badly and the ones where it catches up.
- Relate it to evolution strategies from RL, and ask why perturbing activations scales better than perturbing weights.
- Self check: why might a method with no backward pass matter for analog or neuromorphic hardware even if it costs more FLOPs?
Try it45 min · 5 steps
You need: Linux, Python 3.10 or newer, an NVIDIA GPU with bfloat16 support and a CUDA 12.8 driver. No API key needed.Steps
- Clone
https://github.com/qlabs-eng/dustand create a venv withpython3 -m venv .venv. - Install PyTorch with
python -m pip install torch==2.10.0 --index-url https://download.pytorch.org/whl/cu128, thenpython -m pip install -r requirements.txt. - Prepare data with
python prepare_data.py, which downloads FineWeb and builds the 1M and 10M token sets. - Run Dust with
python dust.py --tokens 1M --population 256 --output runs/1m-p256. - Run the baseline with
python baselines/backprop.py --tokens 1M --output runs/bp-1mand compare validation loss and wall time.
Small angles to try
- Halve and double the population and plot how loss changes.
- Measure peak GPU memory for Dust versus the backprop baseline.
-
ChatGPT signs fake New Yorker cartoons with real artists' names
Nieman Lab found ChatGPT's image generator producing New Yorker style cartoons signed with the names of real contributors, including a viral cartoon signed BLOPER, Brendan Loper's pen name, which he did not draw. At least 15 cartoonists' signatures showed up. After being contacted, ChatGPT began showing a warning for some prompts but kept signing some cartoons.
Heat · HN front page #11, 350 points and 254 commentsWhat people foundOpenAI called it a bug and unintended behavior. The sharper point from the story: a forged signature is an attribution problem, not just style copying, because it misleads readers about who made the work.Learn it15 min · 4 steps
Key ideas
- Memorized signatures
- Image models can learn to reproduce names that appear often in training images, like a signature in a corner.
- Attribution versus style
- Imitating a style is murky, but putting a real person's name on fake work misrepresents who made it.
- Output guardrail
- A filter that checks prompts or outputs for similarity to third party content.
Steps
- Read the Nieman Lab article and note the examples and OpenAI's response.
- Check whether the new warning stops the signature or only flags the prompt.
- Connect it to text models that invent fake citations, the same failure in a different medium.
- Self check: how would you design a release test that catches forged signatures?
Try it30 min · 4 steps
You need: Access to any image generator you already use. No API key needed.Steps
- Write a neutral prompt for a single panel magazine style cartoon with no artist named.
- Generate several images and zoom into the corners to look for any signature or name.
- Repeat with a different magazine style and note any difference.
- Keep the results for your own notes and do not post images that carry real people's names.
Small angles to try
- Compare two or three image models on the same prompt.
- Add the word unsigned to the prompt and see if it helps.
Sources:Nieman Lab -
AI written retro game PC ports split the recomp community
Native PC ports of old console games made by static recompilation are booming, and many are now written mostly by Claude, Qwen or local models. Dot Esports lists recent projects such as Wave Race 64, Mischief Makers and GoldenEye, one of which says every line was written by Claude while humans set goals and tested. Other projects such as DK64: ReKONGpiled now advertise themselves as fully human coded.
Heat · X: posts on AI decompiling games at 13,090 and 7,124 likes in today's scan · Reddit Lies post on AI optimizing software at 20,227 likesWhat people foundExperienced recomp developers quoted by PC Gamer and Dot Esports argue the AI ports are hard to maintain and do not feed shared tooling; supporters answer that a working port is a working port. Human coded is turning into a selling point.Learn it15 min · 4 steps
Key ideas
- Static recompilation
- Translating a game's machine code into native source ahead of time so it runs on PC without an emulator.
- Decompilation
- Recovering readable source code from a compiled binary, often matched byte for byte against the original.
- Maintainability debt
- Code that works today but that nobody understands well enough to fix or extend later.
Steps
- Read the Dot Esports roundup first, then the PC Gamer piece for the background of the fight.
- Note which parts AI does well, such as boilerplate and graphics shims, and what humans still handle.
- Compare with your own experience of agent written code: does it hold up after six months?
- Self check: what evidence would convince you an AI written port is maintainable?
Try it30 min · 5 steps
You need: A C compiler and a coding agent; a local model works too. No API key needed with a local model.Steps
- Compile a tiny C program of your own and strip the binary.
- Ask a coding agent to decompile it back into readable C from a disassembly.
- Recompile the output and compare behavior on the same inputs.
- Have a second session explain one function without seeing the original, and judge whether it is understandable.
- Only use binaries you wrote yourself, not game ROMs.
Small angles to try
- Compare a local Qwen model with a frontier model on the same binary.
- Count how many fix cycles the agent needs before your tests pass.
-
Steam hit Dressmaker treats sewing patterns like UV unwrapping
Dressmaker is a cozy sewing game from Cozy Lives, published by Free Lives, launched on September 21 and rated 98 percent positive across 7,260 Steam reviews. Dexerto reports developer Jonathan Hau-Yoon started it after his wife Sabrina Becker lost her ability to sew because of a medical condition, and she helped design and playtest it. It began as an itch.io prototype built in about three weeks.
Heat · X: Dexerto post 129,237 likes, the second most liked general post in today's scan · Steam: Overwhelmingly Positive, 98% of 7,260 reviewsWhat people foundThe team does not run real cloth physics. Hau-Yoon told Engadget it works like UV unwrapping: where you place a pattern piece on the fabric is exactly what shows on the 3D dress, and a lattice resizes the mannequin and the flat pattern pieces together for each customer.Learn it15 min · 4 steps
Key ideas
- UV unwrapping
- Flattening a 3D surface into 2D pieces so a flat texture can be mapped back onto it.
- Lattice deformation
- A coarse cage of control points bends a mesh smoothly, so resizing the cage resizes the model.
- Seam allowance
- The extra margin around a pattern piece that gets sewn, which the game calculates for you.
Steps
- Read the Engadget Indie Pitch interview with the team for the technical explanation.
- Look for how they keep the 3D dress and the flat patterns in sync when sizes change.
- Connect it to the texture mapping every 3D tool uses, here turned into the gameplay itself.
- Self check: why does a gathered edge stretch the texture more than a straight seam?
Try it30 min · 5 steps
You need: Blender or any free 3D tool with UV editing. No API key needed.Steps
- Model a simple skirt as a cylinder or cone in Blender.
- Mark two seams and unwrap the mesh to get the flat pattern pieces.
- Apply a striped fabric texture and move the pieces in the UV editor to watch the stripes change on the 3D skirt.
- Scale the waist with a lattice modifier and check how much the stripes distort.
- Compare your flat pieces with a real skirt sewing pattern.
Small angles to try
- Write a tiny shader that tints areas that are stretched too much.
- Generate fabric textures with AI and see which ones tile cleanly.
- Compare this trick with the cost of real cloth simulation in a game engine.
-
Cloudflare adds a Web Search API for agents, in beta
Cloudflare added a beta Web Search API that lets Workers and other apps run web searches through AI Gateway, with Ceramic.ai, Exa and Linkup as providers at launch. Requests are billed at each provider's list price with no markup, show up in AI Gateway logs, and all three providers offer zero data retention for this traffic. You can also bring your own provider keys.
Heat · HN front page, 509 points and 234 commentsWhat people foundThe practical win is one billing and logging layer for search grounding, with provider switching per call. HN discussion focused on whether this makes Cloudflare a broker between agents and the open web.Learn it15 min · 4 steps
Key ideas
- Search grounding
- Giving a model fresh search results so it answers from current pages instead of training data.
- AI Gateway
- Cloudflare's proxy layer that logs, caches and bills calls to model and tool providers.
- Zero data retention
- The provider promises not to store your queries or results.
Steps
- Read the Cloudflare changelog post, focusing on the REST endpoint and the Workers binding example.
- Note how provider choice and the result limit are set.
- Compare it with how you wire search into an agent today, for example a direct Exa key.
- Self check: what do you gain and lose by routing all search through one gateway?
Try it30 min · 5 steps
You need: A Cloudflare account with AI Gateway credits, an API token, Node.js and Wrangler for the Worker example.Steps
- Create or reuse an AI Gateway named default in your Cloudflare dashboard.
- In a Worker with the AI binding, call
env.AI.websearch({ gatewayId: "default", query: "...", provider: "exa", limit: 5 }). - Repeat the same query with the other two providers and save the results.
- Or call the REST endpoint
POST https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/websearch/from your machine. - Open the AI Gateway logs and compare latency and cost per provider.
Small angles to try
- Run Japanese language queries and see which provider handles them best.
- Feed results into a model and check how often answers cite the right page.
-
How should agents prove who they are? Web Bot Auth is a start
Nikita Bier argued on X that the most important tech problem of the next five years is a standard for AI agents to identify themselves to the services they use. One existing answer is Web Bot Auth, which Cloudflare supports: an agent signs each HTTP request with an Ed25519 key and points to a public key directory at a well known URL. It builds on two IETF drafts, one for the key directory and one for the signature protocol.
Heat · X: Nikita Bier post 4,563 likes in today's scanWhat people foundBier's framing is his opinion. The practical fact is that signed request headers already exist, but they prove which operator runs a bot, not which human it acts for or what it is allowed to do, and that gap is the open problem.Learn it15 min · 4 steps
Key ideas
- HTTP message signatures
- A standard way to sign selected parts of an HTTP request so the server can check who sent it.
- Signature-Agent header
- Tells the server where to fetch the bot's public keys.
- Delegated identity
- Proving an agent acts for a specific user with specific permissions, which bot signatures alone do not cover.
Steps
- Read Cloudflare's Web Bot Auth reference page and note the three headers.
- Look at which fields are signed and how expiry is handled.
- Compare it with OAuth and ask what agent scopes would need on top.
- Self check: if your agent books a hotel, which layer proves it was allowed to spend your money?
Try it30 min · 5 steps
You need: OpenSSL and a small web server you run on localhost in Node or Python. No API key needed.Steps
- Generate an Ed25519 key pair with OpenSSL on your own machine.
- Convert the public key to JWK format and serve it from your local server at the well known directory path described in the docs.
- Write a small client that signs requests with Signature-Input, Signature and Signature-Agent headers, using a library whose README you have checked.
- Add a verifier to your local server and confirm it accepts good signatures and rejects an expired one.
- Keep all signed traffic on localhost.
Small angles to try
- Add a field for which user the agent acts for and think about how a server could verify it.
- Rotate the key and see what breaks.
-
OpenAI adds invisible text watermarks to ChatGPT and Codex in the EU
OpenAI says it will add an invisible watermark called textGrain to eligible ChatGPT and Codex text output for EU users over the coming weeks. It nudges word choices into a statistical pattern a detector can find; API developers worldwide can opt in, but it is off by default. Detector access is limited to approved researchers and expert organizations for now.
Heat · BleepingComputer report, October 5What people foundOpenAI's own numbers show how fragile it is: swapping 10 percent of words for synonyms cut detection from about 92 percent to 66 percent. Treat it as a compliance signal, not proof that a given text is AI written.Learn it15 min · 4 steps
Key ideas
- Statistical text watermark
- A hidden bias in which words a model picks that readers cannot see but a detector can measure over enough tokens.
- Paraphrase attack
- Rewording or swapping synonyms to wash out the watermark pattern.
- API opt in
- Developers outside the EU decide whether their app's outputs carry the mark.
Steps
- Read the BleepingComputer article and note exactly which products and regions are covered.
- List what is still unclear, such as whether Codex code output is marked.
- Connect it to the EU AI Act transparency rules for synthetic content.
- Self check: why does detection get harder on short texts and on code?
Try it30 min · 4 steps
You need: Any chat model and Python. The detector is not public. No API key needed.Steps
- Generate a few paragraphs from a model and save them.
- Rewrite one copy by swapping about one word in ten for a synonym.
- Compare word frequency between the model text and your own writing with a small word count script.
- Write down what a team in Japan serving EU users would need to tell customers about marked output.
Small angles to try
- Think through whether marked Codex comments could end up in a public repository.
- Compare with image provenance such as C2PA metadata.
Sources:BleepingComputer -
Leviathan lets agents search 1M records in about 436 tokens
Leviathan is a new Apache 2.0 Rust tool that turns JSONL, JSON, CSV, SQLite or database exports into a ranked full text index stored as one SQLite file. Agents query it in plain language and get short cited result cards instead of reading raw data. The README claims a median of 436 tokens per question at 1M records versus about 107,000 for the best grep strategy.
Heat · X: launch post 2,058 likes in today's scan · GitHub repo still smallWhat people foundThe project's benchmark reports 99.0 percent of relevant records in the top 5 and a worst case of 602 tokens against 9.7M for grep; these are the author's numbers on the author's own dataset.Learn it15 min · 4 steps
Key ideas
- Full text index
- A precomputed lookup from words to records so search does not scan every row.
- Context budget
- The number of tokens an agent can spend reading data before cost and quality suffer.
Steps
- Read the Leviathan README, especially the benchmark section.
- Look at what a result card contains and how it cites the source record.
- Connect it to how coding agents already use grep and ripgrep over repositories.
- Self check: when would plain SQL be better than a ranked text index for an agent?
Try it30 min · 5 steps
You need: Rust 1.88 or newer, or a prebuilt binary from the Releases page. No API key needed.Steps
- Install with
cargo install --git https://github.com/elstongun/leviathan leviathan. - Clone the repository and run
cd examples/tickets && leviathan index. - Try
leviathan search -g acme "sso login loop after password reset"and look at the output size. - Export one of your own CSV or SQLite tables and index it the same way.
- Count tokens of a Leviathan result versus a grep result for the same question.
Small angles to try
- Wire it into Claude Code as a shell tool and compare task cost.
- Test it on Japanese text records to see how ranking holds up.
-
Two year Khanmigo trial: small math gains, engagement is the limit
Philip Oreopoulos and Nina Low ran a two year cluster randomized trial in 18 Tennessee middle schools, putting Khan Academy with the Khanmigo AI tutor into daily remedial math sessions. Students gained about 1.3 national percentile ranks per term, roughly 0.06 to 0.08 standard deviations over a school year. Almost every student tried it, but the median student messaged the tutor on only about a third of practice days.
Heat · HN front page #8, 54 points and 39 commentsWhat people foundThe authors conclude the binding constraint is engagement, not the model, a useful check on claims that putting a strong LLM in front of students will change outcomes on its own.Learn it15 min · 4 steps
Key ideas
- Cluster randomized trial
- Randomizing whole groups such as classes or schools instead of individual students.
- Effect size in standard deviations
- A common scale for learning gains; 0.06 to 0.08 for a full year is small but real.
- Intent to treat
- Measuring the effect of being assigned the tool, whether or not students actually used it much.
Steps
- Read the EdWorkingPapers abstract, then the tables on usage and achievement.
- Check how usage is measured and whether heavy users gained more.
- Compare with human tutoring effect sizes, which are usually much larger.
- Self check: if engagement is the bottleneck, what product change would you test first?
Try it30 min · 5 steps
You need: A browser and any chat model you already use. No API key needed.Steps
- Pick three middle school math problems you can check by hand.
- Ask a model to tutor you through each without giving the answer.
- Count how many turns pass before you feel like quitting, and note why.
- Change the instructions to make the tutor shorter and more encouraging, then repeat.
- Write down which version you would actually open again tomorrow.
Small angles to try
- Run the same session in Japanese and compare.
- Add a daily streak or quiz format and see if it changes your motivation.
Sources:EdWorkingPapers study -
HyperFrames Studio: a browser video editor built for agents
HeyGen's HyperFrames is an Apache 2.0 framework that renders HTML, CSS, media and animation into deterministic MP4 files, so coding agents can write videos as web pages. The team is now promoting Studio, a browser composition editor for previewing and editing those videos, which the repository lists as available and still evolving.
Heat · X: HyperFrames Studio post 3,469 likes and about 1.1M views · GitHub about 57.5k starsWhat people foundThe README ships a Claude Code plugin and a skills package, so the intended workflow is an agent writing the HTML timeline and a human fixing it in Studio.Learn it15 min · 4 steps
Key ideas
- Deterministic rendering
- The same HTML input always produces exactly the same frames, which makes agent output reproducible.
- Composition
- A timeline of layers such as text, images and clips that together make one video.
Steps
- Read the HyperFrames README on GitHub, starting with the stack overview.
- Open the Studio page and look at how a composition maps to HTML.
- Connect it to video as code tools like Remotion that you may know.
- Self check: why is HTML a convenient format for an LLM to produce video?
Try it30 min · 4 steps
You need: Node.js 22 or newer and FFmpeg. No API key needed for local rendering.Steps
- Run
npx hyperframes init my-videothencd my-video. - Start the preview with
npx hyperframes previewand edit the HTML. - Render with
npx hyperframes renderand check the MP4. - Optionally add the agent plugin with
claude plugin marketplace add heygen-com/hyperframesandclaude plugin install hyperframes@hyperframes, then ask for a 15 second clip.
Small angles to try
- Have two models write the same video spec and compare the results.
- Make a short explainer with Japanese subtitles for a model launch.