Thursday, October 1, 2026
-
Gemini 4 Argon: Google's coding and cyber defense model, defenders first
Google DeepMind announced Gemini 4 Argon on September 30, aimed at real world coding, enterprise knowledge work and cyber defense, with outputs up to 1M tokens. Access starts with vetted security teams in the Fairwind Program, then paid API customers and Google AI Ultra subscribers. Launch pricing is $2 per 1M input tokens and $10 per 1M output tokens, later rising to $4 and $20.
Heat · HN #1, about 1,140 points · X: Sundar Pichai's post 22K likes, 2.8M viewsWhat people foundGoogle's own table puts Argon at 77.9% on DeepSWE v1.1, and VentureBeat lists GPT-6 Astra at 74.1% and Claude Opus 5.5 at 74.2% on the same test. Treat these as vendor numbers until outside testers can reach the model.Learn it15 min · 5 steps
Key ideas
- Staged release
- A lab gives a strong model to a small trusted group first, here security defenders, before opening it to everyone.
- Output token limit
- The longest answer the model can write in one reply; Argon's 1M is far above the 64K common before.
- Cached input discount
- Repeated prompt prefixes are billed at a much lower rate, here 95% off, which matters for agents that resend the same context.
Steps
- Read Google's Argon post on blog.google and note the four benchmarks it chose to show.
- Read the VentureBeat article and compare the same numbers for GPT-6 Astra and Opus 5.5.
- Look up what DeepSWE and AutomationBench measure, so you know which kind of work each score reflects.
- Connect it to past launches: compare how this rollout differs from a normal public API release.
- Check yourself: why would a lab release a cyber capable model to defenders before attackers can buy it?
Try it30 min · 5 steps
You need: A spreadsheet and the published benchmark tables. No API key needed for the main test.Steps
- Copy Argon's published scores and the VentureBeat comparison numbers into a sheet.
- Compute the cost of a typical agent session of 2M input and 200K output tokens at intro and standard prices.
- Repeat the cost math for the models you use today with their public prices.
- Add a column for the cached input discount and see how much a long agent loop saves.
- Write down which benchmark gap is large enough to matter for your own work.
Small angles to try
- Once you get access, rerun one of your own coding tasks on Argon and log time and tokens.
- Plot price per point of DeepSWE for each model.
- Track how long the defenders only stage lasts and note the date it opens.
-
CONTROL Resonant and #RTXON: when most frames on screen are AI generated
Remedy's CONTROL Resonant launched September 29 and NVIDIA is running custom RTX 5080 giveaways with #RTXON, which became one of the most shared posts on X today. On PC the game uses full path tracing, DLSS 4.5 Ray Reconstruction and Dynamic Multi Frame Generation up to 6X. NVIDIA quotes up to 380 fps at 1440p on an RTX 5090 with 6X frame generation turned on.
Heat · X: NVIDIA giveaway post 33K likes and 33K repostsWhat people foundAt 6X frame generation, five of every six frames are predicted by a neural network rather than rendered, so a big fps number does not mean the game reacts faster to your input. NVIDIA pairs it with Reflex to cut latency.Learn it15 min · 4 steps
Key ideas
- Path tracing
- Simulating many light bounces per pixel for realistic lighting, very heavy for the GPU.
- Frame generation
- A model creates in between frames from rendered ones and motion data, raising fps without rendering them.
- Input latency
- Time from pressing a key to seeing the result; generated frames do not shorten it.
Steps
- Read NVIDIA's CONTROL Resonant post and list each feature it turns on.
- Find the fps table and note which numbers include frame generation.
- Compare it to video interpolation on a TV, which smooths motion the same way but adds delay.
- Check yourself: if a game shows 320 fps with 6X generation, about how many frames per second does the GPU really render?
Try it30 min · 5 steps
You need: Any recent GPU and a game or demo with a frame generation toggle; or Python with OpenCV for the offline version. No API key needed.Steps
- Pick a short gameplay or screen recording at 30 fps that you made yourself.
- Write a small Python script that builds an in between frame by blending two frames.
- Write a second version that uses optical flow from OpenCV to warp one frame toward the next.
- Compare both against the real middle frame from a 60 fps recording using PSNR or SSIM.
- Note where artifacts show up: text, fast edges, UI overlays.
Small angles to try
- If you own a supported GPU, measure fps and felt latency with frame generation off, 2X and 4X.
- Try a learned interpolation model such as RIFE and compare quality.
- Measure how long each method takes per frame on CPU versus GPU.
-
Microsoft Titan accepted unsigned JWTs: admin access to 17 analytics DBs
A researcher found that Titan, an internal Microsoft analytics API, read the claims inside a JWT login token but never checked its signature, so a token marked with the none algorithm could claim to be the admin user. That would have exposed an estimated 17.3 trillion rows across 17 databases. It was reported September 5, locked down September 9, and paid a $5,000 bounty.
Heat · HN front page, about 270 points and 116 commentsWhat people foundHN commenters argue the JWT none algorithm is a design trap that keeps producing this bug, and many criticised the small bounty and Microsoft asking for editorial input on the writeup.Learn it15 min · 4 steps
Key ideas
- JWT
- A signed token that carries claims like user and audience; the signature proves who issued it.
- alg none
- A JWT header value meaning no signature at all; a server must reject it.
- Verify vs decode
- Decoding only reads claims; verifying checks the signature with a key and an allowed algorithm list.
Steps
- Read the researcher's writeup and find the exact point where the server trusted the claims.
- Read the JWT section of the OWASP cheat sheet on JSON Web Tokens about algorithm checks.
- Compare it to session cookies: what stops a user from editing their own cookie?
- Check yourself: why is passing a fixed list of allowed algorithms to your verify call the real fix?
Try it30 min · 5 steps
You need: Python 3 with a JWT library such as PyJWT, on your own laptop. No API key needed.Steps
- Write a tiny local Flask or FastAPI app with one endpoint that returns the user from a bearer token.
- Version A only decodes the token without verification; version B verifies with a secret and a fixed algorithm list.
- Create a token signed with your own secret and confirm both versions accept it.
- Create an unsigned token for your own local app and confirm version A accepts it and version B rejects it.
- Add a unit test that fails if anyone changes B back to decoding only.
Small angles to try
- Repeat in Node with the jsonwebtoken package and compare its defaults.
- Search your own repos for decode calls that skip verification.
- Add a check that the audience and issuer claims match too.
-
Meta's Muse hits #1 on the US App Store amid a Messages sync report
Muse from Meta, an agent app that browses, fills forms and buys things with approval, is #1 in Apple's US top free chart, ahead of ChatGPT and Gemini. AppleInsider, citing Inc.'s Jason Aten, reports that Muse on a Mac synced about 187,000 lines from his Messages database even though he says he had declined that access. A Meta spokesperson said on X that the Messages integration is opt in and needs both Full Disk Access and the connector enabled.
Heat · US App Store top free #1 · X: Unusual Whales post on the report 7.1K likesWhat people foundThe dispute is over what the permission screen showed versus what the app did, which is the core trust question for any agent with desktop access. The claims are a single user's account so far, and Meta disputes them.Learn it15 min · 4 steps
Key ideas
- Full Disk Access
- A macOS permission that lets an app read protected files like the Messages database.
- Connector
- An agent's link to one data source; it can be on or off separately from OS permissions.
- Least privilege
- Give a program only the access it needs, so mistakes or bugs cannot reach more.
Steps
- Read the AppleInsider article and list the exact claims and who made them.
- Read Apple's support page about privacy settings on Mac to see how Full Disk Access works.
- Compare it to a phone app asking for contacts: what can the OS enforce and what does it trust the app to do?
- Check yourself: which logs would prove whether data was read, and who controls them?
Try it30 min · 5 steps
You need: A Mac you own and a test user account. No API key needed.Steps
- Create a fresh local macOS user with no real messages or data.
- Open System Settings, Privacy and Security, and record which apps have Full Disk Access.
- Install any agent app you are evaluating in that test user only, and note each permission prompt.
- Compare what the app's own settings say it can access with what the OS lists.
- Remove the app and confirm its permissions are cleared.
Small angles to try
- Use the macOS unified log to watch file access events while the app runs.
- Repeat with a second agent app and compare prompts.
- Write a checklist you would use before giving any agent desktop access.
-
ChatGPT plugins grow into apps, and Sign in with ChatGPT shares your plan
After DevDay, OpenAI expanded ChatGPT plugins with their own sidebar, interactive panels, a Plugin Creator and support for the proposed MCP Events spec. A plugin bundles skills, MCP servers and optional UI and publishes once to a directory used by ChatGPT and Codex. Sign in with ChatGPT now lets Plus and Pro users spend their plan allowance inside partner apps such as Notion, Warp, Devin and Vercel.
Heat · X: Tibo of OpenAI 8.2K likes on the platform post, Notion 2.6K likesWhat people foundOpenAI's help page says apps get only name, email and profile picture, and subscription sharing gives no access to your chats or memory. For tool makers, the subscription route may replace API billing for many users.Learn it15 min · 4 steps
Key ideas
- MCP server
- A small service that exposes tools and data to a model through the Model Context Protocol.
- Plugin bundle
- OpenAI's package of skills, MCP servers and optional UI that installs as one unit.
- Plan allowance
- Usage included in a subscription; partner apps can now draw from it instead of an API key.
Steps
- Read OpenAI's plugins page on developers.openai.com and draw the three parts of a plugin.
- Read the Sign in with ChatGPT help article and list what data is shared and what is not.
- Compare it to Sign in with Google: identity only versus identity plus paid usage.
- Check yourself: what changes for a small app maker if users pay for AI through their ChatGPT plan?
Try it30 min · 5 steps
You need: A ChatGPT account and a small MCP server you write locally; Node or Python. No OpenAI API key needed for the local part.Steps
- Write a minimal MCP server with one tool, for example a unit converter, following the official MCP quickstart.
- Test it locally with the MCP Inspector.
- Follow OpenAI's plugins docs to connect it to ChatGPT in developer mode and call the tool.
- Note what ChatGPT shows the user before the tool runs.
- Write down what you would need to add before submitting it to the directory.
Small angles to try
- Connect the same server to Codex and compare behavior.
- Add a tiny UI panel and see how it renders.
- Try Sign in with ChatGPT in one partner app and check what it asks for.
-
DoorDash Air drones and order by text: how dispatch picks drone or Dasher
DoorDash showed its own six propeller drone, DoorDash Air, on September 30, piloting in Northern California with Chipotle, Popeyes and others, with flights averaging under 5 minutes. The same day it opened a US beta where you text DoorDash something like order my usual and an AI agent places the order. A dispatch platform chooses a Dasher, ground robot or drone per order using distance, weight, weather and traffic.
Heat · X: Polymarket post on Chipotle by drone 4.5K likes, 545K viewsWhat people foundDoorDash told Axios the drone payload was sized from real order data to carry about 80% of US restaurant orders, a nice example of picking a hardware spec from a data distribution.Learn it15 min · 4 steps
Key ideas
- Dispatch optimization
- Assigning each job to a worker or vehicle to minimise time or cost under constraints.
- Payload sizing by percentile
- Choosing a capacity that covers a target share of real orders, here about 80%.
- Agentic ordering
- An AI turns a casual text into a concrete order using past purchases.
Steps
- Read the Axios piece and list the inputs DoorDash says the dispatcher uses.
- Read Dexerto on the text ordering beta and note how it uses order history.
- Connect it to a ride hailing app matching drivers to riders.
- Check yourself: why would a drone be the wrong choice for a big family order in the rain?
Try it30 min · 5 steps
You need: Python with pandas; synthetic or your own data. No API key needed.Steps
- Generate 1,000 fake orders with weight, distance and a weather flag.
- Find the payload weight that covers 80% of orders, then 90%, and compare.
- Write simple rules that send each order to drone, robot or courier.
- Compute average delivery time and share of orders per mode.
- Change the weather rate and see how the mix shifts.
Small angles to try
- Replace rules with a small linear program using PuLP or OR-Tools.
- Add a battery range limit for drones.
- Use your own past delivery receipts as the order distribution.
-
Magnitude: an open source local inference engine that tunes kernels to your GPU
Magnitude, a YC startup, launched an Apache 2.0 inference engine written in Rust with custom GPU kernels that it tunes on your own machine. It frees memory when agents stop and uses hybrid paged attention so parallel agent sessions share a prefix cache. It ships as a desktop app for macOS, Linux and Windows and connects to agents such as Codex, Claude Code and OpenCode.
Heat · Launch HN on the front page, about 140 points · GitHub about 4.7K starsWhat people foundThe founders report up to 92% faster decoding than llama.cpp on an M4 Pro and 19% on a DGX Spark with Qwen 3.6 35B at 4 bit, using about 27% less memory. HN commenters asked for MLX comparisons and details on how the benchmark was run.Learn it15 min · 4 steps
Key ideas
- Kernel autotuning
- Trying several GPU code variants on your hardware and keeping the fastest.
- Paged attention
- Storing the attention cache in pages so many sessions can share and grow it efficiently.
- Prefill vs decode
- Prefill reads the prompt in parallel; decode writes one token at a time and is usually memory bound.
Steps
- Read the Magnitude README on GitHub and the Launch HN post.
- Read the vLLM paper summary on paged attention to understand the cache idea.
- Compare with llama.cpp, which you may already run, and note what it does not tune.
- Check yourself: why does sharing a prefix cache help when several agents use the same system prompt?
Try it30 min · 5 steps
You need: A Mac with Apple silicon or a PC with an NVIDIA or AMD GPU, 16 GB or more memory. No API key needed.Steps
- Install the desktop app from magnitude.dev/download as the README describes.
- Pick one model in the Discover tab that also runs in llama.cpp on your machine.
- Run the same 2,000 token prompt with a 300 token answer in both and record prefill and decode tokens per second.
- Record memory use in each run.
- Run three sessions at once with the same system prompt and compare throughput.
Small angles to try
- Add MLX or Ollama as a third engine.
- Try a smaller and a larger model size.
- Connect a coding agent and time one real task.
-
栗原40号 trends: Kurihara hits 40 HR, and what Hawk-Eye data shows fans
SoftBank's Ryoya Kurihara reaching 40 home runs trended on X Japan today; no Hawks hitter had reached 40 since 2005, and he hit No. 39 on September 23. Since this season NPB's official app NPB+ shows Sony Hawk-Eye tracking for every pitch, including exit velocity, swing speed and barrel rate.
Heat · X Japan trending #3 · Nishi Nippon and Daily Sports covered the chase to 40What people foundDaily Sports quotes Kurihara saying he became very conscious of 40 after the title was clinched. For data fans, NPB+ now shows the exit velocity of each of his home runs, pitch by pitch.Learn it15 min · 4 steps
Key ideas
- Exit velocity
- How fast the ball leaves the bat, a key predictor of home runs.
- Barrel
- A batted ball with exit speed and launch angle in the range that most often becomes extra base hits.
- Hawk-Eye
- Sony's camera system that tracks ball and players in 3D at high frame rate.
Steps
- Read the Daily Sports article on Kurihara's chase to 40.
- Read the k-tai Watch article on NPB+ to see which Hawk-Eye stats the app shows.
- Connect it to MLB Statcast, which uses the same idea and has public data.
- Check yourself: why can two home runs with the same exit speed travel different distances?
Try it30 min · 5 steps
You need: Python with pandas and matplotlib; public MLB Statcast data via the pybaseball package. No API key needed.Steps
- Install pybaseball and pull one season of batted ball data for a power hitter.
- Plot launch angle against exit velocity and colour home runs.
- Draw the region where most home runs land and count how many balls fall inside it.
- Compare the hitter's home run rate in two parks.
- Write one sentence on what made his best month different.
Small angles to try
- Use NPB+ app screenshots of Kurihara's at bats and log them by hand.
- Add a simple logistic model of home run probability.
- Compare a dome park and an open air park.
-
Netlify moves Edge Functions from V8 isolates to Firecracker microVMs
Netlify now runs every Edge Function in its own Firecracker microVM instead of a V8 isolate. It reports warm calls at about 5 to 6 ms median, down from 25 to 40 ms, P99 47.4% faster, and cold starts averaging 9 ms on about 1.2% of requests. Existing functions keep working with no config change.
Heat · HN front page, about 140 points and 53 commentsWhat people foundNetlify's engineers argue that V8 isolates cannot give the same isolation as a VM, which is a direct jab at the isolate model used by other edge platforms. HN debated whether that security claim holds.Learn it15 min · 4 steps
Key ideas
- V8 isolate
- A separate JavaScript heap inside one process; very cheap but shares the process and kernel.
- microVM
- A tiny virtual machine with a minimal device model that boots in milliseconds.
- Snapshot restore
- Starting a VM from a saved memory image instead of booting from scratch.
Steps
- Read the Netlify post and list each latency number with its percentile.
- Read the Firecracker README on GitHub to see what it leaves out compared to a normal VM.
- Compare it to containers you already use: which share the kernel and which do not.
- Check yourself: why can a VM based platform now be faster than isolates for warm calls?
Try it30 min · 5 steps
You need: A Netlify account on the free tier, Node, and the Netlify CLI; or any HTTP benchmarking tool. No API key needed beyond your own Netlify login.Steps
- Deploy a tiny edge function that returns a fixed JSON body to your own site.
- Send 1,000 sequential requests from your machine and record p50 and p99 latency.
- Wait 30 minutes, send one request, and note the cold start time.
- Deploy the same function on another edge platform you use and repeat.
- Subtract the network round trip measured with a static file to estimate function time.
Small angles to try
- Test from Tokyo and from a US VPS.
- Add an npm dependency and see if latency changes.
- Measure the response header timing data if the platform exposes it.
-
Matthew Green: sandboxes alone will not stop an agent worm
Cryptographer Matthew Green argues that sandboxing AI agents is not enough, because agents still read email, documents, messages and shared package caches written by others. A hostile instruction can hijack one agent, which then carries the same payload to the next agent it writes to. Simon Willison highlighted the post on October 1.
Heat · Featured on simonwillison.net todayWhat people foundGreen's framing: a payload that hijacks an agent plus an agent that delivers it onward are the two halves of a worm, so the defense has to cover the data channels, not only the process boundary.Learn it15 min · 4 steps
Key ideas
- Prompt injection
- Text in data that the model treats as instructions.
- Worm
- Malware that spreads by itself from one system to the next.
- Data channel
- Any input an agent reads, such as email, docs or packages, which can carry hostile text.
Steps
- Read Simon Willison's note and follow it to Green's full post.
- Read Simon's earlier posts on the lethal trifecta for agent security.
- Connect it to email worms of the early 2000s that used address books to spread.
- Check yourself: name one channel in your own agent setup where outside text reaches the model.
Try it30 min · 5 steps
You need: Two local agent scripts using any LLM you have, or a local model with Ollama, and a shared folder. Everything stays on your machine.Steps
- Build agent A that summarises text files in a folder and writes notes to a second folder.
- Build agent B that reads those notes and drafts replies.
- Put a harmless test instruction in one input file, such as asking the agent to add the word BANANA to every output.
- Check whether the marker passes from A's notes into B's drafts.
- Add one defense, such as wrapping inputs as quoted data or a filter step, and measure how often the marker still passes.
Small angles to try
- Try a smaller and a larger model.
- Add a third hop and see if the marker survives.
- Log every place the marker appeared.
Sources:Simon Willison on Matthew Green -
Ideogram 4.5 targets multi turn image edits that do not drift
Ideogram released 4.5, an editing model built to stop the colour shifts, pixel drift and texture artifacts that pile up when you edit the same image many times. It is live in Ideogram, its API and partners such as fal, Krea, Runway and ComfyUI. Open weights were promised but not released at launch.
Heat · X: Ideogram launch post 3K likes, 378K viewsWhat people foundNo independent benchmark yet; the comparisons against GPT Image and Nano Banana models are Ideogram's own demos, so drift over ten or twenty edits is worth measuring yourself.Learn it15 min · 4 steps
Key ideas
- Edit drift
- Small unwanted changes that add up across repeated edits.
- Masked edit
- Changing only a chosen region and keeping the rest identical.
- SSIM
- A score from 0 to 1 for how similar two images look to people.
Steps
- Read Ideogram's 4.5 model page and note what it says is preserved.
- Read RuntimeWire's article for the API endpoints and quality modes.
- Compare with photo editing layers, where untouched pixels never change.
- Check yourself: how would you prove that unedited regions stayed identical?
Try it30 min · 5 steps
You need: An Ideogram account or partner access, Python with Pillow and scikit-image. Uses paid credits.Steps
- Pick one photo you own and plan ten small edits, each on a different region.
- Run the ten edits one after another in Ideogram 4.5, saving each result.
- For each step, compute SSIM between the original and the result outside the edited region.
- Repeat with another editor you use.
- Plot SSIM against edit number for both.
Small angles to try
- Measure average colour shift per step.
- Try text inside the image, which drifts easily.
- Repeat at 2K output.
-
DeepSeek Harness desktop preview for macOS and Windows
DeepSeek's open source agent platform, DeepSeek Harness, now has a desktop app in public preview for Apple silicon Macs and 64 bit Windows. Harness is MIT licensed and built on a plugin first design. Pandaily describes the builds as Electron based preview releases rather than general availability.
Heat · X: DeepSeek Harness post 4K likes, 357K viewsWhat people foundMany unofficial lookalike sites and GitHub wrappers already use the name, so download only from deepseek.com; that is the main practical warning today.Learn it15 min · 4 steps
Key ideas
- Agent harness
- The loop and tools around a model that let it plan, call tools and edit files.
- Plugin architecture
- A core that does little, with features added as separate modules.
- Electron
- A framework that ships a web app with its own Chromium as a desktop app.
Steps
- Read DeepSeek's Harness page on deepseek.com and list what is a plugin.
- Read the Pandaily piece to see what the preview builds include.
- Compare it with Codex CLI or Claude Code, which you may already use.
- Check yourself: how do you confirm a download really comes from DeepSeek?
Try it30 min · 5 steps
You need: Node.js with npx, a DeepSeek API key or a supported model provider, and a throwaway project folder.Steps
- Clone the official repo with
git clone https://github.com/deepseek-ai/deepseek-harnessor start the web UI withnpx @deepseek-ai/dsh web, as listed on the official page. - Open a small test repo and ask it to add one unit test.
- Record time, tokens and whether the test passes.
- Give the same task to another agent you use and compare.
- Check which files each agent touched.
Small angles to try
- Try the desktop app and compare it to the web UI.
- Write a tiny plugin that adds one tool.
- Run it against a local model provider.
-
Mathematicians ask AI labs to hold back on famous open problems
A group of mathematicians published a call for responsible release of AI generated mathematics, asking frontier labs not to use closed models to clear long standing open problems. They warn of a two tier field where labs outrun everyone else, and ask labs to fund mathematicians to interpret the results.
Heat · HN front page, 83 points and 104 commentsWhat people foundHN is split: critics call it gatekeeping, while supporters say outsiders cannot check results from closed models and students lose the problems that train them.Learn it15 min · 4 steps
Key ideas
- Open problem
- A well known question with no accepted proof yet.
- Formal proof
- A proof checked line by line by software such as Lean.
- Reproducibility
- Others can rerun the method and get the same result, hard with closed models.
Steps
- Read the statement on agmai.org and list its concrete requests.
- Skim the HN thread and pick the strongest argument on each side.
- Connect it to open source vs closed code: who can verify a claim.
- Check yourself: would a Lean checked proof answer the main worry, or not?
Try it30 min · 5 steps
You need: Lean 4 with the official VS Code extension, on your own machine. No API key needed.Steps
- Install Lean 4 following the official Lean website instructions.
- Prove a small statement, such as the sum of two even numbers is even.
- Ask an LLM you use for a Lean proof of the same statement.
- Check whether it compiles and fix it by hand.
- Note what a human still had to understand.
Small angles to try
- Try a harder exercise from Mathematics in Lean.
- Compare two models on the same proof.
- Time how long checking takes vs writing.
-
RIDE: distill from the direction RL moved the teacher, not its outputs
A new paper proposes distilling a reinforcement learning tuned teacher by extrapolating how RL changed its hidden states compared with the teacher before RL. Instead of copying output distributions, the student follows that representation shift. The authors report matching or beating the teacher on four base and teacher pairs.
Heat · Hugging Face daily papers #1 for Oct 1What people foundThe interesting claim is that the RL update is a direction you can push further, so a student can end up past the teacher rather than just near it. It is the authors' result and needs replication.Learn it15 min · 4 steps
Key ideas
- Distillation
- Training a smaller student model to imitate a stronger teacher.
- On policy
- The student learns from its own generated samples rather than fixed data.
- Representation residual
- The difference in hidden states between two versions of a model.
Steps
- Read the abstract and figure 1 on the arXiv page.
- Read a short intro to knowledge distillation, such as Hinton's 2015 paper summary.
- Connect it to task vectors: weight differences you can add or scale.
- Check yourself: what would extrapolating too far likely break?
Try it30 min · 5 steps
You need: Python with PyTorch and transformers, one small open model in two versions such as base and instruct. A laptop GPU or Colab is enough.Steps
- Load a small base model and its instruct version.
- Run the same 50 prompts through both and save one middle layer's hidden states.
- Compute the average difference vector between the two.
- Add that vector, scaled 0.5, 1 and 1.5, to the base model's hidden states at that layer during generation.
- Compare outputs by eye and with a simple score.
Small angles to try
- Try a different layer.
- Clone the official RIDE repo and read the training loop.
- Measure when output quality collapses as the scale grows.
-
Gemini app adds skills: saved instructions you run with a slash
Google added skills to the Gemini app, letting you save an instruction once and run it by typing a slash and its name. Gemini can also suggest skills from your chats and trigger them when a prompt matches, and skills can include reference files. They replace Gems and are rolling out to Google AI subscription tiers.
Heat · X: Google's post 3.4K likes, 797K viewsWhat people foundThe same idea now exists across Claude, Codex and Gemini: reusable instruction packs. The practical question is whether a skill written for one tool can move to another with little change.Learn it15 min · 4 steps
Key ideas
- Skill
- A saved instruction pack, sometimes with files, the assistant can reuse on demand.
- Slash command
- Typing a slash and a name to call a saved action.
- Auto trigger
- The assistant picks a skill itself when your prompt matches its description.
Steps
- Read Google's post about skills on blog.google.
- Compare it to skills in Claude or Codex if you use them.
- Connect it to shell aliases: a short name for a longer command.
- Check yourself: what makes a skill description good enough for auto triggering?
Try it30 min · 5 steps
You need: A Gemini app account on a Google AI plan. No API key needed.Steps
- Write one skill for a task you repeat, such as reviewing a PR description.
- Run it with the slash command on three different inputs.
- Write a prompt that should auto trigger it and see if it does.
- Port the same instructions to another assistant's skill format.
- Compare output quality and effort.
Small angles to try
- Add a reference file to the skill.
- Combine two skills in one prompt.
- Convert an old Gem and compare.
-
October 1 in Japan: one beer tax rate, price rises, and your own inflation tracker
From October 1, beer, happoshu and new genre beer all pay the same liquor tax of 54.25 yen per 350 ml, the last step of a plan started in 2017, so beer gets cheaper and the budget types get pricier. Trackers count over a thousand food and drink items rising in October, and major banks raised floating mortgage rates. Price trackers and tax calculators are how people are following it.
Heat · Google Trends Japan: 変動金利 1,000+ searches on Sept 30 · neage tracker lists 46 companies raising prices in OctoberWhat people foundBeer gets about 9 yen cheaper per can in tax and happoshu and new genre about 7 yen pricier, which ends the tax gap that created new genre beer in the first place. Different trackers give different item counts, so cite the source for any number.Learn it15 min · 4 steps
Key ideas
- Tax category design
- Products engineered to fit a cheaper tax bracket, like new genre beer.
- Price index
- A weighted average of price changes across a basket of goods.
- Receipt OCR
- Reading text and numbers from receipt photos automatically.
Steps
- Read the tax summary on ishihide-tax.com and note old and new rates per 350 ml.
- Look at neage tracker's October calendar and pick three items you buy.
- Connect it to the CPI: an official version of the same basket idea.
- Check yourself: why does a single rate remove the reason new genre beer existed?
Try it30 min · 5 steps
You need: Python with pandas; receipt photos you own and any OCR tool or a vision model. No API key needed for the pandas part.Steps
- Collect 10 receipts from September and the first October ones as they come.
- Extract item, price and date with OCR, fixing errors by hand.
- Match the same items across months.
- Compute your personal price change weighted by how much you spend.
- Compare with the tax change for one beer and one happoshu.
Small angles to try
- Use a vision LLM for extraction and compare accuracy to OCR.
- Add a simple chart that updates weekly.
- Compare your number with the official CPI when it is released.
-
Frigate goes viral again: a local AI camera recorder with no cloud fee
Frigate, an open source network video recorder, runs real time object detection for people, cars and animals on your own hardware and is getting a new wave of attention on X. It only runs detection when motion is detected to save compute and supports Coral and other accelerators, plus 24/7 recording and Home Assistant integration.
Heat · X: thread about Frigate 1.5K likes · GitHub about 35.9K starsWhat people foundThe design trick worth copying is cheap motion detection first, expensive neural detection only on frames that changed, which is why it runs on small boxes.Learn it15 min · 4 steps
Key ideas
- NVR
- A network video recorder that stores and indexes camera streams.
- Motion gating
- Running a cheap pixel change check before an expensive model.
- Edge accelerator
- A small chip like Coral that runs neural networks with low power.
Steps
- Read the Frigate README on GitHub and its docs on hardware.
- Find how motion detection and object detection interact in the docs.
- Connect it to a cache: cheap check first, costly work only when needed.
- Check yourself: what happens to detection on a windy day with moving trees?
Try it30 min · 5 steps
You need: Docker on a Linux box or Mac, and one camera you own, or a recorded video file. No API key needed.Steps
- Follow the Frigate docs to run it in Docker with a basic config.
- Point it at your own camera stream or a looping video of your room.
- Turn on object detection for people and record detections for an hour.
- Measure CPU use with and without motion gating settings tuned.
- Count false alarms.
Small angles to try
- Add a Coral or GPU detector and compare CPU load.
- Send events to Home Assistant over MQTT.
- Compare two detection models.