THINGS IN THE CONVERSATION

Mentions

Products, websites and items discussed in our sources. Read who supports, questions or simply mentions them, with the original evidence.

Showing 35 of 35

Tools & products

ABBEL

1 sources · 1 mentions

Supports 1

The authors support ABBEL, reporting that its general reconstruction-based belief grader reduces the performance gap with full-context models by about 50% and significantly lowers memory usage.

Jakob Bjorner, Aly Lidayan, Satvik Golechha, Kartik Goyal, Alane Suhr ·

Read views & evidence

Websites

AlphaFold Database

1 sources · 1 mentions

Mention only 1

Anthony Costa states that NVIDIA and partners are openly releasing predicted viral protein structures via the AlphaFold Database, describing it factually as the distribution channel.

Anthony Costa ·

Read views & evidence

Tools & products

Amazon S3

1 sources · 1 mentions

Mention only 1

Simon Willison observes that while Amazon S3 prices used to drop frequently, there has been no price reduction in the last decade.

Simon Willison ·

Read views & evidence

Tools & products

Azure Search

1 sources · 1 mentions

Mention only 1

Cohere Team uses Azure Search as a baseline comparison for Compass on the High Finance benchmark, reporting its score (64.8) without evaluative judgment beyond relative performance.

Cohere Team ·

Read views & evidence

Tools & products

Baby AGI

1 sources · 1 mentions

Opposes 1

Demetrios Brinkmann recalls Baby AGI getting stuck in recursive loops and generating high API costs, leading him to be skeptical of agents at the time.

Demetrios Brinkmann ·

Read views & evidence

Tools & products

BioNeMo Structure Prediction Pipeline

1 sources · 1 mentions

Mention only 1

Anthony Costa states that NVIDIA is openly releasing the GPU-accelerated BioNeMo Structure Prediction Pipeline for researchers to predict 3D protein structures, without evaluative judgment.

Anthony Costa ·

Read views & evidence

Tools & products

Claude

1 sources · 1 mentions

Mention only 1

Chris Olah notes that forcing specific deception-related features active causes Claude to exhibit lying behavior, illustrating interpretability findings.

Chris Olah ·

Read views & evidence

Tools & products

Claude Sonnet 5.5

1 sources · 4 mentions

Supports 4

Gabriel Grinberg of Base44 reports that Claude Sonnet 5.5 matched Opus 5’s scores across 118 app builds, achieving them in fewer iterations (3.6 vs. 7.7) and with fewer failed tool calls and mid-build interruptions.

Gabriel Grinberg ·

Read views & evidence

Websites

Cloudflare

1 sources · 1 mentions

Mention only 1

Simon Willison states he has wanted Cloudflare to support the HTTP Vary header for years.

Simon Willison ·

Read views & evidence

Tools & products

Codex

1 sources · 1 mentions

Mention only 1

Mollick describes Codex as the tool used by OpenAI to pass best ideas between agent groups during a swarm experiment.

Ethan Mollick ·

Read views & evidence

Tools & products

Compass

1 sources · 1 mentions

Supports 1

Cohere Team reports that Compass achieved a 14–16 point accuracy improvement over Azure Search on the High Finance benchmark, calling the gap decisive for end-user answer quality.

Cohere Team ·

Read views & evidence

Tools & products

Firecracker

1 sources · 1 mentions

Mention only 1

Olivia Johnston mentions Firecracker as a technology that solved isolation for an outdated unit of trust, without expressing support or opposition.

Olivia Johnston ·

Read views & evidence

Tools & products

Gemini

1 sources · 1 mentions

Mention only 1

Gemini is mentioned as an example model within Google DeepMind and Isomorphic Labs' four-step safety process, with no explicit evaluative judgment about the model itself in this excerpt.

Google DeepMind, Isomorphic Labs ·

Read views & evidence

Tools & products

gemini-3.8-flash-tts

1 sources · 1 mentions

Mention only 1

Google released the gemini-3.8-flash-tts model as one of two new Gemini text-to-speech models.

Simon Willison ·

Read views & evidence

Tools & products

Ghostwriter

1 sources · 2 mentions

Supports 2

The authors describe Ghostwriter as having been transformed from a prompted tool into a proactive, asynchronous, creative teammate embedded in Slack and Teams.

Bret Taylor, Clay Bavor ·

Read views & evidence

Tools & products

GitHub Copilot

1 sources · 2 mentions

Mention only 2

The text provides instructional guidance on how to write prompts for GitHub Copilot, describing it as a tool that accepts natural language descriptions.

Kayla Cinnamon ·

Read views & evidence

Tools & products

GPT-6 Astra Ultrafast

1 sources · 1 mentions

Mention only 1

The excerpt states GPT-6 Astra Ultrafast, running on NVIDIA Blackwell GPUs, is available in the OpenAI API and to eligible ChatGPT Work and Codex users.

Dion Harris ·

Read views & evidence

Tools & products

gVisor

1 sources · 1 mentions

Mention only 1

Olivia Johnston mentions gVisor as a technology that solved isolation for an outdated unit of trust, without expressing support or opposition.

Olivia Johnston ·

Read views & evidence

Tools & products

Harvey

1 sources · 3 mentions

Supports 3

The Harvey Team notes that Harvey Review Tables are especially effective for side-by-side comparison of numeric indemnification terms (e.g., caps, thresholds, survival periods), pulling discrete data points across agreements.

Harvey Team ·

Read views & evidence

Tools & products

Holo4

1 sources · 1 mentions

Supports 1

The authors support Holo4, presenting it as a new series of agentic models available in two sizes (27B dense and 35B-A3B MoE) via the H Models API.

Maxime Theillard, Frederic Renard, Vincent Coyette, Emrick Sinitambirivoutin, Avshalom Manevich, Antonio Loison, Antoine Bonnet, Maxime Langevin, Aleix Cambray (H-AI), Léonard Benedetti, Tony Wu, Mats L. Richter, Michael Eickenberg, Sławek Mucha, Matthias Brunel, Daniel Beechey ·

Read views & evidence

Tools & products

KernelBench

1 sources · 1 mentions

Mention only 1

Lilian Weng presents KernelBench as a benchmark of 250 PyTorch tasks assessing LLMs’ ability to generate correct and fast GPU kernels, using fast_p as its primary metric.

Lilian Weng ·

Read views & evidence

Tools & products

MLE-bench

1 sources · 1 mentions

Mention only 1

Lilian Weng describes MLE-bench as a benchmark of 75 curated Kaggle competitions used to evaluate ML engineering agents, with Kaggle public leaderboards serving as human baselines.

Lilian Weng ·

Read views & evidence

Tools & products

MLX

1 sources · 1 mentions

Mention only 1

The authors report performance results for their method on MLX kernels (Attention and Mamba SSM) on Apple Silicon, treating MLX as the target ecosystem—not as a tool under critique or endorsement.

Shiyi Cao, Gal Bloch, Assaf Toledo, Michael Factor, Gil Vernik, Joseph E. Gonzalez ·

Read views & evidence

Tools & products

Node.js

1 sources · 1 mentions

Mention only 1

Levels mentions adding 'Learn Node.js' to his to-do list because it was considered a better language than PHP, but states he never actually learned it due to startup demands.

Pieter Levels ·

Read views & evidence

Websites

Open TTS Leaderboard

1 sources · 1 mentions

Mention only 1

The authors clarify that the Open TTS Leaderboard uses objective metrics like ASR-based WER and speaker similarity as proxies—not direct measures—of intelligibility and voice identity preservation, and does not replace human preference ranking.

Eric Bezzam, Steven Zheng, Eustache Le Bihan, mrfakename ·

Read views & evidence

Tools & products

OpenAI API

1 sources · 1 mentions

Mention only 1

The OpenAI API is mentioned as the source of an $80 bill resulting from Baby AGI's recursive loop errors.

Demetrios Brinkmann ·

Read views & evidence

Tools & products

PHP

1 sources · 1 mentions

Mention only 1

Levels attributes his continued use of PHP to familiarity and lack of time to switch, noting he never learned Node.js despite considering it better.

Pieter Levels ·

Read views & evidence

Tools & products

Qwen3.5 4B

1 sources · 1 mentions

Mention only 1

El Mghari states that Qwen3.5 4B was used as the base model for their new Jev-like classifier on Together's platform.

Hassan El Mghari ·

Read views & evidence

Tools & products

RE-Bench

1 sources · 1 mentions

Mention only 1

Lilian Weng introduces RE-Bench as a benchmark evaluating frontier AI agents across 7 open-ended ML research-engineering environments, with human expert performance data included for comparison.

Lilian Weng ·

Read views & evidence

Tools & products

SAP

1 sources · 1 mentions

Mention only 1

Benedict Evans lists SAP as a representative example of large horizontal enterprise software systems used by major U.S. companies.

Benedict Evans ·

Read views & evidence

Tools & products

Seedance 2.0

1 sources · 1 mentions

Mention only 1

The author describes Seedance 2.0 as a video model that accepts multiple modal inputs—up to 9 images, 3 video clips, 3 audio files, and a text prompt—and assigns distinct creative roles to each input type.

shridharathi ·

Read views & evidence

Tools & products

Terminal-Bench-2

1 sources · 1 mentions

Mention only 1

Lilian Weng reports that a harness evolved on Terminal-Bench-2 transferred successfully to SWE-bench-verified without further evolution, suggesting it encodes general engineering experience rather than benchmark-specific optimization.

Lilian Weng ·

Read views & evidence

Tools & products

Workday

1 sources · 1 mentions

Mention only 1

Benedict Evans lists Workday as a representative example of large horizontal enterprise software systems used by major U.S. companies.

Benedict Evans ·

Read views & evidence

Views belong to a particular speaker and passage. Counts describe this reviewed collection, not product ratings or market consensus.