The authors support ABBEL, reporting that its general reconstruction-based belief grader reduces the performance gap with full-context models by about 50% and significantly lowers memory usage.
Jakob Bjorner, Aly Lidayan, Satvik Golechha, Kartik Goyal, Alane Suhr ·
Anthony Costa states that NVIDIA and partners are openly releasing predicted viral protein structures via the AlphaFold Database, describing it factually as the distribution channel.
Cohere Team uses Azure Search as a baseline comparison for Compass on the High Finance benchmark, reporting its score (64.8) without evaluative judgment beyond relative performance.
Demetrios Brinkmann recalls Baby AGI getting stuck in recursive loops and generating high API costs, leading him to be skeptical of agents at the time.
Anthony Costa states that NVIDIA is openly releasing the GPU-accelerated BioNeMo Structure Prediction Pipeline for researchers to predict 3D protein structures, without evaluative judgment.
Chris Olah notes that forcing specific deception-related features active causes Claude to exhibit lying behavior, illustrating interpretability findings.
Gabriel Grinberg of Base44 reports that Claude Sonnet 5.5 matched Opus 5’s scores across 118 app builds, achieving them in fewer iterations (3.6 vs. 7.7) and with fewer failed tool calls and mid-build interruptions.
Cohere Team reports that Compass achieved a 14–16 point accuracy improvement over Azure Search on the High Finance benchmark, calling the gap decisive for end-user answer quality.
Gemini is mentioned as an example model within Google DeepMind and Isomorphic Labs' four-step safety process, with no explicit evaluative judgment about the model itself in this excerpt.
The authors describe Ghostwriter as having been transformed from a prompted tool into a proactive, asynchronous, creative teammate embedded in Slack and Teams.
The text provides instructional guidance on how to write prompts for GitHub Copilot, describing it as a tool that accepts natural language descriptions.
The excerpt states GPT-6 Astra Ultrafast, running on NVIDIA Blackwell GPUs, is available in the OpenAI API and to eligible ChatGPT Work and Codex users.
The Harvey Team notes that Harvey Review Tables are especially effective for side-by-side comparison of numeric indemnification terms (e.g., caps, thresholds, survival periods), pulling discrete data points across agreements.
The authors support Holo4, presenting it as a new series of agentic models available in two sizes (27B dense and 35B-A3B MoE) via the H Models API.
Maxime Theillard, Frederic Renard, Vincent Coyette, Emrick Sinitambirivoutin, Avshalom Manevich, Antonio Loison, Antoine Bonnet, Maxime Langevin, Aleix Cambray (H-AI), Léonard Benedetti, Tony Wu, Mats L. Richter, Michael Eickenberg, Sławek Mucha, Matthias Brunel, Daniel Beechey ·
Lilian Weng presents KernelBench as a benchmark of 250 PyTorch tasks assessing LLMs’ ability to generate correct and fast GPU kernels, using fast_p as its primary metric.
Lilian Weng describes MLE-bench as a benchmark of 75 curated Kaggle competitions used to evaluate ML engineering agents, with Kaggle public leaderboards serving as human baselines.
The authors report performance results for their method on MLX kernels (Attention and Mamba SSM) on Apple Silicon, treating MLX as the target ecosystem—not as a tool under critique or endorsement.
Shiyi Cao, Gal Bloch, Assaf Toledo, Michael Factor, Gil Vernik, Joseph E. Gonzalez ·
Levels mentions adding 'Learn Node.js' to his to-do list because it was considered a better language than PHP, but states he never actually learned it due to startup demands.
The authors clarify that the Open TTS Leaderboard uses objective metrics like ASR-based WER and speaker similarity as proxies—not direct measures—of intelligibility and voice identity preservation, and does not replace human preference ranking.
Eric Bezzam, Steven Zheng, Eustache Le Bihan, mrfakename ·
Lilian Weng introduces RE-Bench as a benchmark evaluating frontier AI agents across 7 open-ended ML research-engineering environments, with human expert performance data included for comparison.
The author describes Seedance 2.0 as a video model that accepts multiple modal inputs—up to 9 images, 3 video clips, 3 audio files, and a text prompt—and assigns distinct creative roles to each input type.
Lilian Weng reports that a harness evolved on Terminal-Bench-2 transferred successfully to SWE-bench-verified without further evolution, suggesting it encodes general engineering experience rather than benchmark-specific optimization.