
How sources behind AI recommendations are manufactured
An independent analysis of 380 software categories examines how AI-integrated search systems build recommendations. It sent 760 queries—one per category to Perplexity/sonar and one to Perplexity/sonar-pro through OpenRouter—and collected 7,534 citations. 59.8% of cited domains were outside the 100,000 most visited sites, while 23.4% did not appear in Tranco’s top million. The central finding is not that a few famous sites control answers: the ten most-cited sources represented only 17.3% of citations. The striking pattern lies at the periphery, where 751 of 2,055 cited domains were unranked and had more recent Wayback captures. Some appeared designed to be consumed by retrieval systems, with automated pages and descriptions explicitly oriented toward model grounding. The report identifies three apparently related domains that published 215,128 automated “best software” pages, and highlights guideflow.com, whose blog was cited 194 times across 96 categories despite not being a directory or operating in those markets. The authors compare sitemaps, templates, DNS, registration dates, and archived pages, while noting that shared infrastructure is circumstantial evidence, not proof of common ownership. The broader lesson is that AI answers can rely on recent, high-volume content designed for machines without making its authority obvious. The study measured only Perplexity—Google was excluded—so its findings should not automatically be generalized to every search engine or model.
Meta Unveils Muse Spark 1.3: New Open-Source Language Model
Meta has announced the release of Muse Spark 1.3, a new iteration of its open-source language model designed for advanced natural language generation and understanding tasks. This version introduces significant improvements in computational efficiency and output quality, positioning itself as a competitive alternative in the accessible language model ecosystem. Muse Spark 1.3 stands out for its ability to process extended contexts and generate coherent responses in multiple languages, including enhanced support for low-resource languages. The model's architecture has been optimized to reduce carbon footprint during training and inference, aligning with Meta's AI sustainability commitments. The model is available under an open-source license that permits commercial use and modification, fostering community innovation. Researchers and developers can access the model weights and detailed documentation through Meta AI's official repository, driving the development of custom applications and academic research.
FCC plans robocall scorecard to grade phone companies on spam call blocking
The Federal Communications Commission today said it will create a robocall mitigation scorecard to rate phone companies on how effectively they block illegal spam calls. The scorecards could include call-blocking statistics along with data on customer complaints and enforcement actions. The FCC said scorecards could grade providers on a number scale, with letter grades, or by classifying providers as low risk, medium risk, or high risk. "The Scorecard will empower consumers and encourage provide
Massive Driver's License Leak on Dark Web: 153 Million IDs Exposed
An unprecedented data breach has exposed over 153 million driver’s licenses on a new cybercrime platform called Nexus, uncovered by security researcher Brian Krebs. The platform, which operated for hours before disappearing following the exposé, included high-resolution scans of documents belonging to public figures, journalists, and FBI agents. The central victim in the case, a tech journalist, discovered that his license had been scanned by a car rental company and, within hours, appeared for sale on the dark web. According to Krebs, the breach appears linked to IDScan.net, a New Orleans-based ID scanning service that collaborates with companies like Hertz and Planet13. The platform offered not only visible-light images but also infrared and ultraviolet spectrum scans, enabling the creation of nearly undetectable counterfeit IDs. Technically, Nexus operates as a marketplace for stolen identities, integrated with third-party services that process physical documents in real time. The system exploits vulnerabilities in ID scanning APIs to access sensitive data, including metadata and specialized formats. The files include multiple image layers: front, back, infrared, and ultraviolet, suggesting attackers have access to specialized hardware or internal systems of partner companies. The attack architecture indicates advanced automation, with new documents appearing on the platform within minutes or hours of being scanned. This points to near real-time access to databases or cloud services used by businesses that entrust their scans to IDScan.net. From a business and financial ecosystem perspective, Nexus represents a new paradigm in the cybercrime marketplace: specialized platforms offering access to complete identities with advanced verification tools. While no direct investors or funding details have been disclosed, the business model resembles dark web subscription services where users pay for access to updated databases. Competition in this space includes platforms like Genesis Market and XSS, but Nexus stands out for its focus on physical identity documents with cloning capabilities. The rapid emergence and disappearance of the platform suggests a highly coordinated operation, possibly backed by criminal access groups with significant technical resources. Strategically, this incident highlights the fragility of the digital identity supply chain across industries such as travel, retail, and financial services. Companies relying on third parties for document scanning must reassess their security and audit protocols. The U.S. government, including the FBI, is already investigating the case, which may lead to new regulations on handling identification data. In the long term, increased adoption of end-to-end encryption and decentralized identity verification is expected. The cybersecurity market will likely see accelerated growth in identity protection solutions, especially in sectors handling large volumes of sensitive documents.
Spending Deal Pauses Political Control of Grants Until December
The House passed a stopgap funding bill that keeps the government running through early December and includes a provision blocking political control over grant awards. President Trump is expected to sign the measure to avert a shutdown just before the midterm elections. Such short-term deals have become routine, reflecting the difficulty of passing a full-year budget.
TechCrunch Disrupt 2026’s new Real World AI Stage features Nvidia, robots, and extinct animals
TechCrunch Disrupt 2026 will once again place AI at the heart of its program, but with a new twist: a dedicated Real World AI Stage. The event, taking place October 13‑15 at San Francisco’s Moscone West, expands its AI offerings by introducing this fresh stage, which explores the intersection of digital and physical realms. It will highlight how autonomous hardware is moving beyond self‑driving cars to occupy public spaces, battlefields, homes, and even potentially aid extinct species in re‑entering Earth. The lineup features speakers from leading companies such as Shield AI, Colossal Biosciences, FieldAI, Foxglove, and Nvidia, among others still to be announced. The first session tackles the data gap that hampers general‑purpose robotic intelligence. While large language models had the internet and self‑driving cars have millions of hours of road data, robots lack that foundational data. This shortfall is the primary reason most experts believe broad robotic AI is still years away, despite breakthroughs elsewhere in AI. A new wave of startups is racing to close this gap by building data pipelines, simulation environments, and foundation models that could spark a capability explosion similar to that seen with LLMs. Les Karpas, Head of Physical AI at Nvidia, will lead a discussion on what the ChatGPT moment for physical AI truly requires and how close we really are. The second session focuses on safety and readiness when AI enters the physical world. A mistake can ground an aircraft, crash a vehicle, or jeopardize a mission. Leaders building autonomous vehicles, defense technologies, and industrial systems will address one of the toughest questions every hard‑tech founder faces: how to know when a system is safe to deploy. The panel will explore how founders can cultivate a safety culture, test and validate AI, navigate regulatory hurdles, and build trust in high‑stakes environments. Nate Michael, CTO of Shield AI, will participate in this discussion. The third session, a fireside chat, features Ben Lamm, CEO of Colossal Biosciences, a company known for turning de‑extinction from science fiction into a billion‑dollar business. Lamm will discuss the technologies used to revive extinct species, the role AI plays in modern biology, and the growing debate over whether engineering nature is a conservation breakthrough or a distraction from protecting existing wildlife. His talk highlights the ethical, technical, and commercial dimensions of attempting to bring back lost species. The fourth session brings together leaders in edge AI from defense, space, and industrial sectors. Dr. Ali Agha, CEO and founder of FieldAI; Michelle Lee, CEO and founder of Medra; and Aidan Madigan‑Curtis, partner at Eclipse Ventures, will share how they have made AI work where latency matters, connectivity is limited, and failure is not an option. Expect practical lessons on architectural principles, design decisions, and trade‑offs that enable AI‑centric systems to function in real‑world conditions. These insights have direct implications for companies deploying AI in challenging environments. Finally, the fifth session addresses the critical transition from prototype to production and scaling. Founders who have successfully crossed this gap in space hardware, humanoid robotics, and autonomous systems will share what they got wrong, what they would have done earlier, and what the prototype‑to‑production journey looks like when supply chains and manufacturing realities replace lab conditions. John Mackey will contribute his perspective on navigating this complex path. The panel underscores the practical challenges and strategies needed to move from a working prototype to a scalable, profitable business, offering actionable takeaways for deep‑tech startups.
Sequoia-X: quantitative stock selection for China
Sequoia-X V2 is a quantitative stock-selection system for China’s A-share market. It is rebuilt in Python with object-oriented architecture, vectorized calculations, and incremental data updates; it uses baostock for historical and daily data and SQLite for local storage. After market close it can run the selection process and send results to a Feishu group. The project provides two operating modes: a daily mode that updates data and runs strategies with parallel processing, and a backfill mode for the initial historical load. The README documents strategies including TurtleTrade, moving-average and volume breakouts, High Tight Flag, post-limit confirmation, and RPS Breakout, making it an extensible base for experimenting with selection rules. Its main strength is bringing data acquisition, persistence, strategies, and notifications into a reproducible local workflow. The README also spells out requirements—Python 3.10 or newer, environment-based configuration, and an initial historical backfill—which helps users study or modify the system. It is not a profitability guarantee or a system that should be used without validation: strategies depend on data quality, assumptions, and the Chinese market. Before making financial decisions, users should review the code, test out of sample, account for costs and biases, and evaluate the maintenance of baostock. Verdict: a useful educational quantitative starting point, with real risk if treated as advice.
VoiceStudio: Local Voice Cloning, Dubbing, and TTS Suite
VoiceStudio is an open-source application (AGPL-3.0 license) that enables voice cloning, video dubbing, dictation, and long-form audio generation entirely offline, without requiring an account, API key, or subscription. It integrates 16 text-to-speech (TTS) engines, 11 automatic speech recognition (ASR) engines, and supports 646 languages, all runnable on macOS, Windows, Linux, and via Docker containers. Designed for users who prioritize privacy and full control over their data, it serves as a local alternative to cloud-based voice services. The suite is suitable for content creators, developers, and professionals needing to generate audio or dub videos without relying on external services. It operates on local hardware, including NVIDIA GPUs (CUDA), Apple Silicon (MPS/MLX), and AMD (ROCm on Linux), with CPU-only support as well. Minimum requirements include 8 GB of RAM and 10 GB of disk space, though 16 GB and 20 GB are recommended for optimal performance. Notable features include zero-shot voice cloning, integration with models like CosyVoice 3, GPT-SoVITS, and OmniVoice, and a user-friendly interface for switching between engines. Additional tools include vocal isolation, speaker diarization, and an MCP server for custom integrations. While currently in active beta, the latest stable release is recommended for production use.
Ponytail: less code for coding agents
Ponytail is a tool and skill for coding agents focused on reducing generated code in real development tasks. Its approach gives agents more precise operational context so they can produce smaller, faster, and cheaper changes; the README publishes comparative measurements against sessions without the skill. It is relevant for teams using agents on existing repositories and wanting to measure cost, speed, and change size. Its results should be validated in the target workflow before treating the reported figures as a general guarantee.
TimesFM - Time Series Foundation Model
TimesFM is a pretrained time-series foundation model by Google Research designed for forecasting both multivariate and univariate series with native covariate support (past-only and past-and-future). It has achieved rank #1 overall on multiple benchmarks including fev-bench, TIME Benchmark, and GIFT-Eval. The latest version (3.0) improves upon previous releases with 200M parameters, extended 16k context length, optional quantile heads for long-range forecasts, and broader ecosystem integrations like BigQuery ML, Google Sheets, and Vertex AI Model Garden. Notably, while the source code is Apache-2.0 licensed, TimesFM 3.0 pretrained weights are under a separate non-commercial license prohibiting commercial or production use.
{fmt}: A fast and safe C++ text formatting library
{fmt} is an open-source formatting library that provides a fast and safe alternative to traditional C formatting functions like printf and C++ iostreams. It implements the C++20 std::format and C++23 std::print standards, using a syntax similar to Python's format strings for ease of use. Its simple API allows formatting strings with positional arguments, which is particularly useful for application localization. The library stands out for its high performance, outperforming common standard library implementations, and its small code size in both source files and compiled output. It is fully type-safe, with automatic memory management preventing buffer overflows and the ability to catch format string errors at compile time. It is ideal for C++ developers seeking efficiency, portability, and ease of use without external dependencies, under a permissive MIT license. The library features extensive testing, continuous integration across Linux, macOS, and Windows, and optional header-only configuration via the FMT_HEADER_ONLY macro.