Scored 219 articles from 96 feeds; 15 included in digest.
Run ID: run-1786950984567
Generated: August 17, 2026 at 03:32 AM ET
Summaries: claude-sonnet-4-6; enrichment 15/15 succeeded
| Source | Type | Included | Scored | 28d Digest Rate | 28d Avg Score | 28d Hotlist Hit | 7d Article Age | 28d Confidence |
|---|---|---|---|---|---|---|---|---|
| arXiv CompSci CL | research | 3 | 24 | ~5% | ~0.12 | ~0% | 3.5h | Low sample |
| Hacker News | commentary | 2 | 22 | 4% | 0.07 | 0% | 8.5h | Stable |
| Medium AI (keyword) | commentary | 2 | 10 | 15% | 0.15 | 0% | 0.6h | Stable |
| Medium Artificial Intelligence (keyword) | commentary | 2 | 10 | 17% | 0.16 | 0% | 0.5h | Stable |
| Guardian | news | 1 | 25 | 1% | 0.03 | 0% | 7.8h | Stable |
| arXiv CompSci ML | research | 1 | 25 | ~2% | ~0.08 | ~0% | 3.5h | Low sample |
| MyFT | news | 1 | 18 | 10% | 0.12 | 0% | 4.1h | Stable |
| WSJ US Business | news | 1 | 6 | 4% | 0.12 | 1% | 8.6h | Stable |
| WSJ Tech | news | 1 | 2 | 19% | 0.23 | 3% | 7.6h | Stable |
| BIG by Matt Stoller | commentary | 1 | 1 | Collecting data | Collecting data | Collecting data | 2.0h | Collecting |
| Reddit AntiAI | news | 0 | 21 | 5% | 0.08 | 1% | 6.6h | Stable |
| Bloomberg Markets | news | 0 | 20 | 4% | 0.10 | 1% | 3.0h | Stable |
| NYT front page | news | 0 | 13 | 2% | 0.04 | 1% | 5.2h | Stable |
| Seeking Alpha News | commentary | 0 | 7 | 4% | 0.09 | 1% | 0.8h | Stable |
| Outside Law School Scam - Comments | commentary | 0 | 3 | Collecting data | Collecting data | Collecting data | 21.1h | Collecting |
| The Verge | news | 0 | 3 | 5% | 0.10 | 1% | 7.4h | Stable |
| WSJ Social Economy | news | 0 | 2 | 3% | 0.09 | 0% | 5.0h | Stable |
| Daring Fireball | commentary | 0 | 1 | ~8% | ~0.10 | ~0% | 4.3h | Low sample |
| El Reg Offbeat | news | 0 | 1 | Collecting data | Collecting data | Collecting data | 9.7h | Collecting |
| FT Alphaville | news | 0 | 1 | ~3% | ~0.11 | ~0% | 3.5h | Low sample |
| Latent Space | commentary | 0 | 1 | Collecting data | Collecting data | Collecting data | 3.7h | Collecting |
| Reddit FuckAI | news | 0 | 1 | Collecting data | Collecting data | Collecting data | 3.9d | Collecting |
| TechCrunch | news | 0 | 1 | 11% | 0.16 | 0% | 6.7h | Stable |
| Venture Beat | commentary | 0 | 1 | ~67% | ~0.50 | ~0% | 7.7h | Low sample |
Source: arXiv CompSci CL
Type: research
Included: 3
Scored: 24
28d Digest Rate: ~5%
28d Avg Score: ~0.12
28d Hotlist Hit: ~0%
7d Article Age: 3.5h
28d Confidence: Low sample
Source: Hacker News
Type: commentary
Included: 2
Scored: 22
28d Digest Rate: 4%
28d Avg Score: 0.07
28d Hotlist Hit: 0%
7d Article Age: 8.5h
28d Confidence: Stable
Source: Medium AI (keyword)
Type: commentary
Included: 2
Scored: 10
28d Digest Rate: 15%
28d Avg Score: 0.15
28d Hotlist Hit: 0%
7d Article Age: 0.6h
28d Confidence: Stable
Source: Medium Artificial Intelligence (keyword)
Type: commentary
Included: 2
Scored: 10
28d Digest Rate: 17%
28d Avg Score: 0.16
28d Hotlist Hit: 0%
7d Article Age: 0.5h
28d Confidence: Stable
Source: Guardian
Type: news
Included: 1
Scored: 25
28d Digest Rate: 1%
28d Avg Score: 0.03
28d Hotlist Hit: 0%
7d Article Age: 7.8h
28d Confidence: Stable
Source: arXiv CompSci ML
Type: research
Included: 1
Scored: 25
28d Digest Rate: ~2%
28d Avg Score: ~0.08
28d Hotlist Hit: ~0%
7d Article Age: 3.5h
28d Confidence: Low sample
Source: MyFT
Type: news
Included: 1
Scored: 18
28d Digest Rate: 10%
28d Avg Score: 0.12
28d Hotlist Hit: 0%
7d Article Age: 4.1h
28d Confidence: Stable
Source: WSJ US Business
Type: news
Included: 1
Scored: 6
28d Digest Rate: 4%
28d Avg Score: 0.12
28d Hotlist Hit: 1%
7d Article Age: 8.6h
28d Confidence: Stable
Source: WSJ Tech
Type: news
Included: 1
Scored: 2
28d Digest Rate: 19%
28d Avg Score: 0.23
28d Hotlist Hit: 3%
7d Article Age: 7.6h
28d Confidence: Stable
Source: BIG by Matt Stoller
Type: commentary
Included: 1
Scored: 1
28d Digest Rate: Collecting data
28d Avg Score: Collecting data
28d Hotlist Hit: Collecting data
7d Article Age: 2.0h
28d Confidence: Collecting
Source: Reddit AntiAI
Type: news
Included: 0
Scored: 21
28d Digest Rate: 5%
28d Avg Score: 0.08
28d Hotlist Hit: 1%
7d Article Age: 6.6h
28d Confidence: Stable
Source: Bloomberg Markets
Type: news
Included: 0
Scored: 20
28d Digest Rate: 4%
28d Avg Score: 0.10
28d Hotlist Hit: 1%
7d Article Age: 3.0h
28d Confidence: Stable
Source: NYT front page
Type: news
Included: 0
Scored: 13
28d Digest Rate: 2%
28d Avg Score: 0.04
28d Hotlist Hit: 1%
7d Article Age: 5.2h
28d Confidence: Stable
Source: Seeking Alpha News
Type: commentary
Included: 0
Scored: 7
28d Digest Rate: 4%
28d Avg Score: 0.09
28d Hotlist Hit: 1%
7d Article Age: 0.8h
28d Confidence: Stable
Source: Outside Law School Scam - Comments
Type: commentary
Included: 0
Scored: 3
28d Digest Rate: Collecting data
28d Avg Score: Collecting data
28d Hotlist Hit: Collecting data
7d Article Age: 21.1h
28d Confidence: Collecting
Source: The Verge
Type: news
Included: 0
Scored: 3
28d Digest Rate: 5%
28d Avg Score: 0.10
28d Hotlist Hit: 1%
7d Article Age: 7.4h
28d Confidence: Stable
Source: WSJ Social Economy
Type: news
Included: 0
Scored: 2
28d Digest Rate: 3%
28d Avg Score: 0.09
28d Hotlist Hit: 0%
7d Article Age: 5.0h
28d Confidence: Stable
Source: Daring Fireball
Type: commentary
Included: 0
Scored: 1
28d Digest Rate: ~8%
28d Avg Score: ~0.10
28d Hotlist Hit: ~0%
7d Article Age: 4.3h
28d Confidence: Low sample
Source: El Reg Offbeat
Type: news
Included: 0
Scored: 1
28d Digest Rate: Collecting data
28d Avg Score: Collecting data
28d Hotlist Hit: Collecting data
7d Article Age: 9.7h
28d Confidence: Collecting
Source: FT Alphaville
Type: news
Included: 0
Scored: 1
28d Digest Rate: ~3%
28d Avg Score: ~0.11
28d Hotlist Hit: ~0%
7d Article Age: 3.5h
28d Confidence: Low sample
Source: Latent Space
Type: commentary
Included: 0
Scored: 1
28d Digest Rate: Collecting data
28d Avg Score: Collecting data
28d Hotlist Hit: Collecting data
7d Article Age: 3.7h
28d Confidence: Collecting
Source: Reddit FuckAI
Type: news
Included: 0
Scored: 1
28d Digest Rate: Collecting data
28d Avg Score: Collecting data
28d Hotlist Hit: Collecting data
7d Article Age: 3.9d
28d Confidence: Collecting
Source: TechCrunch
Type: news
Included: 0
Scored: 1
28d Digest Rate: 11%
28d Avg Score: 0.16
28d Hotlist Hit: 0%
7d Article Age: 6.7h
28d Confidence: Stable
Source: Venture Beat
Type: commentary
Included: 0
Scored: 1
28d Digest Rate: ~67%
28d Avg Score: ~0.50
28d Hotlist Hit: ~0%
7d Article Age: 7.7h
28d Confidence: Low sample
According to The Wall Street Journal, major technology companies have made massive spending commitments on data-center leases and chips that do not appear on their balance sheets. The article reports that these off-balance-sheet commitments amount to approximately $3 trillion more than publicly visible figures indicate.
Keywords: AI capital expenditure, off-balance-sheet obligations, operating leases vs. capital leases, Big Tech financial leverage, data center infrastructure, semiconductor commitments, circular investment, financial reporting opacity, investment productivity puzzle, balance sheet manipulation
The Financial Times article argues that the widespread adoption of open-source AI models from China will carry broader geopolitical and governance implications. It contends that countries adopting Chinese AI models will also absorb the standards and governance frameworks embedded within them. The piece is categorized under the FT's Artificial Intelligence and Global Economy topics.
Keywords: open-source AI models, China shock, model monoculture, technological standards, geopolitical dependencies, governance frameworks, global economy, technology adoption
This arXiv preprint (cs.CL, submitted August 2026) investigates whether AI agents can simulate the outcomes of A/B tests accurately enough to serve as a pre-screening step before running live experiments. The authors formalize the problem as a "Simulated Randomized Controlled Trial" (S-RCT) and introduce a two-layer error decomposition that separates agent approximation error from subsampling error, allowing each source of error to be addressed independently. The framework is described as agent-agnostic, accommodating any behavioral model from fine-tuned specialists to general-purpose foundation models. Validated against 67 historical marketing A/B tests, a baseline S-RCT using an off-the-shelf foundation model achieves a directional accuracy (sign overlap) of 0.70 but systematically overestimates effect magnitudes. The authors propose two methodological improvements: a two-phase pre-period calibration protocol that reduces squared prediction error (after removing irreducible measurement noise) by approximately 77×, and a within-subject design in which each agent is exposed to both experimental arms, reducing standard errors by approximately 2.4×. The paper also discusses current limitations of the approach and identifies application contexts where agentic simulation signals may be most useful to experimenters.
Keywords: A/B testing, AI agents, simulated experimentation, foundation models, behavioral modeling, causal inference, product development optimization
A Guardian investigation has found an apparent discrepancy between Microsoft's public statements about its AI datacentre capacity and the actual number of advanced AI chips it has in operation. According to internal documents seen by the Guardian, Microsoft currently has approximately 2.2 million AI chips installed across its global datacentres — a figure described as less than half what some industry experts had anticipated given the company's public claims of adding 5 gigawatts of datacentre capacity over the past two years. The investigation highlights the difficulty of independently verifying AI build-out progress, noting that Nvidia does not publicly report how many chips it sells or to whom, and that its tech company clients similarly do not disclose their holdings. Researchers cited in the article, including professors at the University of California, Riverside and the University of Rhode Island, suggest that Microsoft's third-party audited sustainability reports imply a much lower operational AI capacity than its financial filings indicate. Microsoft's CEO Satya Nadella is quoted acknowledging that the core constraint is not chip supply but electrical power and the availability of ready datacentre facilities, saying he has chips 'sitting in inventory' that cannot yet be plugged in. The article also notes that Microsoft's flagship Fairwater datacentre project in Wisconsin had not come fully online as of May 2025, despite earlier announcements. Microsoft disputed the Guardian's findings, stating that the estimates were based on 'incorrect assumptions,' but did not specify which figures were wrong. Nvidia did not respond to a request for comment.
Keywords: Microsoft, AI chips, supply chain, capacity constraints, semiconductor shortage
Published on Medium by TF Business Solutions, this article addresses the concept of transitioning from founder-led to system-led growth. The available text indicates it is aimed at startups and focuses on building systems designed to deliver clarity, speed, and impact without burning out team members. No further article content is available beyond this brief description.
Keywords: organizational restructuring, systems-driven growth, startup scaling, founder-led to scalable systems
According to the article title and Reuters URL, Nvidia has dramatically reduced the amount of OpenAI infrastructure financing it may guarantee, scaling back from a previously reported $250 billion commitment. No further detail is available from the supplied article text, which contains only a link to the Hacker News comments thread.
Keywords: Nvidia, OpenAI, infrastructure financing, capital allocation, AI compute
The article, published on Medium by Y-Consulting, argues that while AI has appeared affordable to end users, the true costs are now becoming visible. According to the piece, these costs include energy consumption, infrastructure demands, dependency on AI systems, and the erosion of human capability.
Keywords: AI infrastructure costs, Energy consumption, Technological dependency, Human capital erosion, Economic sustainability of AI
The article, published on Medium, reflects on the author's experience regenerating an open source tool in 66 minutes using AI assistance, and the approximately three weeks it subsequently took to develop confidence in the result. The piece frames this as a lesson about the true cost of forking open source software in 2026. The available article text is limited to a brief snippet and does not provide further detail about the specific tool, the regeneration process, or the trust-building steps involved.
Keywords: open source, AI code generation, software forking, verification costs, trust, automation gap, software maintenance
Published on Medium, this article discusses how Google's AI Mode is affecting SEO traffic. According to the available excerpt, websites can maintain reasonable search rankings while still losing clicks, as users increasingly receive information directly within Google Search results. The article indicates it goes on to address ways websites can adapt to these changes, though the full text is not available in the supplied excerpt.
Keywords: Google Search, AI-generated answers, SEO traffic, click-through rates, search results, digital advertising, website adaptation
Alibaba Group is selling its videogame business in a deal valued at a minimum of $1.5 billion, according to the Wall Street Journal. The move is part of the Chinese company's broader strategic shift toward artificial intelligence.
Keywords: Alibaba, divestiture, videogame business, capital reallocation, AI investment, strategic pivot, technology conglomerate
In this piece from BIG by Matt Stoller, the author uses the ongoing Paramount-Warner Discovery merger dispute as a lens to examine 'capital strikes' — the use of threatened investment withdrawal to pressure governments into favorable policy outcomes. Stoller describes how Paramount CEO David Ellison threatened to relocate the studio out of California and CBS out of New York City unless state attorneys general dropped their antitrust opposition to the merger. Stoller argues the threat backfired, hardening the presiding judge's resolve and undermining settlement prospects. The article places this episode in a broader historical and contemporary pattern. Stoller draws parallels to 1930s aluminum monopolist Andrew Mellon threatening to offshore Alcoa's production to block tariff reductions, Jeff Bezos canceling Amazon's HQ2 plans in New York after political pushback over subsidies, Ken Griffin threatening to move Citadel operations from New York over a mayoral campaign video, and Microsoft threatening to pull UK investment to win approval for its Activision acquisition — which Stoller says ultimately succeeded in reversing the UK competition authority's position. Stoller argues that as concentrated capital faces more democratic and regulatory challenges, capital strikes will become more common. He contends the structural response would require governments to supply alternative capital, break up monopolies, weaken intellectual property protections, and impose public utility rules on key infrastructure. Brief mentions of additional news items — including antitrust scrutiny of Epic Systems and a crypto banking charter — are noted as reserved for paywalled content.
Keywords: concentrated capital, regulatory threat, capital strike, oligopoly, corporate strategy
This paper argues that large language models used as judges in principle-based regulatory contexts—where standards such as "fair, clear, and not misleading" resist reduction to binary rules—must be evaluated on four axes: accuracy, paraphrase robustness, adversarial robustness, and calibration. The authors introduce Principle-Bench, a benchmark of 168 cryptoasset financial-promotion scenarios mapped to two UK Financial Conduct Authority principles, including paraphrase variants, adversarial keyword-stuffing inputs, and boundary perturbations developed under a pre-registered rubric. They also present Ceca (Calibrated Exemplar-Cluster Assessment), an assessor designed to produce auditable, per-exemplar counterfactual attributions with calibration. Evaluating keyword counting, three sentence-transformer embedders, an open-weight LLM judge, and a calibrated cascade, the authors find no method dominates across all four axes. A 120-billion-parameter LLM judge that performs strongest on benign inputs drops 47 accuracy points (from 0.74 to 0.27) on keyword-stuffed Consumer Duty inputs, a failure the authors term "compliance theatre." Inter-judge agreement on that split reaches only Cohen's kappa of 0.16, localizing the failure to the model rather than the dataset. The paper concludes that deployment-grade LLM judges for principle-based regulation must report per-principle adversarial deception rates and post-hoc calibration in addition to aggregate accuracy.
Keywords: LLM-as-judge, principle-based regulation, financial compliance, adversarial robustness, calibration, cryptoasset promotion, model reliability, regulatory AI
This arXiv paper (cs.CY, submitted August 14, 2026) presents a prespecified randomized algorithm audit examining which factors drive large language model (LLM) recommendations when patients ask AI assistants to choose among physicians. The researchers tested seven models—six open-weight and GPT-4o-mini—across 40,068 scored responses generated from 3,024 choice sets using synthetic family-medicine physician profiles with independently randomized attributes, three patient personas, nine prompt paraphrases, and nine experimental arms. Physician gender and ethnicity were signaled through names using correspondence-audit methodology. Key findings include: reputation signals were the dominant drivers, with a rating increase from 3.9 to 4.7 raising choice probability by 31.4 percentage points, and a fee increase from $90 to $190 lowering it by 20.0 percentage points. Demographic parity was rejected, but the direction differed from findings in human audit studies—female-signaled names gained 2.5 percentage points and Hispanic-, South-Asian-, and Black-signaled names gained 1.3–2.9 percentage points over White-signaled names, effects the authors quantify as equivalent to $7–$14 per visit in fee terms. A first-listed position with no other content was worth an $11 fee equivalent. Critically, the models cited gender or ethnicity in at most 0.03% of their stated reasons and abstained in only 0.39% of trials, meaning the demographic effects were invisible in models' own explanations. The authors argue this makes transparency approaches relying on model self-reporting insufficient for detecting such biases. One reasoning model failed the study's prespecified auditability criteria. The authors propose recurring behavioral audits using frozen experimental designs—allowing any new model to be tested against identical stimuli—as a more suitable monitoring approach.
Keywords: Large language models, Algorithmic bias, Physician recommendation, Demographic parity, AI intermediary, Algorithm audit, Model transparency, Healthcare market access
This arXiv paper (cs.SE) reports early results from an ongoing case study examining the cost of building AI-intensive software with AI assistance. A six-person student team developed a conversational onboarding assistant over one academic term, incorporating features such as RAG-based code chat, guided tours, dependency graphs, and technical-debt analysis, using pervasive AI tools throughout development. The researchers applied a three-layer cost model tracking actual AI expenditure, self-reported human effort, and a human counterfactual estimate. They initially reported a 19.4x cost ratio (AI-assisted vs. unassisted human development), but a subsequent review uncovered two independent measurement errors: incorrectly inferring per-token costs under a flat-rate subscription, and using incorrect regional labor rates for the counterfactual. These errors together inflated the ratio by approximately 2x; the corrected figure is approximately 9.9x. The paper frames this correction as a generalizable finding, arguing that both types of errors are easy to make, invisible in the final reported number, and likely common in similar empirical reports. The authors also outline plans for developing a more robust and replicable costing methodology for AI-intensive software development.
Keywords: AI-intensive software development, development cost measurement, methodology, case study, RAG-based systems, human-AI collaboration, labor cost accounting
SemiAnalysis, in collaboration with Nathan Iyer, reports that its Energy Model team spent six months reverse-engineering PJM's internal 'Reserve Requirement Study' — the model PJM uses to determine how much generation capacity to procure through its annual auctions — and found significant methodological errors with large financial consequences for ratepayers. PJM is the largest electricity market in the United States, serving 66 million residents across multiple states. The article argues that PJM's capacity model underestimates existing generation capacity by approximately 4 gigawatts because it fails to account for two factors: (1) cold, dense winter air increases the output of gas-fired power plants by up to 25%, and (2) roughly 400 of PJM's ~700 gas plants have completed federally mandated winter reliability upgrades since Winter Storm Elliott. As a result, SemiAnalysis estimates PJM overstated its supply shortfall and wasted approximately $12 billion in ratepayer money across the 2025–2027 period through inflated capacity auction prices. The article further argues that PJM's capacity market structure is structurally anti-growth: it sets unrealistically short lead times for new plants, maintains long interconnection queues, and uniquely among global capacity markets does not distinguish between new and existing generation — meaning the premium paid to incentivize new construction is also paid to existing plants that require no such incentive. SemiAnalysis also raises concerns about PJM's planned emergency auction, scheduled for September–October 2025, which would sign capacity contracts running to 2043 funded by anticipated large new loads such as data centers. The article contends this auction proceeds without committed counterparties, and that if large loads opt out or self-supply, residential ratepayers would bear the costs. Internal governance disputes have repeatedly blocked proposed fixes, including seasonal capacity ratings, despite support from PJM's own task forces and market monitor.
Keywords: PJM, electricity market, modeling error, grid operator, pricing, ratepayers, energy infrastructure