Who's afraid of Chinese models?
694 points 495 comments on Hacker News Β· stratechery.com
Showing results for "nvidia"38 results
694 points 495 comments on Hacker News Β· stratechery.com

Across 107 enterprises, AI infrastructure spending is accelerating well ahead of the ability to see or steer its economics. Most organizations run their AI on a familiar base of hyperscalers and model-provider APIs, yet the next dollar is aimed at specialized compute almost none of them use today; a majority intend to switch or add providers within the year, many within a quarter. Buying decisions turn on integration and total cost of ownership rather than headline token price β which is fortunate, because most enterprises cannot yet see their unit economics clearly: GPUs sit at half utilization or less, and fewer than half rigorously track what their compute actually costs. The result is a compute gap β heavy, fast-moving investment running ahead of the visibility needed to control it. This wave of VentureBeat Pulse Research examines enterprise AI infrastructure and compute: where organizations are in their deployment journey, what they run AI on today, how satisfied they are, what would make them switch, where they plan to evaluate their investments, and β most revealingly β how well they can measure and control the economics of the compute underneath it all. The central finding is a compute gap β the distance between how aggressively enterprises are investing in AI infrastructure and how little of its economics they can see. Only about one in five (21%) run AI in production at scale, yet spending intentions are outrunning that maturity: the single largest planned area enterprises plan to evaluate over the next year is AI-specialized clouds (45%), a layer almost none of these enterprises use today. Meanwhile the compute already in place runs cold β 83% report GPU utilization of 50% or less β and fewer than half (44%) can rigorously track what their AI compute costs. Enterprises are buying more infrastructure faster than they can account for what they already own. Enterprises are not settled on their infrastructure vendors, either: A clear majority (64%) plan to switch or add an infrastructure provider within twelve months, and 38% within the next quarter β unusually high churn intent for a category this foundational. When they choose, they choose on integration with the existing stack (41%) and total cost of ownership (35%), not on headline price: cost per million tokens is the deciding factor for just 8%. And the frontier constraint that will shape the next round of decisions β the shift from GPU compute to memory bandwidth as inference scales β is barely on the radar, with roughly one in five enterprises either unaware of it or yet to address it. Methodology VentureBeat fielded this survey as part of its ongoing Pulse Research series, this survey focused on enterprise AI infrastructure, compute, and inference economics. Responses are filtered to organizations with more than 100 employees (n=107; the surveyβs smallest size band, 1β100 employees, is excluded), drawn from a single Q2 2026 (June) wave. Because this is one wave rather than a pooled multi-month sample, the report reads cross-sectionally and does not infer month-over-month trends. Several questions were multiple-select, so those shares can sum to more than 100%. By organization size the sample concentrates in the mid-market: 101β250 employees (36%) and 251β1,000 (27%) lead, with 1,001β5,000 (22%), 5,001β10,000 (8%), and 10,001+ (7%) above them. By role it spans managers (38%), individual contributors (28%), VPs and directors (19%), and the C-suite (13%); on purchasing authority it is buyer-credible, with 45% final decision-makers and another 30% recommenders or influencers for AI solutions. Technology/Software is the largest industry at 26%, followed by Healthcare/Life Sciences (15%), Financial Services (13%), and Retail/E-commerce (12%). At 107 respondents the sample is large enough to read directionally but should be treated as a directional signal rather than a precise measurement; it is self-selected and is not a probability sample. It also skews toward the mid-market and toward earlier-stage adopters, so it is best read as the view from organizations actively building out AI infrastructure rather than from the largest hyperscale operators. Finding 1: Ambition outpaces production Only one in five run AI in production at scale We asked where organizations sit in their AI deployment journey. Most are still building toward production rather than operating at scale. The maturity curve is front-loaded. Three-quarters of enterprises (76%) are either experimenting or running only some workloads in production, and just 21% describe AI in production at scale. This matters for everything that follows: the infrastructure decisions in this report are being made largely by organizations still early in deployment, whose compute footprint β and whose costs β are about to grow. The evaluation and switching intentions in Findings 3 and 4 are the leading edge of that build-out, not the settled preferences of operators who have already found what works. Finding 2: Enterprises run on hyperscalers and model APIs The specialized GPU clouds barely register β today We asked which providers and platforms enterprises currently use to run their AI. The answer is a familiar one: the incumbents. The current stack is hyperscaler-and-API. Google Cloud leads at 48%, and the general-purpose clouds (Google, Microsoft, AWS, Oracle) together with the major model APIs (Gemini, OpenAI, Anthropic) account for essentially all current deployment. The specialized βneocloudβ GPU providers that dominate AI-infrastructure headlines β CoreWeave, Lambda, Crusoe, Nebius and peers β register at or near zero among these enterprises today. Only 6% run their own on-prem GPU clusters and 4% a custom open-source stack. Enterprises are, for now, running AI on the providers they already buy from β which makes the evaluation intentions in Finding 3 all the more striking. (A note on reading these shares. As described in the methodology section, this sample is self-selected and skews mid-market, and this question counted every provider a respondent uses β an average of 2.1 selections each β so the figures measure presence in the stack rather than spending or primary status. A sample built this way will show a different provider mix than a spend-weighted census of the broader market; Google's strength here, for example, is consistent with its long-standing position among smaller enterprises building on AI. Read these shares as a portrait of what this AI-active cohort runs today, and treat gaps between these figures and industry-wide market share estimates as a property of the sample rather than a contradiction of either.) Finding 3: The next dollar goes to infrastructure they donβt yet run AI-specialized clouds top the evaluations list We asked where enterprises planned to evaluate AI infrastructure over the next 12 months. Their answers point away from the stack they run today. Here is the reportβs sharpest tension. The single most-cited planned evaluation area β AI-specialized clouds, at 45% β is the very category almost none of these enterprises use today (Finding 2). Nearly a third (32%) intend to evaluate non-

Onimusha: Way of the Sword is coming to GeForce NOW at launch, with the playable demo available this week. Itβs joined by Denshattack! rolling in with five new games arriving in the cloud. Plus, GeForce NOW officially launches in India, moving from beta to public availability β meaning gamers can sign up without a waitlist. […]

This GFN Thursday brings more games, more power and more ways to play on GeForce NOW. The cloud gaming service is expanding with a new GeForce RTX 5080-powered server in Toronto, bringing dedicated high performance in the cloud closer to members across the region. NTE: Neverness to Everness also gets an update in the cloud, […]
Summer is heating up β and GeForce NOW is taking players along for the ride. Start the month with Monopoly: Star Wars Heroes vs. Villains, bringing a galaxy far, far away to the iconic board-game franchise, alongside 12 new games joining the cloud this month. Plus, donβt let the sun set on the biggest GeForce […]

1 point 0 comments on Hacker News Β· hollisrobbinsanecdotal.substack.com

Teleoperation is necessary for training humanoid robots, but reinforcement learning and simulation are still necessary, says Flexion's CEO. The post How to avoid the teleoperation trap in robotics development appeared first on The Robot Report .

OpenClaw has become one of the most widely adopted agentic frameworks, but it has yet to prove itself at enterprise scale. Agents need real credentials β API keys, OAuth tokens, service accounts β to work effectively, and Brex found that traditional guardrails couldn't contain what those agents were doing with them. Brex set out to overcome these limitations by building an internal platform it calls CrabTrap. The open-source HTTP/HTTPS proxy intercepts all network traffic, examines policy rules, and uses a LLM-as-a-judge to decide whether agent requests should be approved or denied. βWhat we noticed was that the network layer was an untapped enforcement point,β Brex co-founder and CEO Pedro Franceschi told VentureBeat. βEvery request an agent makes is an opportunity to intercept, reason about, and make a policy decision.β The takeaway Franceschi wants IT leaders to draw: agent governance should shift from SDK-level permissions and model guardrails toward a centralized network control plane that enforces and learns from real in-the-wild agent behavior. How Brex targeted the transport layer The βobvious fixβ (at least initially) to the agent security gap was guardrails, and much of the early work has centered on scoped tools, per-action permissions, and human-in-the-loop approvals. But as agents evolve, each new capability means thereβs another API to tune or surface to audit, Franceschi noted. βAny agentic system with multiple tools and access to the open internet creates an immediate tension for builders: The more capable you make an agent, the more dangerous it becomes, and the safer you make it, the less useful it is,β he said. Existing solutions to this tradeoff were βweakβ: Fine-grained API tokens help at the margins but can still be misused and constrain functionality. Semantic guardrails (such as context, skills, or prompt steering) are easily bypassed by prompt injection, especially for agents connected to the internet. Agents can be βdefangedβ when given read-only access or limited toolsets, but then they can't do meaningful work, Franceschi said. On the other hand, granting broad write access and a large tool surface can result in hallucinations and real production consequences. Model context protocol (MCP) gateways enforce policy at the protocol layer β but only for traffic using MCP. Meanwhile, guardrails from LLM providers are tied to a single model and can be βopaqueβ to customize with enterprise-specific policies. And powerful tools like

Moonshot AI, the Beijing-based artificial intelligence startup backed by Alibaba, on Thursday released Kimi K3 β a 2.8-trillion-parameter model that the company says is now the largest open-source AI model in the world, and one that benchmarks show performs neck-and-neck with the most powerful proprietary systems from Anthropic and OpenAI . The release, timed to land just ahead of the 2026 World Artificial Intelligence Conference in Shanghai, is a dramatic escalation in the global AI arms race and a watershed moment for the open-source AI movement. It also marks a remarkable comeback for a company whose market position had eroded significantly over the past 18 months following DeepSeek's meteoric rise. Full model weights are scheduled to be released on July 27, according to details shared by researchers who reviewed the company's technical documentation. If you want to take Kimi K3 for a spin right now, you can β just head to kimi.com , sign up with a Google account or phone number (no credit card required), and start chatting with what may be the most powerful open-source model ever built. Inside the architecture that powers the world's largest open-source AI model Kimi K3 is a frontier-class large language model with 2.8 trillion total parameters β roughly 75 percent larger than DeepSeek's V4 Pro , which the company's own timeline chart shows at approximately 1.6 trillion parameters. The model features a 1-million-token context window, native visual understanding capabilities, and an always-on reasoning mode that the company calls "thinking mode." The model is built on two key architectural innovations developed internally at Moonshot AI: Kimi Delta Attention , a hybrid linear attention mechanism, and Attention Residuals , which the company describes as a drop-in replacement for residual connections that delivers consistent scaling gains. Both techniques were previously published as open research by the Moonshot team on GitHub . On the API side , Kimi K3 is compatible with the OpenAI SDK , lowering the integration barrier for developers already building on OpenAI or Anthropic toolchains. The model is priced at $3 per million input tokens and $15 per million output tokens, with cached input tokens dropping to just $0.30 per million β pricing that positions it roughly in line with mid-tier offerings from Western labs, but at a performance level the company claims approaches the top of the market. A promotional top-up rebate running through August 12 offers up to 30 percent back in vouchers for API credits of $1,000 or more. As Xinhua reported , a Moonshot AI executive explained the significance of the parameter count in simple terms: parameters are like neural connections in the human brain, and nearly 3 trillion of them means the model can "store more knowledge and patterns in its brain, understand more, think deeper, and answer more accurately." Benchmark results show Kimi K3 trading blows with Claude and GPT at the top of the leaderboard The benchmark results, drawn from public leaderboard data and a private evaluation by analytics firm Artificial Analysis, tell a striking story. On GDPval-AA v2 , a benchmark measuring real-world tasks across 44 occupations and 9 major industries, Kimi K3 scored 1,687 β placing it third overall, behind only Claude Fable 5 Max (1,815) and GPT-5.6 Sol Max (1,747.8), and ahead of Claude Opus 4.8 (1,600). On AA-Briefcase , a private agentic benchmark from Artificial Analysis designed to test long-horizon knowledge work, K3 climbed to second place with a score of 1,527 β beating GPT-5.6 Sol Max (1,495) and trailing only Fable 5 Max (1,587). Perhaps most impressively, K3 achieved a state-of-the-art score of 91.2 out of 100 on BrowseComp , a benchmark for long-horizon, high-difficulty information seeking. The company says it accomplished this in a single-agent setup using its 1-million-token context window, without any context compression or additional context management techniques β a feat that suggests raw context length, when paired with strong retrieval capabilities, may be more powerful than elaborate multi-agent workarounds. As one widely followed AI commentator put it on social media: "Open source is no longer lagging six months behind Western closed-source models. Read that again, and think about what it all means." That observation captures the significance of the moment. For much of the past three years, open-source models have typically trailed their proprietary counterparts by a meaningful margin. Kimi K3 appears to have closed that gap almost entirely. How a 48-hour autonomous chip design demo reveals Moonshot's real ambitions Beyond raw benchmarks, Moonshot AI showcased a proof-of-concept that may be even more revealing of K3's capabilities and the company's strategic direction. In a demonstration documented in the company's technical materials, Kimi K3 was tasked with designing a physical chip to run a nano-scale version of itself. Over 48 hours of continuous autonomous agent operation, K3 independently completed the chip's full construction pipeline β from architectural design through optimization and verification β using open-source electronic design automation tools. The result was a tiny but functional chip design, just 4 square millimeters, that achieved timing convergence at 100 MHz and could decode more than 8,700 tokens per second in simulation. This is not a production chip. It is a demonstration of what Moonshot AI clearly views as the next competitive frontier: long-range autonomous agent capabilities. The ability to sustain coherent, multi-step technical work over a 48-hour window β reading documentation, making design decisions, running verification loops, and iterating on failures β represents a qualitative leap beyond the kind of single-turn question-answering that defined the first generation of large language models. The company also highlighted a case in computational astrophysics, where K3 reportedly reproduced the universal I-Love-Q relation β a complex calculation that typically takes a senior researcher one to two weeks β in approximately two hours, reading and cross-validating more than 20 papers and implementing a complete numerical pipeline along the way. Moonshot AI's fall and rise tells the story of China's brutal AI market To understand why Kimi K3 matters, you need to understand where Moonshot AI was 18 months ago β and how far it fell. Founded in 2023 by Yang Zhilin , a Tsinghua University graduate who previously conducted research at Google and Meta, Moonshot AI quickly became one of China's most prominent AI startups. The company gained early traction in 2024 when users flocked to its Kimi platform for its long-text analysis capabilities and AI search functions. By early 2026, it had raised roughly $1.5 billion across multiple rounds, with its valuation climbing from $2.5 billion to $4.3 billion and the company reportedly seeking a new round at $5 billion . Then DeepSeek happened. The release of DeepSeek's low-cost R1 model in January 2025 disrupted the entire Chinese AI landscape, and Moonshot AI was among the hardest hit. Kimi, which had ranked third in monthly active users in China, slid to seventh. The company's strategic pivot to open-source models β beginning with Kimi K2 in July 2025 and accelerating with K2.5 in January 2026 β was in large part an effort to reclaim relevance. Kimi K3 is the culmination of that effort β and the sheer scale of the model suggests that Moonshot AI has been planning this move for some time. Training a 2.8-trillion-parameter model requires enormous computational resources and months of preparation, which means the architectural and infrastructure decisions behind K3 were likely locked in well before the model reached the public. Why open-sourcing the world's biggest model is a geopolitical chess move The decision to release K3's full weights on July 27 is strategically significant and worth parsing carefully. The company's own timeline chart of open-source frontier model scale positions K3 as a dramatic outlier, towering above competitors like DeepSeek (1.6T), Xiaomi (1.02T), and Alibaba (397B). By releasing the world's largest open-source model, Moonshot AI is making a bid to become the center of gravity for the global open-source AI developer community. This follows a broader trend among Chinese AI companies. As Reuters noted , open-sourcing allows companies to "showcase their technological capabilities and expand developer communities as well as their global influence, a strategy likely to help China counter U.S. efforts to limit Beijing's tech progress." DeepSeek, Alibaba, Tencent, and Baidu have all released open-source models. But none have released anything at this parameter count. For enterprise technology leaders, the implications are concrete. A 2.8-trillion-parameter open-source model that performs at near-frontier levels creates new options for companies that want to fine-tune, self-host, or build proprietary systems on top of a capable base model β without being locked into API contracts with OpenAI or Anthropic. The trade-off, of course, is that running a model of this size requires substantial GPU infrastructure. Inference at 2.8 trillion parameters is not something that runs on a single server rack. That said, Moonshot AI has signaled awareness of this challenge. Its Mooncake project, which won the Best Paper award at FAST 2025, pioneered KV-cache-centric disaggregated serving for large language models β an architecture designed specifically to make inference at extreme scale more practical and cost-efficient. Kimi Code and a three-tier model lineup form the foundation of Moonshot's enterprise play Alongside K3, Moonshot AI continues to invest heavily in its coding agent ecosystem. Kimi Code , the company's open-source coding tool that competes with Anthropic's Claude Code and Google's Gemini CLI, received two major updates on the same day as K3's launch β versions 0.25.0 and 0.26.0 β adding features like expanded subagent tooling, background task management, and security fixes. The Kimi Code CLI has accumulated over 3,100 stars on GitHub and features integration with VSCode, Cursor, and Zed. The latest release expanded the "coder subagent" tool set to include background tasks, todo lists, plan mode, skill invocation, and nested agents β effectively turning the coding agent into a multi-layered autonomous system capable of managing complex software engineering projects with minimal human intervention. This is not incidental. Coding tools have become a critical revenue driver for AI labs. As Anthropic disclosed in January, Claude Code reached $1 billion in annualized recurring revenue . By building Kimi Code as an open-source alternative that defaults to Kimi's own models β but supports other providers β Moonshot AI is positioning itself to capture developer workflows and, eventually, enterprise contracts. The company's model lineup now includes three tiers: K3 as the flagship ($3/$15 per million tokens for input/output), K2.7 Code as a specialized coding model ($0.95/$4), and K2.6 as a general-purpose option ($0.95/$4). All three support context windows of 256,000 tokens or above, with K3 offering the full 1-million-token window. Context caching is automatic β no cache ID, TTL, or extra parameter is required β a small but meaningful developer-experience advantage over competitors that require explicit cache management. What Kimi K3 means for the future of enterprise AI and the global model landscape Kimi K3's release forces a recalibration of several assumptions that have guided enterprise AI strategy. The performance gap between open-source and proprietary models has functionally closed at the frontier. If K3's benchmark numbers hold up under independent evaluation β and particularly once the open weights are available for community testing on July 27 β it will be difficult for closed-source providers to justify premium pricing purely on the basis of capability. The locus of AI innovation, meanwhile, continues to shift. China's AI ecosystem, which many Western observers questioned after early struggles with chip export restrictions, has now produced a model that competes with the best systems from companies with direct access to
Recorded at HumanX, Ryan sits down with Garima Kapoor and Anand Babu Periasamy, co-founders and co-CEOs of MinIO, to chat about eliminating the storage bottlenecks that leave GPUs underutilized, their partnership with NVIDIA on the new STX reference architecture, and why modern AI infrastructure is converging on S3-compatible object storage. βββββ β ββββββ β βββββββββ ββββββ ββββββ βββββββ β ββββββ ββββββ βββ ββββ βββββββ ββββββ ββββββββββ βββββββββββββββ βββββββββββ βββ βββ βββ β β ββββ ββ βββ ββ ββ β ββ ββ β β βββββββββ βββ ββ β βββββββ ββββββββ βββ β β ββ ββββ ββ ββ ββββββ ββ ββββββββ ββ ββββ βββββββββββββ ββββ ββ βββ βββββββββ ββ βββ βββββββ β ββββββ ββ βββββββ ββββββββ ββ ββ β βββββββββ ββββββ β β βββ βββ β β βββ βββ ββββ ββββββββ β βββ ββββ ββ βββ β ββββββββββ ββ βββ βββ ββββββββ ββββββ βββ βββ βββ βββ β β β β βββ βββ βββββββ βββ β β βββ βββ βββββββ βββββββ βββ βββ ββββββ β β β βββ β β βββββββ βββββββ ββββββ βββββ βββββ βββ βββ βββ βββββββββ β βββββββββ ββββ ββ ββββββ βββ βββ βββ β βββββ β βββββββββ βββββββββββ βββββββββ ββ ββ β ββ ββ β β βββββββββ βββ ββ β βββββββ ββββββββ βββ β β ββ ββββ ββ βββββββββββββ ββββββ β β βββ βββ β β βββ βββ ββββ ββββββββ β βββ ββββ ββ βββ β ββββββββββ ββ βββ βββ ββββββββ ββββββ βββ βββ βββ βββ β β β β βββ βββ βββββββ βββββββ βββ βββ βββββββ βββββββ βββ βββ ββββββ β β β βββββββ βββββββ βββββββ ββββββ βββββ βββββ βββ βββ βββββββ βββββββ βββ β β βββββββββ β βββββββ βββββββ ββ βββ ββββββββ ββββββ β βββββββββββββββββ β
This Week