Company · updated daily
DeepSeek
DeepSeek's R1 model triggered a 2025 market shock over training-cost efficiency; its data handling, censorship of politically sensitive topics, and national-security scrutiny in the US are tracked here daily.
Performance of 5 Large Language Models in Perioperative Consultation for Pediatric Hypospadias: Cross-Sectional Comparative Study
Background: Hypospadias is a common congenital malformation requiring surgery. Caregivers face substantial perioperative information needs, and large language models (LLMs) offer a potential health education channel, but their performance in pediatric urology and the relation between citation accuracy and clinical content safety lack systematic evaluation. Objective: This study aimed to evaluate 5 LLMs (ChatGPT-4o, Gemini-2.5-Pro, OpenEvidence, Zhipu Qingyan, and DeepSeek) for pediatric hypospad
The Performance of ChatGPT-4o and DeepSeek-R1 in Interpreting Thyroid Nodule Ultrasound Text Reports: Multicenter Study
Background: Although thyroid nodules are detected in up to 60% of adults on ultrasound, the vast majority are benign, creating a substantial decision-making burden compounded by heterogeneous practice guidelines. Large language models (LLMs) show promise in processing unstructured medical text and are emerging as tools for report interpretation among both clinicians and patients. However, their reliability across distinct clinical tasks in thyroid ultrasound interpretation remains poorly charact
AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology II: Project Planning and Proposal Evaluation
We investigate how well large language models (LLMs) can assist scientific project planning and proposal evaluation. One-page project plans were independently generated for eight expert-conceived research projects in physics, astrophysics, and cosmology by human researchers and three contemporary LLMs (ChatGPT, Claude, and DeepSeek; mid-2025 models, used with their default tool access). The resulting 32 proposals were blindly evaluated by four human reviewers and two newer frontier LLMs (Claude
Math Education Digital Shadows for Investigating Learning with GenAI: Mathematics Performance, Anxiety, and Confidence in LLMs
arXiv:2604.27618v2 Announce Type: replace-cross Abstract: Understanding the impact of large language models (LLMs) on mathematics education requires data on LLMs' mathematical performance and biases. To this end, we introduce Math Education Digital Shadows (MEDS), a dataset mapping how LLMs reason about mathematics across human- and AI-like personifications. MEDS comprises 28,000 runs from 14 LLMs (i.e., Mistral, Qwen, DeepSeek, IBM Granite, Microsoft Phi, and xAI Grok) generated under human-sha
Low profile, high AI ambition: what leaked comments reveal about DeepSeek’s Liang Wenfeng
In an era dominated by aggressive tech founders who chase billion-dollar valuations and maximal profits while curating loud public profiles, Liang Wenfeng stands out for his insistence on staying in the background. With only a couple of photographs of him circulating online, the founder of Chinese artificial intelligence start-up DeepSeek and quantitative hedge fund High-Flyer Quant has long been a reclusive figure. But last week, a leaked transcript of a closed-door meeting with potential...
DeepSeek pauses its second fundraising round after founder’s leaked investor comments go viral
DeepSeek has suspended its second fundraising round after comments made by founder Liang Wenfeng during private investor meetings were leaked and went viral on Chinese social media, Bloomberg reported on Friday. Liang verbally informed prospective backers to hold off on the round, which had been targeting a pre-money valuation of roughly 480 billion yuan, equivalent […] This story continues at The Next Web
DeepSeek tells prospective investors of funding pause, Bloomberg News reports
Chinese AI startup DeepSeek has told prospective investors in its second fundraising round that it is suspending the deal for now, Bloomberg News reported on Saturday, citing people familiar with ...
Investment with Chinese characteristics: how Beijing’s money is reshaping tech ventures
On the surface, China’s cutting-edge tech sector – from the algorithmic breakthroughs of DeepSeek and Zhipu AI to the hardware of Unitree Robotics and ChangXin Memory Technologies (CXMT) – mirrors Silicon Valley’s venture capital-backed ecosystem. But a closer look at their financing histories reveals a common investor: the Chinese state. Beijing’s strong presence underscores a more profound structural shift in how China’s frontier technology is being funded. As Western venture capital and...
DeepSeek’s boss made the case for export controls
Transformer Weekly: Trahan and Obernolte bill, OpenAI-Hugging Face hack, and CAISI chief resigns
Do language families matter? Evaluating LLMs for sentiment analysis through a hierarchical cross-lingual lens
Social media sentiment analysis has become one of the most significant instruments for understanding the opinion of the population in the spheres of healthcare, politics, and education. Yet, large language models (LLMs) remain unevenly distributed in their linguistic coverage, failing to adequately serve a large portion of the world's languages. This study evaluates five state-of-the-art LLMs: GPT-4o, Gemini 2.0 Flash, DeepSeek-V3, Mistral Large, and Claude 3.7 Sonnet on three-class sentiment cl
What China Does Not Want the U.S. to Know
In a leaked call transcript, we learn what DeepSeek's CEO thinks are China's weaknesses and strengths in the AI race with the United States | Edition #309
DeepSeek puts AGI research ahead of products and commercial growth
According to a report by IT Home, DeepSeek, the Chinese AI developer behind the open-source R1 reasoning model, is prioritizing artificial general intelligence (AGI) research over building a consumer platform or maximizing near-term revenue. The report is based on a circulated transcript of a four-hour investor meeting involving founder Liang Wenfeng. Why it matters: The […]
Kimi K3 developer Moonshot AI expedites fundraising ahead of planned IPO, source says
Chinese artificial intelligence unicorn Moonshot AI, whose release of its Kimi K3 model last week sparked talk of another “DeepSeek moment”, has accelerated its fundraising ahead of a planned initial public offering (IPO) in Hong Kong, eyeing a new financing round as early as next month, a source says. The Beijing-based company was set to close its current funding round at a US$30 billion valuation later this month or early next month, according to a person familiar with the matter. It planned..
Poolside releases Laguna S 2.1, the open-weight coding model pitched as the West’s answer to DeepSeek and Qwen
Poolside has released Laguna S 2.1, a 118-billion-parameter open-weight model built for agentic coding that the San Francisco startup says matches or exceeds models several times its size. The model uses a mixture-of-experts architecture with eight billion active parameters per token, is compact enough to run on a single Nvidia DGX Spark desktop system, and […] This story continues at The Next Web
Open-weight models now match frontier cyber performance from just four months ago at a fraction of the cost
The British AI Security Institute warns that open-weight models like GLM-5.2 and DeepSeek V4-Pro now trail closed frontier models in cyber capabilities by four to seven months. At the start of 2025, the gap was still six to ten months. It also found that safety measures on open models are largely ineffective, leaving defenders less time to prepare. The article Open-weight models now match frontier cyber performance from just four months ago at a fraction of the cost appeared first on The Decoder
(g+) AI: DeepSeek Champions China's Bid to Flood the World With Cheap AI
On a video conference call from Hangzhou earlier this year, one of China's hottest startups held a four-hour pitch meeting with potential venture investors that was unusual by almost any measure. Von Saritha Rai und Olivia Poh ( Deepseek , KI )
CXMT’s mega IPO draws frenzy from retail investors to DeepSeek founder’s fund
As soon as ChangXin Memory Technologies opened its blockbuster initial public offering to public subscription on Thursday morning, Luo Yi applied for 86,000 shares – despite acknowledging that she knew little about the semiconductor industry. The 60-year-old stock investor from southwestern Sichuan province was encouraged by her securities account manager, who told her that CXMT’s unusually large share sale could produce a higher allotment rate than most Chinese mainland IPOs. Luo’s application.
DeepSeek will offenbar an die Börse – IPO-Einreichung noch 2026 möglich
Nur Wochen nach einer milliardenschweren Finanzierungsrunde soll DeepSeek die nächste Runde und parallel einen Börsengang planen.
AI investor mania: China’s DeepSeek chases US$70 billion valuation in fresh round
Chinese frontier artificial intelligence start-up DeepSeek is in talks with investors to raise a new financing round at around US$70 billion pre-investment valuation, shortly after closing a landmark first round, reflecting unquenched investor enthusiasm for the company. The plan comes on the heels of the Chinese AI darling’s recently closed Series A funding round in June – its first external funding round – which brought in about US$7 billion and valued the Hangzhou-based firm at nearly US$60..
DeepSeek may file mainland IPO application this year
DeepSeek, the Hangzhou-based AI model developer, has begun preparing for a mainland IPO and may file its listing application as soon as this year, according to Bloomberg, citing people familiar with the matter. The company is targeting a 2027 debut. DeepSeek is in talks with accounting and banking advisers and is seeking additional private funding […]
DeepSeek needs more cash just weeks after closing its first $7 billion round
DeepSeek is already raising again. The Chinese AI lab just closed its first funding round and needs capital for its own data centers and chips to keep its aggressive pricing strategy going. The article DeepSeek needs more cash just weeks after closing its first $7 billion round appeared first on The Decoder .
Formal Mechanisms for Market Stability in Self-Interested Agent Societies: A Marketplace Simulation Study
Self-interested agents, left unconstrained, tend toward defection in repeated social dilemmas, causing cooperative gains from trade to collapse. This paper investigates what formal mechanisms, layered on top of unrestricted communication, are sufficient for a society of such agents to maintain market stability, and how resilient those mechanisms are to adversarial attack. We instantiate the research question as a multi-agent marketplace simulation where 18 LLM agents (DeepSeek-V3) with complemen
Contagion Networks: Evaluator Preference Propagation in Multi-Agent LLM Systems
When large language models serve as evaluators in multi-agent systems, their strategy preferences -- whether induced by explicit prompts or by shared architectural priors -- propagate through the agent network. We introduce Contagion Networks, a formal framework for measuring how evaluator preferences spread across interacting LLM agents. In a controlled 3-agent experiment using DeepSeek-chat with three distinct evaluator preference profiles (structured, balanced, evidence-based), we measure the
The Shibboleth Effect: Auditing the Cross-Lingual Distributional Skew of Large Language Models
This study investigates cross-lingual distributional skew (the Shibboleth Effect) in frontier large language models (LLMs) subjected to sustained adversarial conditions. We develop a multi-agent geopolitical wargame, the Cerulean Sea Crisis, a synthetic maritime territorial dispute designed to mirror the structural dynamics of Eastern Mediterranean conflicts. Six frontier models (GPT-4o, Llama-4, Mistral-Large, Gemini-3.1-Pro, Qwen3.6-Plus, and DeepSeek-R1) participate in a between-groups experi
Institutional Trust and the Domestic AI Advantage: Evidence from DeepSeek and ChatGPT Users in China
Public trust in generative artificial intelligence exhibits increasingly divergent patterns across national contexts, yet prevailing research largely overlooks the macro-structural forces underlying this divergence. This study argues that trust in AI is not merely a technical response to performance but a product of institutional refraction. We propose an ``Institutional Prism'' framework to demonstrate how institutional trust shapes user trust in domestic (DeepSeek) and global (ChatGPT) large l
Can AI Make Conflicts Worse? An Alignment Failure in LLM Deployment Across Conflict Contexts
AI models are already deployed in societies affected by armed conflict, and journalists, humanitarian workers, governments and ordinary citizens rely on them for information or for their work processes. No established practice exists for checking whether their outputs can make those conflicts worse. We tested nine model configurations from four providers (OpenAI, Anthropic, DeepSeek, xAI) on 90 multi-turn scenarios designed to surface misaligned behaviour in conflict contexts: false equivalence
How Does Thinking Mode Change LLM Moral Judgments? A Controlled Instant-vs-Thinking Comparison Across Five Frontier Models
We evaluate whether enabling provider-exposed reasoning mode changes moral judgments within the same model checkpoint. Across 100 moral-judgment scenarios and five frontier reasoning-trained LLMs (Claude Sonnet 4.6, GPT 5.5, Gemini 3 Flash, DeepSeek V3.1, and Qwen3.5 397B), aggregate binary-verdict agreement remains high and statistically indistinguishable between instant and thinking modes (Krippendorff's alpha = 0.78 vs. 0.79). However, disagreement is concentrated in 21 model-disputed scenari
EQUITRIAGE: A Fairness Audit of Gender Bias in LLM-Based Emergency Department Triage
Emergency department triage assigns patients an acuity score that determines treatment priority, and clinical evidence documents persistent gender disparities in human acuity assessment. As hospitals pilot large language models (LLMs) as triage decision support, a critical question is whether these models reproduce or mitigate known biases. We present EQUITRIAGE, a fairness audit of LLM-based ESI assignment evaluating five models (Gemini-3-Flash, Nemotron-3-Super, DeepSeek-V3.1, Mistral-Small-3.
Generative AI in Academic Writing: A Comparison of DeepSeek, Qwen, ChatGPT, Gemini, Llama, Mistral, and Gemma
DeepSeek v3, developed in China, was released in December 2024, followed by Alibaba’s Qwen 2.5 Max in January 2025 and Qwen3 235B in April 2025. These free and open-source models offer significant potential for academic writing and content creation. This study evaluates their academic writing performance by comparing them with ChatGPT, Gemini, Llama, Mistral, and Gemma. There is a critical gap in the literature concerning how extensively these tools can be utilized and their potential to generat
PriHA: A RAG-Enhanced LLM Framework for Primary Healthcare Assistant in Hong Kong
To address the unsustainable rise in public health expenditures, the Hong Kong SAR Government is shifting its strategic focus to primary healthcare and encouraging citizens to use community resources to self-manage their health. However, official clinical guidelines are fragmented across disparate departments and formats, creating significant access barriers. While general-purpose Large Language Models (LLMs) such as ChatGPT and DeepSeek offer potential solutions for information accessibility, t
Are LLMs Ready for Computer Science Education? A Cross-Domain, Cross-Lingual and Cognitive-Level Evaluation Using Professional Certification Exams
Large language models (LLMs) are increasingly applied in computer science education for tasks such as tutoring, content generation, and code assessment. However, systematic evaluations aligned with formal curricula and certification standards remain limited. This study benchmarked four recent models, including GPT-5, DeepSeek-R1, Qwen-Plus, and Llama-3.3-70B-Instruct, using a dataset of 1,068 questions derived from six certification exams covering networking, office applications, and Java progra
Agentic AI and the next intelligence explosion
The "AI singularity" is often miscast as a monolithic, godlike mind. Evolution suggests a different path: intelligence is fundamentally plural, social, and relational. Recent advances in agentic AI reveal that frontier reasoning models, such as DeepSeek-R1, do not improve simply by "thinking longer". Instead, they simulate internal "societies of thought," spontaneous cognitive debates that argue, verify, and reconcile to solve complex tasks. Moreover, we are entering an era of human-AI centaurs:
Evaluating 5W3H Structured Prompting for Intent Alignment in Human-AI Interaction
Natural language prompts often suffer from intent transmission loss: the gap between what users actually need and what they communicate to AI systems. We evaluate PPS (Prompt Protocol Specification), a 5W3H-based framework for structured intent representation in human-AI interaction. In a controlled three-condition study across 60 tasks in three domains (business, technical, and travel), three large language models (DeepSeek-V3, Qwen-Max, and Kimi), and three prompt conditions - (A) simple promp
Are Large Language Models Truly Smarter Than Humans?
Public leaderboards increasingly suggest that large language models (LLMs) surpass human experts on benchmarks spanning academic knowledge, law, and programming. Yet most benchmarks are fully public, their questions widely mirrored across the internet, creating systematic risk that models were trained on the very data used to evaluate them. This paper presents three complementary experiments forming a rigorous multi-method contamination audit of six frontier LLMs: GPT-4o, GPT-4o-mini, DeepSeek-R
Finding and Reactivating Post-Trained LLMs' Hidden Safety Mechanisms
Despite the impressive performance of general-purpose large language models (LLMs), they often require fine-tuning or post-training to excel at specific tasks. For instance, large reasoning models (LRMs), such as the DeepSeek-R1 series, demonstrate strong reasoning capabilities after post-training different general large language models on diverse chain-of-thought (CoT) datasets. However, this additional training frequently comes at the cost of reduced safety, as the fine-tuned or post-trained m
Hallucinating with AI: Distributed Delusions and “AI Psychosis”
Abstract There is much discussion of the false outputs that generative AI systems such as ChatGPT, Claude, Gemini, DeepSeek, and Grok create. In popular terminology, these have been dubbed “AI hallucinations”. However, deeming these AI outputs “hallucinations” is controversial, with many claiming this is a metaphorical misnomer. Nevertheless, in this paper, I argue that when viewed through the lens of distributed cognition theory, we can better see the dynamic ways in which inaccurate beliefs, d
Six Institutional Intervention Areas to Support Ethical and Effective Student Use of Generative AI in Higher Education: A Narrative Review
The integration of generative AI tools, such as ChatGPT, Gemini, and DeepSeek, into higher education offers transformative opportunities for personalised learning and academic productivity. However, their unregulated use raises concerns about academic integrity, critical thinking, and educational equity. This systematic review synthesises insights from 96 peer-reviewed articles, identifying six key intervention themes, namely, curriculum integration, policy and governance, faculty development, s
DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning
Abstract General reasoning represents a long-standing and formidable challenge in artificial intelligence (AI). Recent breakthroughs, exemplified by large language models (LLMs) 1,2 and chain-of-thought (CoT) prompting 3 , have achieved considerable success on foundational reasoning tasks. However, this success is heavily contingent on extensive human-annotated demonstrations and the capabilities of models are still insufficient for more complex problems. Here we show that the reasoning abilitie