Archive
678 posts
-
Meta’s AI agent has been blocked from using Amazon.com
-
Higgsfield AI ships new video features in a day with GPT-6 Astra
-
Show HN: Mini-AGI – Dynamic continual learning model trained on 8GB VRAM
-
Introducing the Australian Youth Safety Blueprint
-
PrismML hopes its tiny LLM will change how we all use AI
-
Base Labs launches an open-weight AI safety partnership with Hugging Face and Goodfire
-
GLM Built Its Own Inference Infrastructure
-
Introducing Astra for Law
-
OpenSpec – A lightweight and configurable AI spec framework
-
Our framework for reporting model misalignment
-
Nvidia announces native GPU programming in Rust
-
Show HN: Capsule – Single-file web apps that save their data into SQLite
-
Salesforce and Nvidia’s new reasoning model is everything the AI labs should fear
-
Mecka AI nears $500M valuation in Sequoia-led deal amid rush for robot training data
-
Y Combinator’s Garry Tan wants US open-weight AI labs to ‘distill’ frontier models, too
-
GrapheneOS' rewritten Messages app is released
-
Introducing ChatGPT for Financial Services
-
Build more natural voice experiences with GPT‑Live‑1 in the API
-
Introducing the Agents API
-
IBM releases SOTA Granite Time Series PatchTST-FM-r2 model with commercial-friendly license
-
Instacart launches an AI grocery shopping assistant called Clementine
-
Suno replaces its AI models with a new one trained on licensed music as copyright suits pile up
-
Introducing ChatGPT Images 2.5
-
Mistral raises €3B
-
OpenAI confirms ‘wiki incident,’ says it’s ‘working on a framework’ for more disclosure
-
Ask HN: Fable hacked my piano, can I release the results?
-
Project HydraFusion: Frontier quality via multi-model orchestration
-
Show HN: TERMy – A fast terminal assistant that does not use LLMs
-
Qwen 3.8 27B available on Cerebras at 1500 tokens/s
-
OpenAI launches Astra, its powerful (and controversial) new model
-
Introducing WeatherNext 3, our most advanced and accurate global weather AI model
-
Porting my 1993 Amiga game to Godot, with an LLM reading the 68000 assembly
-
Daybreak for Frontline Defenders: $1B to protect essential services
-
NeoMME: an efficient Multimodal-native and Multilingual Encoder
-
Safety overview: GPT-6 Astra
-
Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps
-
The efficient frontier of LLM inference
-
BenchMIRT: What are LLM benchmarks actually measuring?
-
OpenAI’s Astra model is on the way — and very good at breaking into computer systems
-
The latest AI news we announced in August 2026
-
Launch HN: Nori Robotics (YC S26) – A low-cost humanoid robot for development
-
Path to Astra: critical capabilities and frontier safeguards
-
Introducing @huggingface/kernels: 200+ WebGPU Kernels for Local AI
-
Open-weight AI companies are the Valley’s hottest acquisition targets
-
GLM-5.3 is now open-weight
-
Surprise: Z.ai is the AI lab behind the mysterious Ox Alpha model
-
OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show
-
Apple introduces M6 and M5 Ultra
-
Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original
-
Jalapeño’s first results show industry-leading speed and efficiency in AI inference
-
Introducing the Admin plugin for ChatGPT Work and Codex
-
Valor, Point72 back General Intuition at $6B valuation as AI startup pushes into robotics
-
Advancing price-performance for developers with GPT‑5.6 in Kiro
-
Nvidia just showed that the harness, not the AI model, is now the real hero
-
How Hugging Face Inference Endpoints, Jobs, and Buckets Power Search on Papers with Code
-
Measuring benchmark optimization in speech recognition
-
Up to 3.2x Faster Inference with LFM2.5-DSpark
-
Ramp launches its own AI model router, called Router
-
Introducing AI Futures
-
Offering Zero Data Retention for frontier models
-
LFM2.5 Q4\_0 Checkpoints from Quantization-Aware Distillation
-
Replit expands access to software creation with GPT-5.6 Luna
-
Cursor capitalizes on GitHub frustration, launches rival hosting platform
-
Strengthening democratic oversight in national security
-
Introducing ChatGPT for Teens: Built for learning, backed by protections
-
Launch HN: Speko (YC S26) – OpenRouter for Voice AI
-
Does Mark Zuckerberg really believe AI is ‘for everyone’?
-
Meta’s ‘open’ AI, and a $250M deal gone very wrong
-
Writer introduces new AI model and upgraded harness to contain token costs
-
Gemini 3.7 Flash
-
How Fyxer built an AI executive assistant people trust
-
The builder’s guide to GPT‑5.6
-
Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed
-
Introducing OlmoEarth embeddings: Custom embedding exports from OlmoEarth Studio for downstream analysis
-
DeepSeek V4 Pro 0813
-
Putting sign language AI into users’ hands
-
Launch HN: Discovered Materials (YC P26) – AI agents to discover new materials
-
Daybreak models are now available on AWS
-
NVIDIA Nemotron 3.5 Lightning
-
As AI-led attacks multiply, OpenAI launches a new cyber model
-
Using the GitHub Copilot SDK for Java
-
Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots
-
Build Low-Latency Multilingual Voice Agents: Open Weights & Full Deployment Control with NVIDIA Magpie TTS
-
Meta’s new Glimmer AI model offers a hint at Zuckerberg’s personal intelligence vision
-
Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
-
Meta is back with Muse Glimmer: local, agentic, multimodal, and open source
-
Muse Glimmer from Meta Superintelligence Labs is now available
-
After Rippling blew millions on AI in months, it built an employee ROI tool
-
AMD acquires Taalas to boost inference performance by etching models in silicon
-
Improving GPT‑5.6 Sol in ChatGPT—and expanding access to GPT-5.6 Luna for free users
-
Baseten on Hugging Face Inference Providers 🔥
-
MacPaw taps Liquid AI to offer on-device inference to devs building for its app store
-
Open-weight AI models are catching up to the frontier. The safety gap remains.
-
Mistral's Shieldstral: 3B open-weights model for multimodal moderation
-
The latest AI news we announced in July 2026
-
Is the future of data centers portable? Runware builds a pod to find out
-
AirLLM 70B inference with single 4GB GPU
-
Circles powers telco personalization with OpenAI technology
-
Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it
-
LinkedIn adds a button to report AI-generated ‘slop’
-
Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
-
How enabling two settings tripled our scores on the ARC-AGI-3 benchmark
-
As AI content floods the internet, Pangram raises $9M to detect it
-
How GPT-5.6 fuses frontier intelligence with frontier efficiency
-
The OlmoEarth Platform: Geospatial inference at planetary scale
-
Gemini API Managed Agents: 3.6 Flash, hooks, and more
-
LFM2.5-Encoders for Fast Long-Context Inference on CPU
-
A $500 RL fine-tune of a 9B open model beat frontier models on catalog review
-
Anthropic’s Dario Amodei responds: doesn’t oppose open-weight models, but fears Chinese AI
-
Microsoft launches its first cybersecurity model, plus a new agentic cybersecurity system
-
Running a 28.9M parameter LLM on an $8 microcontroller
-
Open-weight AI is having its Kubernetes moment
-
As US weighs response to Chinese AI, industry urges against broad open-weight restrictions
-
Nvidia, Microsoft, Meta warn against overregulating open-weight models
-
Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models
-
Runway launches AI model router as generative media gets crowded
-
AI chip startup Etched defies skeptics, hits $10.3B valuation from big-name investors
-
Show HN: I simulated closing the Strait of Hormuz on real oil trade data
-
Bringing Nunchaku 4-bit Diffusion Inference to Diffusers
-
Copilot vs. raw API access: What are you actually paying for?
-
Building AI infrastructure with the Effingham County community
-
Synthesia’s AI training platform is moving beyond videos into live coaching
-
Introducing OpenAI Presence
-
Introducing the ChatGPT for small business program
-
Kimi: Threat or menace?
-
Fine-tune video and image models at scale with NVIDIA NeMo Automodel and 🤗 Diffusers
-
Introducing Gemini 3.5 Flash Cyber
-
Why the first GPU financiers are turning to inference chips in a $400 million deal
-
A scorecard for the AI age
-
The AI compute gap: Enterprises are buying infrastructure faster than they can measure what it costs
-
The agent security gap: 54% of enterprises have already had an AI agent incident, and most still let agents share credentials
-
Roblox launches an AI-powered game-creation feature in its mobile app
-
The AI context gap: Enterprise AI organizations have a trust problem, not a retrieval problem — and most are still building the fix
-
The agent evaluation gap: Enterprise AI organizations have a reality-alignment problem, not a coverage problem — and most are shipping to production anyway
-
Agentic orchestration: Enterprise AI organizations have a deployment problem, not a platform problem — and most are calling chatbots agents
-
The US is advancing AI safety through state and federal action
-
Introducing Real World VoiceEQ: Measuring the human quality of voice AI
-
Empowering India’s next generation of innovators with ATL Saathi
-
Ollama: all aboard open models
-
Separating signal from noise in coding evaluations
-
Introducing GPT-Live
-
Expanding Managed Agents in Gemini API: background tasks, remote MCP and more
-
The latest AI news we announced in June 2026
-
ScarfBench: Benchmarking AI Agents for Enterprise Java Framework Migration
-
Introducing GeneBench-Pro
-
DiScoFormer: One transformer for density and score, across distributions
-
Faster Gemma 4 on MLX with multi-token prediction
-
HP Inc. launches Frontier strategic partnership with OpenAI
-
Accelerating Transformers Fine-Tuning with NVIDIA NeMo AutoModel
-
OpenAI and Broadcom unveil LLM-optimized inference chip
-
Introducing the FFASR Leaderboard: Benchmarking ASR in the Real World
-
Experimenting with the proposed Cross-Origin Storage API in Transformers.js
-
Daybreak: Tools for securing every organization in the world
-
Patch the Planet: a Daybreak initiative to support open source maintainers
-
New usage analytics and updated spend controls for enterprises
-
Beyond LoRA: Can you beat the most popular fine-tuning technique?
-
Introducing LifeSciBench
-
Predicting model behavior before release by simulating deployment
-
Introducing the OpenAI Partner Network
-
New OpenAI Academy courses for the next era of work
-
Introducing Gemma 4 12B: a unified, encoder-free multimodal model
-
Introducing the OpenAI Economic Research Exchange
-
The latest AI news we announced in May 2026
-
Improved performance and model support with GGUF
-
Nemotron 3.5 Content Safety: Customizable Multimodal Safety for Global Enterprise AI
-
Dreaming: Better memory for a more helpful ChatGPT
-
Introducing new capabilities to GPT-Rosalind
-
A blueprint for democratic governance of frontier AI
-
Introducing Mellum2: A 12B Mixture-of-Experts Model by JetBrains
-
OpenAI frontier models and Codex are now available on AWS
-
Strengthening societal resilience with Rosalind Biodefense
-
OpenAI’s Frontier Governance Framework
-
OpenJarvis: a local-first personal AI is now available to run with Ollama
-
How Virgin Atlantic ships faster with Codex
-
Introducing OpenAI for Singapore
-
Google just redesigned the search box for the first time in 25 years — here’s why it matters more than you think.
-
Introducing the Ettin Reranker Family
-
Simulate real-world places with Project Genie and Street View
-
Databricks brings GPT-5.5 to enterprise agent workflows
-
Co-Scientist: A multi-agent AI partner to accelerate research
-
What Parameter Golf taught us about AI-assisted research
-
Building Blocks for Foundation Model Training and Inference on AWS
-
OpenAI launches DeployCo to help businesses build around intelligence
-
Advancing voice intelligence with new models in the API
-
Introducing Trusted Contact in ChatGPT
-
Introducing ChatGPT Futures: Class of 2026
-
Unlocking large scale AI training networks with MRC (Multipath Reliable Connection)
-
Enabling a new model for healthcare with AI co-clinician
-
Introducing Advanced Account Security
-
DeepInfra on Hugging Face Inference Providers 🔥
-
Introducing NVIDIA Nemotron 3 Nano Omni: Long-Context Multimodal Intelligence for Documents, Audio and Video Agents
-
OpenAI models, Codex, and Managed Agents come to AWS
-
OpenAI available at FedRAMP Moderate
-
Modeling an AI jobs transition
-
Introducing GPT-5.5
-
Speeding up agentic workflows with WebSockets in the Responses API
-
Introducing workspace agents in ChatGPT
-
Introducing OpenAI Privacy Filter
-
Introducing ChatGPT Images 2.0
-
QIMMA قِمّة ⛰: A Quality-First Arabic LLM Leaderboard
-
Scaling Codex to enterprises worldwide
-
Introducing GPT-Rosalind for life sciences research
-
Accelerating the cyber defense ecosystem that protects us all
-
Training and Finetuning Multimodal Embedding & Reranker Models with Sentence Transformers
-
Gemini 3.1 Flash TTS: the next generation of expressive AI speech
-
The next evolution of the Agents SDK
-
Trusted access for the next era of cyber defense
-
Our response to the Axios developer tool compromise
-
Multimodal Embedding & Reranker Models with Sentence Transformers
-
Introducing the Child Safety Blueprint
-
Welcome Gemma 4: Frontier multimodal intelligence on device
-
Granite 4.0 3B Vision: Compact Multimodal Intelligence for Enterprise Documents
-
Ollama is now powered by MLX on Apple Silicon in preview
-
Inside our approach to the Model Spec
-
Introducing the OpenAI Safety Bug Bounty program
-
Helping developers build safer AI experiences for teens
-
Update on the OpenAI Foundation
-
Powering product discovery in ChatGPT
-
A New Framework for Evaluating Voice Agents (EVA)
-
Measuring progress toward AGI: A cognitive framework
-
Introducing GPT-5.4 mini and nano
-
OpenAI Japan announces Japan Teen Safety Blueprint to put teen safety first
-
From model to agent: Equipping the Responses API with a computer environment
-
New ways to learn math and science in ChatGPT
-
Introducing Storage Buckets on the Hugging Face Hub
-
Introducing GPT-5.4
-
Introducing ChatGPT for Excel and new financial data integrations
-
Introducing the Adoption news channel
-
Introducing Modular Diffusers - Composable Building Blocks for Diffusion Pipelines
-
Understanding AI and learning outcomes
-
Introducing the Stateful Runtime Environment for Agents in Amazon Bedrock
-
Pacific Northwest National Laboratory and OpenAI partner to accelerate federal permitting
-
OpenAI announces Frontier Alliance Partners
-
Introducing OpenAI for India
-
Introducing EVMbench
-
Introducing Lockdown Mode and Elevated Risk labels in ChatGPT
-
Introducing GPT-5.3-Codex-Spark
-
Bringing ChatGPT to GenAI.mil
-
Transformers.js v4: Now Available on NPM!
-
Introducing SyGra Studio
-
Introducing Trusted Access for Cyber
-
Introducing OpenAI Frontier
-
Introducing GPT-5.3-Codex
-
Unlocking the Codex harness: how we built the App Server
-
Introducing the Codex app
-
Retiring GPT-4o, GPT-4.1, GPT-4.1 mini, and OpenAI o4-mini in ChatGPT
-
Introducing Daggr: Chain apps programmatically, inspect visually
-
The next chapter for AI in the EU
-
Introducing Prism
-
Unrolling the Codex agent loop
-
Railway secures $100 million to challenge AWS with AI-native cloud infrastructure
-
Introducing Edu for Countries
-
Differential Transformer V2
-
Our approach to age prediction
-
Introducing Waypoint-1: Real-time interactive video diffusion from Overworld
-
Claude Code costs up to $200 a month. Goose does the same thing for free.
-
A business that scales with the value of intelligence
-
Introducing ChatGPT Go, now available worldwide
-
Claude Code with Anthropic API compatibility
-
Strengthening the U.S. AI supply chain through domestic manufacturing
-
OpenAI Codex with Ollama
-
OpenAI partners with Cerebras
-
Zenken boosts a lean sales team with ChatGPT Enterprise
-
Salesforce rolls out new Slackbot AI agent as it battles Microsoft and Google in workplace AI
-
Anthropic launches Cowork, a Claude Desktop agent that works in your files — no coding required
-
Nous Research's NousCoder-14B is an open-source coding model landing right in the Claude Code moment
-
Introducing ChatGPT Health
-
Introducing Falcon-H1-Arabic: Pushing the Boundaries of Arabic Language AI with Hybrid Architecture
-
Announcing OpenAI Grove Cohort 2
-
AprielGuard: A Guardrail for Safety and Adversarial Robustness in Modern LLM Systems
-
Evaluating chain-of-thought monitorability
-
Deepening our collaboration with the U.S. Department of Energy
-
Introducing GPT-5.2-Codex
-
Introducing OpenAI Academy for News Organizations
-
Developers can now submit apps to ChatGPT
-
Gemma Scope 2: helping the AI safety community deepen understanding of complex language model behavior
-
Evaluating AI’s ability to perform scientific research tasks
-
Measuring AI’s capability to accelerate biological research
-
The new ChatGPT Images is here
-
BBVA and OpenAI collaborate to transform global banking
-
Introducing GPT-5.2
-
The Walt Disney Company and OpenAI reach landmark agreement to bring beloved characters to Sora
-
FACTS Benchmark Suite: Systematically evaluating the factuality of large language models
-
Introducing swift-huggingface: The Complete Swift Client for Hugging Face
-
Introducing OpenAI for Australia
-
We Got Claude to Fine-Tune an Open Source LLM
-
Announcing the initial People-First AI Fund grantees
-
Mixpanel security incident: what OpenAI users need to know
-
Expanding data residency access to business customers worldwide
-
OVHcloud on Hugging Face Inference Providers 🔥
-
Introducing shopping research in ChatGPT
-
20x Faster TRL Fine-tuning with RapidFire AI
-
Early experiments in accelerating science with GPT-5
-
Introducing AnyLanguageModel: One API for Local and Remote LLMs on Apple Platforms
-
Building more with GPT-5.1-Codex-Max
-
Introducing OpenAI for Ireland
-
SIMA 2: An Agent that Plays, Reasons, and Learns With You in Virtual 3D Worlds
-
Introducing group chats in ChatGPT
-
Introducing GPT-5.1 for developers
-
GPT-5.1: A smarter, more conversational ChatGPT
-
Introducing the Teen Safety Blueprint
-
Introducing IndQA
-
Introducing Aardvark: OpenAI’s agentic security researcher
-
Introducing gpt-oss-safeguard
-
gpt-oss-safeguard technical report
-
Doppel’s AI defense system stops attacks before they spread
-
MiniMax M2
-
MedGemma: Our most capable open models for health AI development
-
Gemini 2.5 Flash-Lite is now ready for scaled production use
-
Aeneas transforms how historians connect the past
-
Strengthening our Frontier Safety Framework
-
Introducing CodeMender: an AI agent for code security
-
Try Deep Think in the Gemini app
-
Introducing Gemma 3 270M: The compact model for hyper-efficient AI
-
VaultGemma: The world's most capable differentially private LLM
-
Consensus accelerates research with GPT-5 and Responses API
-
Building the Open Agent Ecosystem Together: Introducing OpenEnv
-
The next chapter for UK sovereign AI
-
Introducing ChatGPT Atlas, the browser with ChatGPT built in
-
Codex is now generally available
-
Introducing apps in ChatGPT and the new Apps SDK
-
AMD and OpenAI announce strategic partnership to deploy 6 gigawatts of AMD GPUs
-
Introducing AgentKit, new Evals, and RFT for agents
-
OpenAI announces strategic collaboration with Japan’s Digital Agency
-
Introducing RTEB: A New Standard for Retrieval Evaluation
-
Sora 2 System Card
-
Introducing parental controls
-
Measuring the performance of our models on real-world tasks
-
Introducing ChatGPT Pulse
-
Web search
-
New model scheduling
-
SyGra: The One-Stop Framework for Building Data for LLMs and SLMs
-
Scaleway on Hugging Face Inference Providers 🔥
-
Public AI on Hugging Face Inference Providers 🔥
-
Introducing Stargate UK
-
Introducing upgrades to Codex
-
Addendum to GPT-5 system card: GPT-5-Codex
-
Introducing the Palmyra-mini family: Powerful, lightweight, and ready to reason!
-
Fine-tune Any LLM from the Hugging Face Hub with Together AI
-
OpenAI and Greek Government launch ‘OpenAI for Greece’
-
Introducing gpt-realtime and Realtime API updates
-
Supporting nonprofit and community innovation
-
NVIDIA Releases 6 Million Multi-Lingual Reasoning Dataset
-
Introducing AI Sheets: a tool to work with datasets using open AI models!
-
Introducing GPT-5 for developers
-
Introducing GPT-5
-
gpt-oss-120b & gpt-oss-20b Model Card
-
Introducing gpt-oss
-
Open Weights and AI for All
-
Estimating worst case frontier risks of open weight LLMs
-
📚 3LM: A Benchmark for Arabic LLMs in STEM and Code
-
Introducing Stargate Norway
-
Introducing study mode in ChatGPT
-
Introducing Trackio: A Lightweight Experiment Tracking Library from Hugging Face
-
TimeScope: How Long Can Your Video Large Multimodal Model Go?
-
Fast LoRA inference for Flux with Diffusers and PEFT
-
OpenAI’s new economic analysis
-
ChatGPT agent System Card
-
Introducing ChatGPT agent
-
Asynchronous Robot Inference: Decoupling Action Prediction and Execution
-
Efficient MultiModal Data Pipeline
-
No-code personal agents, powered by GPT-4.1 and Realtime API
-
(LoRA) Fine-Tuning FLUX.1-dev on Consumer Hardware
-
Toward understanding and preventing misalignment generalization
-
Introducing OpenAI for Government
-
Groq on Hugging Face Inference Providers 🔥
-
How Long Prompts Block Other Requests - Optimizing LLM Performance
-
Featherless AI on Hugging Face Inference Providers 🔥
-
Introducing Training Cluster as a Service - a new collaboration with NVIDIA
-
Scaling security with responsible disclosure
-
How we’re responding to The New York Times’ data demands in order to protect user privacy
-
Addendum to OpenAI o3 and o4-mini system card: OpenAI o3 Operator
-
Introducing Stargate UAE
-
New tools and features in the Responses API
-
Exploring Quantization Backends in Diffusers
-
Introducing Codex
-
Ollama's new engine for multimodal models
-
Blazingly fast whisper transcriptions with Inference Endpoints
-
Introducing HealthBench
-
Introducing data residency in Asia
-
Introducing OpenAI for Countries
-
Introducing AI stories: daily benefits shine a light on bigger opportunities
-
Introducing AutoRound: Intel’s Advanced Quantization for LLMs and VLMs
-
Introducing our latest image generation model in the API
-
Prefill and Decode for Concurrent Requests - Optimizing LLM Performance
-
Introducing OpenAI o3 and o4-mini
-
Cohere on Hugging Face Inference Providers 🔥
-
Introducing HELMET: Holistically Evaluating Long-context Language Models
-
OpenAI announces nonprofit commission advisors
-
Our updated Preparedness Framework
-
Introducing GPT-4.1 in the API
-
Visual Salamandra: Pushing the Boundaries of Multimodal Understanding
-
BrowseComp: a benchmark for browsing agents
-
Arabic Leaderboards: Introducing Arabic Instruction Following, Updating AraGen, and More
-
The NLP Course is becoming the LLM Course
-
Efficient Request Queueing – Optimizing LLM Performance
-
PaperBench: Evaluating AI’s Ability to Replicate AI Research
-
🚀 Accelerating LLM Inference with TGI on Intel Gaudi
-
Introducing 4o Image Generation
-
Introducing Gradio's new Dataframe!
-
The New and Fresh analytics in Inference Endpoints
-
Introducing next-generation audio models in the API
-
New in ChatGPT for Business: March 2025
-
Welcome Gemma 3: Google's all new multimodal, multilingual, long context open LLM
-
Detecting misbehavior in frontier reasoning models
-
LLM Inference on Edge: A Fun and Easy Guide to run LLMs via React Native on your Phone!
-
Introducing NextGenAI
-
Deep research System Card
-
Minions: where local and cloud LLMs meet
-
Remote VAEs for decoding with Inference Endpoints 🤗
-
Introducing the SWE-Lancer benchmark
-
Introducing Three New Serverless Inference Providers: Hyperbolic, Nebius AI Studio, and Novita 🔥
-
Fixing Open LLM Leaderboard with Math-Verify
-
The Open Arabic LLM Leaderboard 2
-
Introducing the Intelligence Age
-
Introducing data residency in Europe
-
DABStep: Data Agent Benchmark for Multi-step Reasoning
-
Introducing deep research
-
OpenAI o3-mini System Card
-
How to deploy and fine-tune DeepSeek models on AWS
-
Welcome to Inference Providers on the Hub 🔥
-
Introducing Operator
-
SmolVLM Grows Smaller – Introducing the 256M & 500M Models!
-
Trading inference-time compute for adversarial robustness
-
Introducing multi-backends (TRT-LLM, vLLM) support for Text Generation Inference
-
CO₂ Emissions and Models Performance: Insights from the Open LLM Leaderboard
-
Introducing smolagents: simple agents that write actions in code.
-
Deliberative alignment: reasoning enables safer language models
-
Finally, a Replacement for BERT: Introducing ModernBERT
-
Bamba: Inference-Efficient Hybrid Mamba2 Model
-
OpenAI o1 and new tools for developers
-
Introducing the Synthetic Data Generator - Build Datasets with Natural Language
-
Sora is here
-
Introducing ChatGPT Pro
-
OpenAI o1 System Card
-
OpenAI and Future partner on specialist content
-
Rethinking LLM Evaluation with 3C3H: AraGen Benchmark and Leaderboard
-
Investing in Performance: Fine-tune small models with LLM insights - a CFM case study
-
Building smarter maps with GPT-4o vision fine-tuning
-
Letting Large Models Debate: The First Multilingual LLM Debate Competition
-
Introducing the Open Leaderboard for Japanese LLMs!
-
Rox goes “all in” on OpenAI
-
Argilla 2.4: Easily Build Fine-Tuning and Evaluation Datasets on the Hub — No Code Required
-
Introducing ChatGPT search
-
Introducing SimpleQA
-
Expert Support case study: Bolstering a RAG app with LLM-as-a-Judge
-
Introducing SynthID Text
-
Introducing HUGS - Scale your AI with Open Models
-
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
-
Introducing the AMD 5th Gen EPYC™ CPU
-
Introducing the Open FinLLM Leaderboard
-
Introducing canvas, a new way to write and code with ChatGPT.
-
Introducing the Realtime API
-
Introducing vision to the fine-tuning API
-
Prompt Caching in the API
-
Model Distillation in the API
-
🇨🇿 BenCzechMark - Can your LLM Understand Czech?
-
Upgrading the Moderation API with our new multimodal moderation model
-
Llama 3.2 goes small and multimodal
-
Introducing Verdi, an AI dev platform powered by GPT-4o
-
Genmab launches “AI Everywhere”
-
Fine-tuning LLMs to 1.58bit: extreme quantization made easy
-
Reduce hallucinations with Bespoke-Minicheck
-
Introducing the SQL Console on Datasets
-
Introducing Community Tools on HuggingChat
-
Introducing OpenAI o1
-
Fine-tuning now available for GPT-4o
-
Introducing SWE-bench Verified
-
Introducing Structured Outputs in the API
-
Introducing TextImage Augmentation for Document Images
-
Google releases Gemma 2 2B, ShieldGemma and Gemma Scope
-
Serverless Inference with Hugging Face and NVIDIA NIM
-
LAVE: Zero-shot VQA Evaluation on Docmatix with LLMs - Do We Still Need Fine-Tuning?
-
New compliance and administrative tools for ChatGPT Enterprise
-
Our Transformers Code Agent beats the GAIA benchmark 🏅
-
Welcome Gemma 2 - Google’s new open LLM
-
XLSCOUT Unveils ParaEmbed 2.0: a Powerful Embedding Model Tailored for Patents and IP with Expert Support from Hugging Face
-
Fine-tuning Florence-2 - Microsoft's Cutting-edge Vision Language Models
-
Going multimodal: How Prezi is leveraging the Hub and the Expert Support Program to accelerate their ML roadmap
-
Introducing the Hugging Face Embedding Container for Amazon SageMaker
-
Introducing NPC-Playground, a 3D playground to interact with LLM-powered NPCs
-
Introducing OpenAI for Nonprofits
-
Automating customer support agents
-
Benchmarking Text Generation Inference
-
CyberSecEval 2 - A Comprehensive Evaluation Framework for Cybersecurity Risks and Capabilities of Large Language Models
-
Introducing Spaces Dev Mode for a seamless developer experience
-
Google announces Firebase Genkit with Ollama support
-
Unlocking Longer Generation with Key-Value Cache Quantization
-
Ilya Sutskever to leave OpenAI, Jakub Pachocki announced as Chief Scientist
-
Introducing the Open Arabic LLM Leaderboard
-
Introducing GPT-4o and more tools to ChatGPT free users
-
Spring Update
-
License to Call: Introducing Transformers Agents 2.0
-
Introducing the Model Spec
-
Understanding the source of what we see and hear online
-
API Partnership with Stack Overflow
-
Introducing the Open Leaderboard for Hebrew LLMs!
-
Bringing the Artificial Analysis LLM Performance Leaderboard to Hugging Face
-
Powerful ASR + diarization + speculative decoding with Hugging Face Inference Endpoints
-
GPT-4 API general availability and deprecation of older models in the Completions API
-
Introducing ChatGPT and Whisper APIs
-
Introducing more enterprise-grade features for API customers
-
Introducing the Open Chain of Thought Leaderboard
-
Jack of All Trades, Master of Some, a Multi-Purpose Transformer Agent
-
The Open Medical-LLM Leaderboard: Benchmarking Large Language Models in Healthcare
-
Welcome Llama 3 - Meta's new open LLM
-
Llama 3
-
Introducing the LiveCodeBench Leaderboard - Holistic and Contamination-Free Evaluation of Code LLMs
-
Introducing Idefics2: A Powerful 8B Vision-Language Model for the community
-
Introducing OpenAI Japan
-
Introducing improvements to the fine-tuning API and expanding our custom models program
-
Text2SQL using Hugging Face Dataset Viewer API and Motherduck DuckDB-NSQL-7B
-
Blazing Fast SetFit Inference with 🤗 Optimum Intel on Xeon
-
Bringing serverless GPU inference to Hugging Face users
-
Binary and Scalar Embedding Quantization for Significantly Faster & Cheaper Retrieval
-
Embedding AI into developer software
-
Introducing the Chatbot Guardrails Arena
-
Reimagining the email experience with AI
-
Building a data-driven, efficient culture with AI
-
Quanto: a PyTorch quantization backend for Optimum
-
OpenAI announces new members to board of directors
-
Using AI to improve patient access to clinical trials
-
Introducing ConTextual: How well can your Multimodal model jointly reason over text and image in text-rich scenes?
-
Fine-Tuning Gemma Models in Hugging Face
-
Introducing the Red-Teaming Resistance Leaderboard
-
Welcome Gemma - Google’s new open LLM
-
Introducing the Open Ko-LLM Leaderboard: Leading the Korean LLM Evaluation Ecosystem
-
Video generation models as world simulators
-
Windows preview
-
From OpenAI to Open LLMs with Messages API on Hugging Face
-
OpenAI compatibility
-
Hugging Face Text Generation Inference available for AWS Inferentia2
-
Patch Time Series Transformer in Hugging Face
-
Building an early warning system for LLM-aided biological threat creation
-
Introducing the Enterprise Scenarios Leaderboard: a Leaderboard for Real World Use Cases
-
An Introduction to AI Secure LLM Safety Leaderboard
-
New embedding models and API updates
-
Python & JavaScript Libraries
-
Fine-Tune W2V2-Bert for low-resource ASR with 🤗 Transformers
-
Accelerating SD Turbo and SDXL Turbo Inference with ONNX Runtime and Olive
-
Introducing ChatGPT Team
-
Introducing the GPT Store
-
Make LLM Fine-tuning 2x faster with Unsloth and 🤗 TRL
-
Delivering LLM-powered health solutions
-
Speculative Decoding for 2x Faster Whisper Inference
-
Optimum-NVIDIA Unlocking blazingly fast LLM inference in just 1 line of code
-
Goodbye cold boot - how we made LoRA Inference 300% faster
-
Open LLM Leaderboard: DROP deep dive
-
OpenAI announces leadership transition
-
Introducing Prodigy-HF: a direct integration with Hugging Face
-
Introducing GPTs
-
New models and developer products announced at DevDay
-
Introducing Storage Regions on the HF Hub
-
Deploy Embedding Models with Hugging Face Inference Endpoints
-
DALL·E 3 is now available in ChatGPT Plus and Enterprise
-
Building LLM-Powered Web Apps with Client-Side Technology
-
Ollama is now available as an official Docker image
-
🧨 Accelerating Stable Diffusion XL Inference with JAX on Cloud TPU v5e
-
Deploying the AI Comic Factory using the Inference API
-
Llama 2 on Amazon SageMaker a Benchmark
-
Inference for PROs
-
Leveraging LLMs in your Obsidian Notes
-
Optimizing your LLM in production
-
Introducing OpenAI Dublin
-
Introducing Würstchen: Fast Diffusion for Image Generation
-
Fine-tuning Llama 2 70B using PyTorch FSDP
-
Overview of natively supported quantization schemes in 🤗 Transformers
-
Introducing ChatGPT Enterprise
-
OpenAI partners with Scale to provide support for enterprises fine-tuning models
-
GPT-3.5 Turbo fine-tuning and API updates
-
Introducing SafeCoder
-
Introducing IDEFICS: An Open Reproduction of State-of-the-art Visual Langage Model
-
Fine-tune Llama 2 with DPO
-
Deploy MusicGen in no time with Inference Endpoints
-
Stable Diffusion XL on Mac with Advanced Core ML Quantization
-
Introducing Agents.js: Give tools to your LLMs using JavaScript
-
Custom instructions for ChatGPT
-
Open-Source Text Generation & LLM Ecosystem at Hugging Face
-
Fine-tuning Stable Diffusion models on Intel CPUs
-
Deploy LLMs with Hugging Face Inference Endpoints
-
Introducing OpenAI London
-
What's going on with the Open LLM Leaderboard?
-
Fine-Tune MMS Adapter Models for low-resource ASR
-
Function calling and other API updates
-
Introducing the Hugging Face LLM Inference Container for Amazon SageMaker
-
Introducing BERTopic Integration with the Hugging Face Hub
-
Making LLMs even more accessible with bitsandbytes, 4-bit quantization and QLoRA
-
Introducing the ChatGPT app for iOS
-
Introducing RWKV - An RNN with the advantages of a transformer
-
StarCoder: A State-of-the-Art LLM for Code
-
How to Install and Use the Hugging Face Unity API
-
Introducing HuggingFace blog for Chinese speakers: Fostering Collaboration with the Chinese AI community
-
Fast Inference on Large Language Models: BLOOMZ on Habana Gaudi2 Accelerator
-
Accelerating Stable Diffusion Inference on Intel CPUs
-
GPT-4
-
Fine-tuning 20B LLMs with RLHF on a 24GB consumer GPU
-
Why we’re switching to Hugging Face Inference Endpoints, and maybe you should too
-
Parameter-Efficient Fine-Tuning using 🤗 PEFT
-
Introducing ⚔️ AI vs. AI ⚔️ a deep reinforcement learning multi-agents competition system
-
Introducing ChatGPT Plus
-
Using LoRA for Efficient Stable Diffusion Fine-Tuning
-
Forecasting potential misuses of language models for disinformation campaigns and how to reduce risk
-
Fine-tuning GPT-3 to scale video creation
-
Faster Training and Inference: Habana Gaudi®2 vs Nvidia A100 80GB
-
Introducing ChatGPT
-
An overview of inference solutions on Hugging Face
-
Introducing our new pricing
-
DALL·E API now available in public beta
-
Fine-Tune Whisper For Multilingual ASR with 🤗 Transformers
-
MTEB: Massive Text Embedding Benchmark
-
Getting Started with Hugging Face Inference Endpoints
-
Optimization story: Bloom inference
-
Introducing DOI: the Digital Object Identifier to Datasets and Models
-
DALL·E now available without waitlist
-
Introducing Whisper
-
Incredibly Fast BLOOM Inference with DeepSpeed and Accelerate
-
Train your first Decision Transformer
-
DALL·E: Introducing outpainting
-
Introducing Skops
-
New and improved content moderation tooling
-
Train and Fine-Tune Sentence Transformers Models
-
Introducing the Private Hub: A New Way to Build With Machine Learning
-
Introducing new audio and vision documentation in 🤗 Datasets
-
A hazard analysis framework for code synthesis large language models
-
DALL·E now available in beta
-
Introducing The World's Largest Open Multilingual Language Model: BLOOM
-
Learning to play Minecraft with Video PreTraining
-
The Annotated Diffusion Model
-
Introducing Pull Requests and Discussions 🥳
-
Powering next generation applications with OpenAI Codex
-
Accelerated Inference with Optimum and Transformers Pipelines
-
Introducing Hugging Face for Education 🤗
-
Habana Labs and Hugging Face Partner to Accelerate Transformer Model Training
-
Introducing Decision Transformers on Hugging Face 🤗
-
Fine-Tune a Semantic Segmentation Model with a Custom Dataset
-
Accelerate BERT inference with Hugging Face Transformers and AWS Inferentia
-
New GPT-3 capabilities: Edit & insert
-
Fine-Tune ViT for Image Classification with 🤗 Transformers
-
Introducing text and code embeddings
-
Deploy GPT-J 6B for inference using Hugging Face Transformers and Amazon SageMaker
-
Customizing GPT-3 for your application
-
Introducing Snowball Fight ☃️, our first ML-Agents environment
-
Introducing the Data Measurements Tool: an Interactive Tool for Looking at Datasets
-
Accelerating PyTorch distributed fine-tuning with Intel technologies
-
OpenAI’s API now available with no waitlist
-
Fine-Tune XLSR-Wav2Vec2 for low-resource ASR with 🤗 Transformers
-
Scaling up BERT-like model Inference on modern CPU - Part 2
-
Introducing Optimum: The Optimization Toolkit for Transformers at Scale
-
OpenAI Codex
-
Introducing Triton: Open-source GPU programming for neural networks
-
Improving language model behavior by training on a curated dataset
-
Few-shot learning in practice: GPT-Neo and the 🤗 Accelerated Inference API
-
Scaling-up BERT Inference on CPU (Part 1)
-
Introducing 🤗 Accelerate
-
GPT-3 powers the next generation of apps
-
Fine-Tune Wav2Vec2 for English ASR in Hugging Face with 🤗 Transformers
-
Multimodal neurons in artificial neural networks
-
How we sped up transformer inference 100x for 🤗 API customers
-
CLIP: Connecting text and images
-
Transformer-based Encoder-Decoder Models
-
Procgen and MineRL Competitions
-
Image GPT
-
OpenAI API
-
Jukebox
-
OpenAI Microscope
-
OpenAI standardizes on PyTorch
-
Procgen Benchmark
-
GPT-2: 1.5B release
-
Fine-tuning GPT-2 from human preferences
-
MuseNet
-
Generative modeling with sparse transformers
-
Introducing Activation Atlases
-
OpenAI Five Benchmark: Results
-
OpenAI Five Benchmark
-
Gym Retro
-
Gotta Learn Fast: A new benchmark for generalization in RL
-
Machine Learning Unconference
-
Introducing OpenAI