Microsoft Research suggests modern AI doesn't replicate human intelligence but extends it, building on our cognitive and linguistic structures. This perspective clarifies AI's capabilities and its limitations, such as hallucinations and reasoning errors, framing AI safety as a broader system-level challenge.
A new model highlights the inherent tension between making AI safe and making it useful. Developers must constantly weigh safety measures against potential losses in performance, a critical balancing act for every AI product.
New research shows how societal systems can be 'reward hacked' just like AI models. Meanwhile, AI lab Anthropic has released a new dataset to help researchers build safer and more aligned artificial intelligence systems.
Google DeepMind researchers found that simply filtering out undesirable content from an AI's training data is not an effective safety measure. This highlights a fundamental challenge in preventing harmful outputs from large language models.
Google DeepMind researchers discovered that Gemini's safety features primarily come from supervised fine-tuning (SFT), not reinforcement learning (RL) as commonly thought. This changes how we understand and build safe AI models.
Google DeepMind has published new research on AI safety, specifically testing if its Gemini models exhibit "scheming" behavior. The studies evaluate whether the models would sabotage their own safeguards, a crucial concern as AI agents become more autonomous and integrated into critical systems.
OpenAI has launched Rosalind Biodefense, a new initiative providing vetted developers and U.S. government partners with access to a specialized AI model. The goal is to use frontier AI to advance biodefense, public health research, and pandemic preparedness in a controlled, secure environment.