Anthropic has announced a new initiative under its AI for Science program, focusing on rare genetic diseases. The program offers grants of up to $50,000 in Claude credits to researchers and biotechs. The goal is to use AI to model diseases, detect patterns, and accelerate drug development. This initiative could transform how rare diseases are understood and treated, addressing the challenges of limited data and slow therapeutic development.
Read originalOpenAI is shedding light on the challenges and lessons learned from deploying long-running AI models. As these models operate over extended periods, new safety risks and potential failures have emerged, prompting the need for improved safeguards. OpenAI emphasizes the importance of iterative deployment to address these issues effectively. This approach not only enhances the safety of AI systems but also contributes to the broader understanding of AI alignment in complex, real-world scenarios.
© MIT Technology Review AIRecent research indicates that large language models (LLMs) like ChatGPT and Claude may develop biases more readily than humans in hiring scenarios. In a simulated hiring game, these models began to stereotype job applicants based on early observations, assigning candidates to roles based on perceived group traits. This tendency to generalize from limited data presents a significant challenge as AI systems gain memory and personalization capabilities. The study suggests that offering incentives for diverse hiring or providing more personal information can mitigate these biases, though the issue remains complex and unresolved. As AI systems increasingly influence decisions in hiring, loans, and parole, understanding and addressing these biases becomes crucial. The findings underscore the need for careful design and goal-setting in AI systems to prevent unintended discrimination.
© WIRED AIResearchers at Tracebit have ingeniously repurposed prompt injections as a defensive mechanism against AI hacking agents. By embedding these prompts alongside sensitive data in AWS environments, they can induce a refusal response in large language models, effectively halting potential attacks. This method, known as 'context bombing,' has demonstrated remarkable effectiveness, significantly lowering the success rate of AI-driven attacks in trials. This approach not only introduces a novel use for prompt injections but also provides a promising new avenue for strengthening AI security defenses.