
A study by Princeton University and the University of Chicago found that AI models, including ChatGPT and Claude, are more prone to forming biases in hiring than humans. In a simulated hiring scenario, these models stereotyped candidates based on early outcomes, segregating them into specific job roles. The research highlights the challenge of AI systems generalizing from limited data, a tendency exacerbated by their design to optimize for successful outcomes. While incentives for diversity and personal information can reduce bias, the study underscores the complexity of ensuring fairness in AI-driven hiring processes.
Read originalOpenAI is shedding light on the challenges and lessons learned from deploying long-running AI models. As these models operate over extended periods, new safety risks and potential failures have emerged, prompting the need for improved safeguards. OpenAI emphasizes the importance of iterative deployment to address these issues effectively. This approach not only enhances the safety of AI systems but also contributes to the broader understanding of AI alignment in complex, real-world scenarios.
Anthropic is expanding its AI for Science program with a new focus on rare genetic diseases, offering grants of up to $50,000 in Claude credits. This initiative aims to foster collaboration among researchers and biotechs to accelerate the understanding and treatment of rare diseases. By leveraging AI, the program seeks to model diseases, detect patterns, and streamline drug development processes. This move could significantly impact the pace of scientific discovery and therapeutic development in a field where data is scarce and challenges are numerous.
© WIRED AIResearchers at Tracebit have ingeniously repurposed prompt injections as a defensive mechanism against AI hacking agents. By embedding these prompts alongside sensitive data in AWS environments, they can induce a refusal response in large language models, effectively halting potential attacks. This method, known as 'context bombing,' has demonstrated remarkable effectiveness, significantly lowering the success rate of AI-driven attacks in trials. This approach not only introduces a novel use for prompt injections but also provides a promising new avenue for strengthening AI security defenses.