To more fully understand how private sector actors can send costly signals, it is worth considering two examples of leading AI companies going beyond public statements to signal their commitment to develop AI responsibly: OpenAI’s publication of a “system card” alongside the launch of its GPT-4 model, and Anthropic’s decision to delay the release of its chatbot, Claude. Both of these examples come from companies developing LLMs, the type of AI system that burst into the spotlight with OpenAI’s release of ChatGPT in November 2022.147 LLMs are distinctive in that, unlike most AI systems, they do not serve a single specific function. They are designed to predict the next word in a text, which has proven to be useful for tasks as varied as translation, programming, summarization, and writing poetry. This versatility makes them useful, but also makes it more challenging to understand and mitigate the risks posed by a given LLM, such as fabricating information, perpetuating bias, producing abusive content, or lowering the barriers to dangerous activities.
In March 2023, California-based OpenAI released the latest iteration in their series of LLMs. Named GPT-4 (with GPT standing for “generative pre-trained transformer,” a phrase that describes how the LLM was built), the new model demonstrated impressive performance across a range of tasks, including setting new records on several benchmarks designed to test language understanding in LLMs. **From a signaling perspective, however, the most interesting part of the GPT-4 release was not the technical report detailing its capabilities, but the 60-page so-called “system card” laying out safety challenges posed by the model and mitigation strategies that OpenAI had implemented prior to the release.** 148
The system card provides evidence of several kinds of costs that OpenAI was willing to bear in order to release GPT-4 safely. These include the time and financial cost of producing the system card as well as the possible reputational cost of disclosing that the company is aware of the many undesirable behaviors of its model. The document states that OpenAI spent six months on “safety research, risk assessment, and iteration” between the development of an initial version of GPT-4 and the eventual release. Researchers at the company used this time to carry out a wide range of tests and evaluations on the model, including engaging external experts to assess its capabilities in areas that pose safety risks. These external “red teamers” probed GPT-4’s ability to assist users with undesirable activities, such as carrying out cyberattacks, producing chemical or biological weapons, or making plans to harm themselves or others. They also investigated the extent to which the model could pose risks of its own accord, for instance through the ability to replicate and acquire resources autonomously. The system card documents a range of strategies OpenAI used to mitigate risks identified during this process, with before-and-after examples showing how these mitigations resulted in less risky behavior. It also describes several issues that they were not able to mitigate fully before GPT-4’s release, such as vulnerability to adversarial examples.
Returning to our framework of costly signals, OpenAI’s decision to create and publish the GPT4 system card could be considered an example of tying hands as well as reducible costs. By publishing such a thorough, frank assessment of its model’s shortcomings, OpenAI has to some extent tied its own hands—creating an expectation that the company will produce and publish similar risk assessments for major new releases in the future. OpenAI also paid a price in terms of foregone revenue from the period in which the company could have launched GPT-4 sooner. These costs are reducible in as much as OpenAI is able to end up with greater market share by credibly demonstrating its commitment to developing safe and trustworthy systems. As explored above, the types of costs in question for OpenAI as a commercial actor differ somewhat from those that might be paid by states or other actors.
While the system card itself has been well received among researchers interested in understanding GPT-4’s risk profile, it appears to have been less successful as a broader signal of OpenAI’s commitment to safety. The reason for this unintended outcome is that the company took other actions that overshadowed the import of the system card: most notably, the blockbuster release of ChatGPT four months earlier. Intended as a relatively inconspicuous “research preview,” the original ChatGPT was built using a less advanced LLM called GPT-3.5, which was already in widespread use by other OpenAI customers. GPT-3.5’s prior circulation is presumably why OpenAI did not feel the need to perform or publish such detailed safety testing in this instance. **Nonetheless, one major effect of ChatGPT’s release was to spark a sense of urgency inside major tech companies.** 149 **To avoid falling behind OpenAI amid the wave of customer enthusiasm about chatbots, competitors sought to accelerate or circumvent internal safety and ethics review processes, with Google creating a fast-track “green lane” to allow products to be released more quickly.** 150 **This result seems strikingly similar to the raceto-the-bottom dynamics that OpenAI and others have stated that they wish to avoid. OpenAI has also drawn criticism for many other safety and ethics issues related to the launches of ChatGPT and GPT-4, including regarding copyright issues, labor conditions for data annotators, and the susceptibility of their products to “jailbreaks” that allow users to bypass safety controls.** 151 This muddled overall picture provides an example of how the messages sent by deliberate signals can be overshadowed by actions that were not designed to reveal intent.
**A different approach to signaling in the private sector comes from Anthropic, one of OpenAI’s primary competitors. Anthropic’s desire to be perceived as a company that values safety shines through across its communications, beginning from its tagline: “an AI safety and research company.”** 152 **A careful look at the company’s decision-making reveals that this commitment goes beyond words. A March 2023 strategy document published on Anthropic’s website revealed that the release of Anthropic’s chatbot Claude, a competitor to ChatGPT, had been deliberately delayed in order to avoid “advanc[ing] the rate of AI capabilities progress.”** 153 The decision to begin sharing Claude with users in early 2023 was made “now that the gap between it and the public state of the art is smaller,” according to the document—a clear reference to the release of ChatGPT several weeks before Claude entered beta testing. In other words, Anthropic had deliberately decided not to productize its technology in order to avoid stoking the flames of AI hype. Once a similar product (ChatGPT) was released by another company, this reason not to release Claude was obviated, so Anthropic began offering beta access to test users before officially releasing Claude as a product in March.
Anthropic’s decision represents an alternate strategy for reducing “race-to-the-bottom” dynamics on AI safety. Where the GPT-4 system card acted as a costly signal of OpenAI’s emphasis on building safe systems, Anthropic’s decision to keep their product off the market was instead a costly signal of restraint. By delaying the release of Claude until another company put out a similarly capable product, Anthropic was showing its willingness to avoid exactly the kind of frantic corner-cutting that the release of ChatGPT appeared to spur. Anthropic achieved this goal by leveraging installment costs, or fixed costs that cannot be offset over time. In the framework of this study, Anthropic enhanced the credibility of its commitments to AI safety by holding its model back from early release and absorbing potential future revenue losses. The motivation in this case was not to recoup those losses by gaining a wider market share, but rather to promote industry norms and contribute to shared expectations around responsible AI development and deployment.
Yet where OpenAI’s attempt at signaling may have been drowned out by other, even more conspicuous actions taken by the company, Anthropic’s signal may have simply failed to cut through the noise. By burying the explanation of Claude’s delayed release in the middle of a long, detailed document posted to the company’s website, Anthropic appears to have ensured that this signal of its intentions around AI safety has gone largely unnoticed. Taken together, these two case studies therefore provide further evidence that signaling around AI may be even more complex than signaling in previous eras.
번역 해오자
https://www.reddit.com/r/OpenAI/comments/180y6pn/the_publication_that_ignited_the_feud_between_sam/