I'm actually getting between a 90 and 100% fix rate on my evals so far--average is maybe 95%... the evals  are not too extensive yet so I don't know how well it will generalize to all kinds of languages and projects. i'm working on beefing up the evals but it will take some time. I think though that with right approach to retries I *might* actually be able to get gpt-4o to consistently fix 100% of its own syntax errors in my smaller set of evals at least within 3-4 attempts. I was getting more like 70-80% with gpt-4-turbo no matter how many retries.

I think this model could actually be a game-changer for reliability.

Gpt4 가 긴 텍스트의 코드를 70퍼센트의 완벽함을 보여줬는데

Gpt4o 이건 95%정도 라고함

3 4번 더돌리면 에러없는 프로그램도 가능할거라고함.

깃헙 코파일럿 가져다 버려라 오픈소스가 답이다 성능은