It's difficult for an mean personification to understand the complexity of these tasks. I'm nary mathematician, and I don't spot a quality betwixt e.g., results 3 and 10. So I had GPT-5.6 Sol Pro and Fable 5 Max categorize these utilizing
@EpochAIResearchOpenMath's rubric: — "Solid Result": A beardown interrogator successful the area would beryllium happy if their median output addressed problems of this caliber. Still, the problem would astir apt not get overmuch engagement extracurricular of its subfield. — "Major Advance": The median personification moving successful a wide area of mathematics (on the standard of number mentation aliases chart theory) would return note, and would apt make the clip to understand astatine slightest the outline of the solution. — "Breakthrough": The median mathematician would want to cognize astir this result, moreover if it was extracurricular their area. It would beryllium a campaigner for 1 of the champion results of the twelvemonth successful each of mathematics === Both Fable and Sol work together #3 is simply a Breakthrough (which explains why
@SebastienBubeckopens his tweet pinch it). They besides work together that astatine slightest 7 are Major Advancements. Fable thinks #7 is conscionable a Solid Result, while Sol assigns the "Major Advance" label. What's besides absorbing is that the charismatic Epoch.AI rubrics opportunity this: > When aggregate tiers seemed plausible for a problem, we erred successful the blimpish direction. It would beryllium disappointing to downgrade a problem’s notability aft it was solved, whereas we tin ever item immoderate unexpectedly absorbing elements of a solution. And Fable 5 thinks that astatine slightest 3 of the results are "Borderline Breakthrough" (#1, #4, and #9).
@AcerFurimmoderate thoughts connected this
English (US) ·
Indonesian (ID) ·