Chain of thought geometric reasoning

The technique of Chain-of-Thought (CoT) prompting evaluates the ability of an AI model to break complex problems into a series of logical steps. To validate the integrity and reproducibility of the expected results, two volunteers successfully completed the exercise embedded in this paper, prior to querying several large language models (LLMs).
The LLMs were prompted to evaluate a straightedge and compass construction illustrating Definitions 3 & 4 from Euclid’s “Elements”, Book III. Two open-source models ran in a laboratory environment on a single GPU isolated from the Internet. The remaining models were tested online.
No, none of the straight lines 1, 2, 3, or 4 are at a greater distance from the center of the circle. Due to the symmetric nature of the construction.
— Claude 3.5 (The construction is actually rectangular)
Many of the LLMs correctly parsed the logic but were unable to reach a conclusion absent exact measurement. Note that Euclid’s approach in “Elements” was to emphasize general principles and logical reasoning rather than numerical calculation.
Some of the LLMs gave contradictory answers within the same response. One model produced an SVG diagram of a squared-off construction, whereas the instructions specifically stated rectangular.
The steps are easy to complete. The answer to both questions is yes. None of the LLMs successfully answered the prompt.
Understanding Euclidean Geometric Construction

Euclid is one of the most influential figures in the history of mathematics. He is known for a comprehensive description of geometry, “Elements”. Its proofs are known for their logical rigor and their systematic approach. In this essay I illustrate simple straightedge and compass construction.
Clarifying definitions by illustration
The essay in Part 1 of this series required both mathematical understanding and a high level of reading comprehension. This helped us to evaluate Artificial Intelligence (AI) Large Language Models (LLMs) on both of those challenges.
In this part we want to make geometric concepts as easy to understand as possible.
“Elements” Book III begins with a number of Definitions. However, reading through them we find there’s a lot to unpack here. Let’s try to make these easier to understand with a drawing (Figure1).
The definitions
4. In a circle straight lines are said to be equally distant from the center when the perpendiculars drawn to them from the center are equal.
5. And that straight line is said to be at a greater distance on which the greater perpendicular falls.
— Elements, Book 3 Definitions
Geometry exercise
Tools:
- Blank sheet of paper
- Straightedge
- Compass
- Red marker
- Orange marker
Using a straightedge and compass complete the following steps
- On a blank sheet of rectangular paper draw a line from the top-left corner to the bottom right corner of the page.
- Draw a line from the bottom left corner to the top right corner. The lines create an intersection at the center of the page.
- Set the compass to a width that allows a complete circle small enough to fit on the page. From the center of the page draw a circle.
- Label the point at the top-left where the line crosses the circle, point A.
Label the point at the top-right where the line crosses the circle, point B.
Label the point at the bottom-right where the line crosses the circle, point C.
Label the point at the bottom-left where the line crosses the circle, point D. - Without changing the compass, use it to mark two points on the circle, starting from point B.
Label the new point between A and B as point E.
Label the new point between B and C as point F. - Without changing the compass, use it to mark two points on the circle starting from point D.
Label the new point between C and D as point G.
Label the new point between D and A as point H. - Draw a straight line from point E to point F and label it line 1.
Draw a straight line from point G to point H and label it line 2.
Draw a straight line from point B to point G and label it line 3.
Draw a straight line from point E to point H and label it line 4. - Using the compass, bisect lines 1, 2, 3 and 4 to find the center of each line.
- From the center point of lines 1, 2, 3, and 4 draw perpendicular lines to the center point of the circle.
- For this exercise make the perpendiculars from lines 1, 2, and 3 red.
Make the perpendicular from line 4 orange.
Answer the following questions
- Are any of the straight lines 1, 2, 3, or 4 at a greater distance from the center of the circle?
- Is the orange line longer than any of the red?

Appendix A: Drawings by volunteers


Appendix B: CoT prompt 2

Appendix C: Anthropic Claude 3
“No, the orange line (perpendicular to line 4) is not longer than any of the red lines.”

The construction is squared-off, producing errors.

Appendix D: X Grok
“The perpendiculars from the center to these chords would be equal if the chords were at equal distances from the center, which they are not in this setup.” Ergo, the answer is yes and not “we can’t definitively say.”

Appendix E: OpenAI ChatGPT
“The answer here would depend on the relative lengths of these perpendiculars, which could vary depending on the exact configuration of points on your paper.” Dodges the question.

Appendix F: Google Gemini
“No. All the lines 1, 2, 3, and 4 are equidistant from the center of the circle.” Contradicts the answer to question 2, “Yes. The orange line (line 4) is longer than any of the red lines… The orange line is a diameter of the circle.”

Appendix G: Meta Llama 3.2 (open-source. lab test)
“No. based on the instructions, the orange line is not longer than any of the red lines..”

Appendix H: Microsoft Phi-3 (open source. lab test)
“The orange line (which is the perpendicular from line 4) is not necessarily longer than any of the red lines.”

Appendix I: Summary of LLM release notes and benchmarks
Anthropic Claude 3
Opus, our most intelligent model, outperforms its peers on most of the common evaluation benchmarks for AI systems, including undergraduate level expert knowledge (MMLU), graduate level expert reasoning (GPQA), basic mathematics (GSM8K), and more. It exhibits near-human levels of comprehension and fluency on complex tasks, leading the frontier of general intelligence.
GSM8K is a dataset of linguistically diverse grade school math word problems requiring multi-step reasoning.
X Grok
Both Grok-2 and Grok-2 mini demonstrate significant improvements over our previous Grok-1.5 model. They achieve performance levels competitive to other frontier models in areas such as graduate-level science knowledge (GPQA), general knowledge (MMLU, MMLU-Pro), and math competition problems (MATH). Additionally, Grok-2 excels in vision-based tasks, delivering state-of-the-art performance in visual math reasoning (MathVista) and in document-based question answering (DocVQA).
MathVista is a benchmark designed to evaluate the mathematical reasoning capabilities of AI models, particularly in visual contexts.
OpenAI ChatGPT
We’ve created GPT-4, the latest milestone in OpenAI’s effort in scaling up deep learning. GPT-4 is a large multimodal model (accepting image and text inputs, emitting text outputs) that, while less capable than humans in many real-world scenarios, exhibits human-level performance on various professional and academic benchmarks.
SAT (Scholastic Assessment Test) is a standardized test widely used for college admissions in the United States. The Math section assesses mathematical skills and problem-solving abilities, including geometry.
Google Gemini
We designed Gemini to be natively multimodal, pre-trained from the start on different modalities. Then we fine-tuned it with additional multimodal data to further refine its effectiveness. This helps Gemini seamlessly understand and reason about all kinds of inputs from the ground up.
The MATH benchmark focuses on a broad range of mathematical concepts, including algebra, geometry, and calculus.
Meta Llama 3
In the development of Llama 3, we looked at model performance on standard benchmarks and also sought to optimize for performance for real-world scenarios. To this end, we developed a new high-quality human evaluation set. This evaluation set contains 1,800 prompts that cover 12 key use cases: asking for advice, brainstorming, classification, closed question answering, coding, creative writing, extraction, inhabiting a character/persona, open question answering, reasoning, rewriting, and summarization.
The ARC Challenge (Abstraction and Reasoning Corpus) is a benchmark designed to evaluate an AI model’s ability to solve problems requiring abstract reasoning and pattern recognition.
Microsoft Phi-3
Phi-3 models significantly outperform language models of the same and larger sizes on key benchmarks.
The OpenBookQA benchmark is a dataset designed to evaluate an AI model’s understanding of core science facts and their application to novel situations.