Technologies
AI Is Bad at Sudoku. It’s Even Worse at Showing Its Work
Researchers did more than ask chatbots to play games. They tested whether AI models could describe their thinking. The results were troubling.
Chatbots are genuinely impressive when you watch them do things they’re good at, like writing a basic email or creating weird, futuristic-looking images. But ask generative AI to solve one of those puzzles in the back of a newspaper, and things can quickly go off the rails.
That’s what researchers at the University of Colorado at Boulder found when they challenged large language models to solve sudoku. And not even the standard 9×9 puzzles. An easier 6×6 puzzle was often beyond the capabilities of an LLM without outside help (in this case, specific puzzle-solving tools).
A more important finding came when the models were asked to show their work. For the most part, they couldn’t. Sometimes they lied. Sometimes they explained things in ways that made no sense. Sometimes they hallucinated and started talking about the weather.
If gen AI tools can’t explain their decisions accurately or transparently, that should cause us to be cautious as we give these things more control over our lives and decisions, said Ashutosh Trivedi, a computer science professor at the University of Colorado at Boulder and one of the authors of the paper published in July in the Findings of the Association for Computational Linguistics.
“We would really like those explanations to be transparent and be reflective of why AI made that decision, and not AI trying to manipulate the human by providing an explanation that a human might like,” Trivedi said.
Don’t miss any of our unbiased tech content and lab-based reviews. Add CNET as a preferred Google source.
The paper is part of a growing body of research into the behavior of large language models. Other recent studies have found, for example, that models hallucinate in part because their training procedures incentivize them to produce results a user will like, rather than what is accurate, or that people who use LLMs to help them write essays are less likely to remember what they wrote. As gen AI becomes more and more a part of our daily lives, the implications of how this technology works and how we behave when using it become hugely important.
When you make a decision, you can try to justify it, or at least explain how you arrived at it. An AI model may not be able to accurately or transparently do the same. Would you trust it?
Why LLMs struggle with sudoku
We’ve seen AI models fail at basic games and puzzles before. OpenAI’s ChatGPT (among others) has been totally crushed at chess by the computer opponent in a 1979 Atari game. A recent research paper from Apple found that models can struggle with other puzzles, like the Tower of Hanoi.
It has to do with the way LLMs work and fill in gaps in information. These models try to complete those gaps based on what happens in similar cases in their training data or other things they’ve seen in the past. With a sudoku, the question is one of logic. The AI might try to fill each gap in order, based on what seems like a reasonable answer, but to solve it properly, it instead has to look at the entire picture and find a logical order that changes from puzzle to puzzle.Â
Read more: 29 Ways You Can Make Gen AI Work for You, According to Our Experts
Chatbots are bad at chess for a similar reason. They find logical next moves but don’t necessarily think three, four or five moves ahead — the fundamental skill needed to play chess well. Chatbots also sometimes tend to move chess pieces in ways that don’t really follow the rules or put pieces in meaningless jeopardy.Â
You might expect LLMs to be able to solve sudoku because they’re computers and the puzzle consists of numbers, but the puzzles themselves are not really mathematical; they’re symbolic. “Sudoku is famous for being a puzzle with numbers that could be done with anything that is not numbers,” said Fabio Somenzi, a professor at CU and one of the research paper’s authors.
I used a sample prompt from the researchers’ paper and gave it to ChatGPT. The tool showed its work, and repeatedly told me it had the answer before showing a puzzle that didn’t work, then going back and correcting it. It was like the bot was turning in a presentation that kept getting last-second edits: This is the final answer. No, actually, never mind, this is the final answer. It got the answer eventually, through trial and error. But trial and error isn’t a practical way for a person to solve a sudoku in the newspaper. That’s way too much erasing and ruins the fun.
AI struggles to show its work
The Colorado researchers didn’t just want to see if the bots could solve puzzles. They asked for explanations of how the bots worked through them. Things did not go well.
Testing OpenAI’s o1-preview reasoning model, the researchers saw that the explanations — even for correctly solved puzzles — didn’t accurately explain or justify their moves and got basic terms wrong.Â
“One thing they’re good at is providing explanations that seem reasonable,” said Maria Pacheco, an assistant professor of computer science at CU. “They align to humans, so they learn to speak like we like it, but whether they’re faithful to what the actual steps need to be to solve the thing is where we’re struggling a little bit.”
Sometimes, the explanations were completely irrelevant. Since the paper’s work was finished, the researchers have continued to test new models released. Somenzi said that when he and Trivedi were running OpenAI’s o4 reasoning model through the same tests, at one point, it seemed to give up entirely.Â
“The next question that we asked, the answer was the weather forecast for Denver,” he said.
(Disclosure: Ziff Davis, CNET’s parent company, in April filed a lawsuit against OpenAI, alleging it infringed Ziff Davis copyrights in training and operating its AI systems.)
Explaining yourself is an important skill
When you solve a puzzle, you’re almost certainly able to walk someone else through your thinking. The fact that these LLMs failed so spectacularly at that basic job isn’t a trivial problem. With AI companies constantly talking about “AI agents” that can take actions on your behalf, being able to explain yourself is essential.
Consider the types of jobs being given to AI now, or planned for in the near future: driving, doing taxes, deciding business strategies and translating important documents. Imagine what would happen if you, a person, did one of those things and something went wrong.
“When humans have to put their face in front of their decisions, they better be able to explain what led to that decision,” Somenzi said.
It isn’t just a matter of getting a reasonable-sounding answer. It needs to be accurate. One day, an AI’s explanation of itself might have to hold up in court, but how can its testimony be taken seriously if it’s known to lie? You wouldn’t trust a person who failed to explain themselves, and you also wouldn’t trust someone you found was saying what you wanted to hear instead of the truth.Â
“Having an explanation is very close to manipulation if it is done for the wrong reason,” Trivedi said. “We have to be very careful with respect to the transparency of these explanations.”
Technologies
Venezuela awards 100-year concessions across 17 oil fields to U.S.-backed company NABEP, according to the White House
The White House announced that Venezuela has granted century-long concessions across 17 oil fields to NABEP, a U.S.-backed energy firm. The agreement marks a major expansion of American energy interests in the country.
Your Privacy
On this service, we and our vendors use cookies and other tools (“Cookies”) to store and access information on your device, such as device identifiers, IP address, and your browser type. Your data may be used to save and communicate your privacy choices; ensure security, prevent fraud, and debug our products and services; personalize advertising and content; for advertising and content measurement; to conduct audience research and services development; so we can improve our services and develop new ones; to match and combine offline data with your online activity; and for social features. We may share this data with select vendors with your consent. You can adjust your choices any time through the “Cookie Preferences” link in the footer of relevant Versant sites or in-app settings.
We will also use other Cookies that are essential or related to our services, including for security and fraud prevention. To learn more about some of our vendors, see the IAB and Google Vendors under each purpose. Visit our Cookie Notice and Privacy Policy to learn more.
These Cookies and SDKs are required for Service functionality, including security and fraud prevention, and to enable any purchasing capabilities. You can set your browser to block these tracking technologies, but some parts of the site may not function properly.
The choices you make regarding the purposes and entities listed in this notice are saved and made available to those entities in the form of digital signals (such as a string of characters). This is necessary in order to enable both this service and those entities to respect such choices.
Your data can be used to monitor for and prevent unusual and possibly fraudulent activity (for example, regarding advertising, ad clicks by bots), and ensure systems and processes work properly and securely. It can also be used to correct any problems you, the publisher or the advertiser may encounter in the delivery of content and ads and in your interaction with them.
Certain information (like an IP address or device capabilities) is used to ensure the technical compatibility of the content or advertising, and to facilitate the transmission of the content or ad to your device.
Cookies, device or similar online identifiers (e.g. login-based identifiers, randomly assigned identifiers, network based identifiers) together with other information (e.g. browser type and information, language, screen size, supported technologies etc.) can be stored or read on your device to recognise it each time it connects to an app or to a website, for one or several of the purposes presented here.
– Use limited data to select advertising 781 partners can use this purpose Advertising presented to you on this service can be based on limited data, such as the website or app you are using, your non-precise location, your device type or which content you are (or have been) interacting with (for example, to limit the number of times an ad is presented to you).
– Create profiles for personalised advertising 629 partners can use this purpose Information about your activity on this service (such as forms you submit, content you look at) can be stored and combined with other information about you (for example, information from your previous activity on this service and other websites or apps) or similar users. This is then used to build or improve a profile about you (that might include possible interests and personal aspects). Your profile can be used (also later) to present advertising that appears more relevant based on your possible interests by this and other entities.
– Use profiles to select personalised advertising 633 partners can use this purpose Advertising presented to you on this service can be based on your advertising profiles, which can reflect your activity on this service or other websites or apps (like the forms you submit, content you look at), possible interests and personal aspects.
– Create profiles to personalise content 259 partners can use this purpose Information about your activity on this service (for instance, forms you submit, non-advertising content you look at) can be stored and combined with other information about you (such as your previous activity on this service and other websites or apps) or similar users. This is then used to build or improve a profile about you (which might for example include possible interests and personal aspects). Your profile can be used (also later) to present content that appears more relevant based on your possible interests, such as by adapting the order in which content is shown to you, so that it is even easier for you to find content that matches your interests.
– Use profiles to select personalised content 231 partners can use this purpose Content presented to you on this service can be based on your content personalisation profiles, which can reflect your activity on this or other services (for instance, the forms you submit, content you look at), possible interests and personal aspects. This can for example be used to adapt the order in which content is shown to you, so that it is even easier for you to find (non-advertising) content that matches your interests.
– Measure advertising performance 905 partners can use this purpose Information regarding which advertising is presented to you and how you interact with it can be used to determine how well an advert has worked for you or other users and whether the goals of the advertising were reached. For instance, whether you saw an ad, whether you clicked on it, whether it led you to buy a product or visit a website, etc. This is very helpful to understand the relevance of advertising campaigns.
– Measure content performance 404 partners can use this purpose Information regarding which content is presented to you and how you interact with it can be used to determine whether the (non-advertising) content e.g. reached its intended audience and matched your interests. For instance, whether you read an article, watch a video, listen to a podcast or look at a product description, how long you spent on this service and the web pages you visit etc. This is very helpful to understand the relevance of (non-advertising) content that is shown to you.
– Understand audiences through statistics or combinations of data from different sources 572 partners can use this purpose Reports can be generated based on the combination of data sets (like user profiles, statistics, market research, analytics data) regarding your interactions and those of other users with advertising or (non-advertising) content to identify common characteristics (for instance, to determine which target audiences are more receptive to an ad campaign or to certain contents).
– Develop and improve services 678 partners can use this purpose Information about your activity on this service, such as your interaction with ads or content, can be very helpful to improve products and services and to build new products and services based on user interactions, the type of audience, etc. This specific purpose does not include the development or improvement of user profiles and identifiers.
– Use limited data to select content 180 partners can use this purpose Content presented to you on this service can be based on limited data, such as the website or app you are using, your non-precise location, your device type, or which content you are (or have been) interacting with (for example, to limit the number of times a video or an article is presented to you).
These Cookies and SDKs are used to collect data about your browsing habits, use of the Services, your preferences, and your interaction with advertisements across platforms and devices for the purpose of delivering targeted advertising content, both on our Services and on third party sites. Third-party sites and services also use Targeting Cookies to deliver content, including advertisements relevant to your interests on the Services. If you reject these Cookies or SDKs, you will see less relevant advertising.
Data collected under this category through Cookies and SDKs can also be used to select and deliver personalized content, such as news articles and videos.
Technologies
Tehan Calls for Return to June Agreement, Crude Climbs as Trump Pledges to Strike Iran Forcefully
Oil prices climbed Tuesday amid fears of escalation after fresh U.S.-Iran military strikes and unrest in the Strait of Hormuz, as Tehran urged Washington to return to the June interim deal and Trump promised to hit Iran hard.
Oil prices extended gains Tuesday as traders mulled the threat of escalation, following the resumption of U.S.-Iran military strikes and more turmoil in the Strait of Hormuz.
A tanker was struck by three unknown projectiles while transiting Hormuz on Monday, according to an incident report from the the U.K. Maritime Trade Operations Centre.
Brent
The tanker was sailing in the southern shipping lane close to the Omani coast, the UKMTO said, adding that no casualties were reported.
Iran launched an attack on two American bases in Jordan on Monday in retaliation for the U.S. strike on its Larak Island. American forces targeted two Iranian rocket launchers on Larak Island on Sunday, reportedly killing three, saying that Tehran intended to launch rockets carrying sea mines into the Hormuz Strait.
The small island, located in the strait, has been a critical military and shipping control point for Iranian forces, helping them keep a firm grip on vessel traffic through one of the world’s most critical maritime routes.
Iranian President Masoud Pezeshkian told the Shanghai Cooperation Organisation Summit on Tuesday that Tehran would immediately reciprocate if Washington agreed to return to its commitments under the interim deal signed in June, according to the Iranian Student News Agency.
The back-and-forth hostilities marked the first time that the U.S. and Iran traded strikes in over a month.
While neither side appeared to be seeking a return to full-scale war, both signaled they were prepared to respond to further attacks. “We are going to hit them hard,” President Donald Trump told Fox News on Monday, saying that “there will be a response” to Iran’s attacks on U.S. military bases in the region.
Analysts largely view the U.S. strike on Larak Island as an attempt to break a deadlock rather than a shift in strategy. “By hitting the launchers rather than broader Iranian military infrastructure, the U.S. appears to be punishing a specific behaviour rather than, at least for now, broadening its war aims,” said Ali Vaez, deputy program director at International Crisis Group.
“It is enforcing the blockade,” said Jason Brodsky, policy director of United Against Nuclear Iran, adding that the Trump administration’s goal was to further degrade Tehran’s capabilities to mine the Strait of Hormuz, while focusing on economic coercive measures as the midterm elections approach.
Washington has ramped up pressure to squeeze Iran’s already-torn economy with “secondary sanctions” that punish nations and businesses buying Iranian crude oil. U.S. Treasury Secretary Scott Bessent said Monday, on the sidelines of the Group of 20 finance ministers gathering, that Iran was “lashing out kinetically” because the new sanctions were taking a toll on its economy.
Speaking from the Oval Office on Monday, Trump reportedly said Iran’s financial systems, armed forces, and governing body have been largely degraded. “It doesn’t mean we won’t smack them to see what happens,” the president said.
The war, now stretching into its seventh month, has disrupted global energy supplies and sent shock waves through global financial markets.
“This is fundamentally an endurance contest,” said Brodsky, as Trump has demonstrated an “unpredictability” that should concern the Iranians, and Tehran may lash out more aggressively militarily as the economic pressure mounts.
Technologies
Fed Governor Barr Backs Rate Increase Should Inflation Fail to Cool
Fed Governor Michael Barr said he would support raising interest rates if inflation fails to ease, expressing concern that price pressures have remained elevated above the Fed’s 2% target for nearly 5½ years.
Federal Reserve Governor Michael Barr stated Tuesday that he would be willing to back an interest rate increase if inflation fails to come down.
Addressing a banking forum in Washington, the policymaker expressed concern over “broader price pressures taking hold,” noting that inflation has lingered above the Fed’s 2% target for nearly 5½ years.
“If trends in the data give me some confidence that inflation is moderating on a path to 2%, then I think we can take a bit more time to assess our policy stance,” Barr said in prepared remarks. “However, if inflation appears not to be moderating sufficiently, then I think we should act decisively to raise rates.”
The remarks arrive at a pivotal moment for policy amid a wider environment of persistent inflation and climbing Treasury yields. As a governor, Barr holds a permanent voting seat on the rate-setting Federal Open Market Committee.
Amid renewed concerns over the volatile Middle East situation, yields climbed again Tuesday, with the benchmark 10-year note reaching a level not seen since mid-January 2025.
Meanwhile, Fed Chairman Kevin Warsh last week delivered remarks that markets broadly read as leaning toward a rate hike, possibly as early as the next policy meeting in two weeks. Barr backed the July decision to hold the benchmark funds rate at 3.5%-3.75%, but markets Tuesday morning were pricing in roughly a 66% probability of an increase this month, according to the CME Group’s FedWatch.
Barr gave the economy positive marks even with inflation running high.
“Consumer spending to date has been largely resilient,” he said. “But inflation remains too high — and has been for over five years,” he added.
The latest data indicated headline prices up 3.7% over the past year, or 3.3% excluding food and energy. The Fed will receive one additional look at inflation when the consumer and producer price indexes are published next week.
-
Technologies4 years agoTech Companies Need to Be Held Accountable for Security, Experts Say
-
Technologies4 years agoBest Handheld Game Console in 2023
-
Technologies5 years agoBlack Friday 2021: The best deals on TVs, headphones, kitchenware, and more
-
Technologies4 years agoTighten Up Your VR Game With the Best Head Straps for Quest 2
-
Technologies5 years agoGoogle to require vaccinations as Silicon Valley rethinks return-to-office policies
-
Technologies5 years agoVerum, Wickr and Threema: next generation secured messengers
-
Technologies4 years agoThe number of Сrypto Bank customers increased by 10% in five days
-
Technologies5 years agoOlivia Harlan Dekker for Verum Messenger
