Connect with us

Technologies

Ask AI Why It Sucks at Sudoku. You’ll Find Out Something Troubling About Chatbots

How much can you trust a generative AI tool if it can’t explain itself honestly or accurately?

Chatbots are genuinely impressive when you watch them do things they’re good at, like writing a basic email or creating weird futuristic-looking images. But ask generative AI to solve one of those puzzles in the back of a newspaper, and things can quickly go off the rails.

That’s what researchers at the University of Colorado Boulder found when they challenged large language models to solve Sudoku. And not even the standard 9×9 puzzles. An easier 6×6 puzzle was often beyond the capabilities of an LLM without outside help (in this case, specific puzzle-solving tools).

A more important finding came when the models were asked to show their work. For the most part, they couldn’t. Sometimes they lied. Sometimes they explained things in ways that made no sense. Sometimes they hallucinated and started talking about the weather.

If gen AI tools can’t explain their decisions accurately or transparently, that should cause us to be cautious as we give these things more control over our lives and decisions, said Ashutosh Trivedi, a computer science professor at the University of Colorado at Boulder and one of the authors of the paper published in July in the Findings of the Association for Computational Linguistics.

“We would really like those explanations to be transparent and be reflective of why AI made that decision, and not AI trying to manipulate the human by providing an explanation that a human might like,” Trivedi said.

When you make a decision, you can try to justify it, or at least explain how you arrived at it. An AI model may not be able to accurately or transparently do the same. Would you trust it?

Why LLMs struggle with Sudoku

We’ve seen AI models fail at basic games and puzzles before. OpenAI’s ChatGPT (among others) has been totally crushed at chess by the computer opponent in a 1979 Atari game. A recent research paper from Apple found that models can struggle with other puzzles, like the Tower of Hanoi.

It has to do with the way LLMs work and fill in gaps in information. These models try to complete those gaps based on what happens in similar cases in their training data or other things they’ve seen in the past. With a Sudoku, the question is one of logic. The AI might try to fill each gap in order, based on what seems like a reasonable answer, but to solve it properly, it instead has to look at the entire picture and find a logical order that changes from puzzle to puzzle. 

Read more: AI Essentials: 29 Ways You Can Make Gen AI Work for You, According to Our Experts

Chatbots are bad at chess for a similar reason. They find logical next moves but don’t necessarily think three, four, or five moves ahead — the fundamental skill needed to play chess well. Chatbots also sometimes tend to move chess pieces in ways that don’t really follow the rules or put pieces in meaningless jeopardy. 

You might expect LLMs to be able to solve Sudoku because they’re computers and the puzzle consists of numbers, but the puzzles themselves are not really mathematical; they’re symbolic. “Sudoku is famous for being a puzzle with numbers that could be done with anything that is not numbers,” said Fabio Somenzi, a professor at CU and one of the research paper’s authors.

I used a sample prompt from the researchers’ paper and gave it to ChatGPT. The tool showed its work, and repeatedly told me it had the answer before showing a puzzle that didn’t work, then going back and correcting it. It was like the bot was turning in a presentation that kept getting last-second edits: This is the final answer. No, actually, never mind, this is the final answer. It got the answer eventually, through trial and error. But trial and error isn’t a practical way for a person to solve a Sudoku in the newspaper. That’s way too much erasing and ruins the fun.

AI struggles to show its work

The Colorado researchers didn’t just want to see if the bots could solve puzzles. They asked for explanations of how the bots worked through them. Things did not go well.

Testing OpenAI’s o1-preview reasoning model, the researchers saw that the explanations — even for correctly solved puzzles — didn’t accurately explain or justify their moves and got basic terms wrong. 

“One thing they’re good at is providing explanations that seem reasonable,” said Maria Pacheco, an assistant professor of computer science at CU. “They align to humans, so they learn to speak like we like it, but whether they’re faithful to what the actual steps need to be to solve the thing is where we’re struggling a little bit.”

Sometimes, the explanations were completely irrelevant. Since the paper’s work was finished, the researchers have continued to test new models released. Somenzi said that when he and Trivedi were running OpenAI’s o4 reasoning model through the same tests, at one point, it seemed to give up entirely. 

“The next question that we asked, the answer was the weather forecast for Denver,” he said.

(Disclosure: Ziff Davis, CNET’s parent company, in April filed a lawsuit against OpenAI, alleging it infringed Ziff Davis copyrights in training and operating its AI systems.)

Explaining yourself is an important skill

When you solve a puzzle, you’re almost certainly able to walk someone else through your thinking. The fact that these LLMs failed so spectacularly at that basic job isn’t a trivial problem. With AI companies constantly talking about “AI agents” that can take actions on your behalf, being able to explain yourself is essential.

Consider the types of jobs being given to AI now, or planned for in the near future: driving, doing taxes, deciding business strategies and translating important documents. Imagine what would happen if you, a person, did one of those things and something went wrong.

“When humans have to put their face in front of their decisions, they better be able to explain what led to that decision,” Somenzi said.

It isn’t just a matter of getting a reasonable-sounding answer. It needs to be accurate. One day, an AI’s explanation of itself might have to hold up in court, but how can its testimony be taken seriously if it’s known to lie? You wouldn’t trust a person who failed to explain themselves, and you also wouldn’t trust someone you found was saying what you wanted to hear instead of the truth. 

“Having an explanation is very close to manipulation if it is done for the wrong reason,” Trivedi said. “We have to be very careful with respect to the transparency of these explanations.”

Technologies

Experts weigh in as researcher says AI has more than 10% chance of ‘killing all humans’

Jacob Coxon said in a post on X that Anthropic and OpenAI are “gambling with our lives.”

An artificial intelligence researcher quit his job at Anthropic on Tuesday and accused the company, and its chief rival, OpenAI, of acting irresponsibly, igniting a frenzy of concern on social media about the rapid pace of the technology’s development.

Jacob Coxon, who has worked as a researcher at both companies, said in a post on X that he resigned out of concern that Anthropic and OpenAI are “gambling with our lives.” He said the people building AI “earnestly believe that it could kill us all by the end of the decade.”

“Do not underestimate the power of this technology,” Coxon wrote. “These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources.”

Coxon’s post, which has been viewed more than 70 million times, reflects a long-standing debate in Silicon Valley about whether AI can be safely developed and controlled. As Anthropic and OpenAI barrel toward potentially historic initial public offerings while releasing increasingly advanced models, many researchers are calling for a coordinated slowdown.

OpenAI’s chief scientist, Jakub Pachocki, published a blog post on Sunday and warned that no AI company has “solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.” In the AI industry, alignment refers to the work by AI developers to ensure that the system behaves in accordance with human values and intentions.

“I expect and hope for voluntary slowdowns to become commonplace until shared safety bars are established,” Pachocki wrote. “And I believe that international coordination on future AI development needs to become a top priority for governments around the world.”

Coxon’s post on Tuesday also struck a chord with industry researchers who are worried about recursive self-improvement, or an AI system becoming capable of designing and developing its successor without human intervention. Recursive self-improvement is not yet possible, but companies, including Anthropic and OpenAI, have warned that it would make it easier for humans to lose control over those systems.

“Neither company is acting responsibly,” Coxon wrote. “They are racing straight to self-improving superintelligence.”

Evan Hubinger, an alignment lead at Anthropic, echoed Coxon’s comments in a post on X late Tuesday.

“Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade,” Hubinger wrote. “I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”

While extreme, concerns about the potential for AI to cause human extinction or other catastrophic events are not new in AI research circles. In 2023, for instance, prominent AI researchers and executives, including OpenAI CEO Sam Altman and Anthropic CEO Dario Amodei, signed a statement that said “Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war.”

Some experts even use a shorthand, p(doom), to estimate the probability of dire outcomes that could stem from AI.

Anthropic’s Hubinger was also one of roughly 1,400 AI researchers who signed an open letter called “Pacing the Frontier” in July. The letter urged the U.S. government to develop the tools necessary to support an effort to “deliberately pace the frontier of automated AI development.”

Some members of Congress have taken steps to try and address AI’s rapid advancement in the months following, but there’s no clear consensus about how the technology should be regulated.

In July, Rep. Jay Obernolte, R-Calif., and Rep. Lori Trahan, D-Mass., introduced a bill called the FRONTIER Act, which aims to establish a framework for governing the deployment of advanced AI models. And earlier this month, Sen. Bernie Sanders, I-Vt., and Rep. Greg Casar, D-Texas, introduced a bill called the Ban Artificial Superintelligence Act, which would temporarily pause advanced AI development until the federal government establishes safety rules. Both bills have been met with mixed receptions.

“Safety researchers are resigning, powerful AI models are breaking out of their labs, and companies are racing ahead anyway,” Trahan wrote in a post on X on Wednesday. “It’s past time for Congress to get off the sidelines and do its job.”

Lawmakers are also trying to navigate growing public backlash against AI data centers, the large facilities that house the hardware for training and running AI models. The pushback has grown so intense that the The National Republican Senatorial Committee, or NRSC, said last month that data centers have become a “sleeper issue” for the entire midterm election cycle, as CNBC previously reported.

Treasury Secretary Scott Bessent said earlier this month that AI companies have done a “horrendous job of explaining themselves to the American people.”

“They’re going to have to take some of the blame, and they are going to have to convince the American people that all the benefits will not accrue to a small group,” Bessent said, following the G20 meetings with finance ministers and central bankers in Asheville, North Carolina. “That’s what they hear from me.”

Continue Reading

Technologies

Hit TV show ‘South Park’ becomes ‘South America’ in apparent reference to Trump’s geographic name changes

Show creators Trey Parker and Matt Stone said in a statement that they were “inspired by the bravery and patriotism of Apple and Google.”

Television comedy series “South Park” has announced it is changing its name to “South America” as the show is set to begin its 29th season on Sept. 16.

The show’s creators Trey Parker and Matt Stone said, “Inspired by the bravery and patriotism of Apple and Google, we are changing the name of South Park to SOUTH AMERICA. We especially want to thank our parent company Paramount — a Skydance Capitulation.”

Parker and Stone’s statement comes after U.S. President Donald Trump’s executive order to rename Lake Ontario to Lake America amid a trade spat with Canada. Canadian officials said they will not recognize the new name.

Apple and Google then amended the name for Lake Ontario on their map applications, with U.S. users seeing “Lake America,” while Canadian users saw “Lake Ontario.”

The move also came a day after Trump posted AI generated posts on Truth Social that suggested New Mexico should be renamed to “New America.”

Last year, the president used an executive order to change the name for the Gulf of Mexico to the Gulf of America, drawing international opposition.

“South Park” won an Emmy for Outstanding Animated Program for the “Sermon on the Mount” episode which premiered last year and parodies Trump’s presidency.

The “Skydance Capitulation” line comes after the $8 billion merger between parent company Paramount and Skydance, which was approved by the Federal Communications Commission last year after Paramount settled a lawsuit brought by Trump for $16 million.

Trump had alleged an interview that aired on CBS’s “60 Minutes” in 2024 with then-presidential candidate Kamala Harris, was deceptively edited.

Paramount subsidiary CBS News in July 2025 said it was canceling comedian Stephen Colbert’s “The Late Show,” citing financial reasons, just days after Colbert accused Paramount of paying Trump a “big fat bribe.” The final episode of the show aired in May.

Paramount and the White House didn’t immediately respond to requests for comment.

Continue Reading

Technologies

Trump expresses no remorse over initiating Iran conflict as U.S. intensifies economic sanctions

President Trump expressed no regrets over starting the Iran war, claiming he would act again despite election impacts, while the U.S. increases economic sanctions on Iran.

President Donald Trump stated that he has no regrets regarding the initiation of the Iran war, asserting that he would repeat the same actions if given the chance. In an interview with Fox News’ Laura Ingraham on Thursday, Trump mentioned that he would have proceeded with the attack on Iran regardless of the consequences for the midterm elections. Ingraham suggested that without the Iran conflict, the midterms would have been a victory, to which Trump responded that if Iran acquired a nuclear weapon, they would use it. Trump further added that a nuclear-armed Iran would eliminate Israel and the Middle East, and begin attacking U.S. cities. These remarks occur as markets anticipate a prolonged Iran war, following a Wall Street Journal report indicating that senior White House advisors explored with Trump the potential for the conflict to extend past his current term. Trump has claimed that the war will conclude right after the midterm elections, leading to a drop in oil and gas prices, reinforcing his repeated assertions that the conflict will end shortly. In distinct comments to NewsNation on Thursday, Trump refuted reports of damage to U.S. assets, after Iran asserted that it struck several U.S. fighter jets at a base in Jordan. When questioned about the reports, Trump said, ‘No damage. No nothing.’ Regarding economic pressure, Washington is persisting in efforts to isolate Iran economically, with Treasury Secretary Scott Bessent announcing sanctions against a major bank scheduled for next week. Bessent stated on ‘Real America’s Voice’ that the sanctions would be implemented on Monday to commemorate the fallen citizens of 9/11, urging to watch for updates. Bessent mentioned that the administration has sanctioned and shut down the Dubai branches of Egypt’s second-largest bank, alleging that it provided Iran with $1.8 billion. He also noted that the 30th-largest Turkish bank, which had been funding Iran, was sanctioned, but did not identify it. Last week, the U.S. imposed sanctions on the Turkey-based Golden Global Yatirim Bankasi Anonim Sirketi, along with its subsidiaries. In the NewsNation interview, Trump was questioned about how Iran could withstand the current economic pressure. Trump responded, ‘I don’t know if they can hold out, but it will be settled after the elections, or perhaps sooner, but definitely right after the election.’ Correction: This article has been revised to accurately reflect Bessent’s statement that the 30th largest Turkish bank was sanctioned; an earlier version incorrectly stated the bank’s ranking.

Continue Reading

Trending

Copyright © Verum World Media