(Credit: © patpitchaya - stock.adobe.com)
The Secret to Better Research Questions? An AI That Pushes Back
In A Nutshell
- A North Carolina State University study had 45 undergraduates use an eight-step, AI-assisted workflow to develop research questions for an ecology course.
- Instead of letting AI hand students a finished question, one key step used ChatGPT to challenge students to defend their ideas against real published research.
- Students rated every step of the process as “helped a lot” on average, and no single step stood out as more valuable than the rest.
- Nearly half the students used language in their written reflections showing they were actively reconsidering or revising their thinking.
Staring at a blank page and trying to turn a vague interest, maybe corals, maybe bees, into an actual question a scientist could test is one of the hardest, least-taught skills in research. Most students freeze right there. The tough part isn’t finding an answer, it’s figuring out the right question to ask in the first place.
A new study out of North Carolina State University tried a different approach. Instead of letting students ask a chatbot to hand them a research question, professors built an eight-step process where AI helped students explore and refine ideas, but at a key stage switched roles and made them defend those ideas against the research literature. Researchers called it the “Socratic Challenger,” a nod to the old teaching method of asking pointed questions instead of giving answers.
Forty-five undergraduates in an ecology course at North Carolina State University went through the full process and rated how much each step helped. Their answers, combined with written reflections published in the journal Frontiers in Education, offer a detailed look at how students experienced AI being folded into one of the messiest parts of research: the beginning.
The Eight-Step Workflow
Students moved through eight stages, switching back and forth between AI-assisted brainstorming and old-fashioned checking. First, they picked a broad topic on their own, no AI allowed. Then they used ChatGPT to brainstorm subtopics and search terms, followed by a hand-done literature review compared against another AI tool, NotebookLM.
Step five is where the “challenger” idea kicked in. ChatGPT was prompted to generate tough questions testing whether a student’s idea was actually new, doable, and scientifically sound. Students then had to answer those challenge questions themselves, using scholarly sources, with no AI help. After that came another round of refining the question with ChatGPT, a final manual literature check, and feedback from peers and instructors.
After grades were finalized, all 45 students agreed to let researchers study their coursework, rating how much each step helped them develop their research question on a four-point scale from no impact to dramatically enhanced. Researchers also read through the written reflections, looking for patterns in how students described the experience.
Every Step Scored Well, No Single Step Dominated
Across every single step, the typical rating landed on “helped a lot.” Students rated the literature comparison step and the closing feedback step a bit higher than earlier stages like picking a topic or refining questions with ChatGPT, but once researchers checked whether any specific step actually beat another, none of those individual differences held up. Someone who found one step useful tended to find the others useful too, as if the whole thing worked as one connected process rather than eight separate assignments.
Students Say AI Worked Best as a Thinking Partner
Written reflections filled in the picture the ratings couldn’t show. First, students described AI as a thinking partner that helped them narrow a fuzzy idea into something specific, not a tool that handed them a finished product. One student wrote that AI helped them “get into the specifics of my project rather than keeping the topic broad.” Another said AI helped them “find those knowledge gaps, and from there, I was able to critique my work in order to fill in those gaps of missing questions.”
Second, many students said the assignment’s structure, not the AI, kept them on track. Staged deadlines stopped procrastination. One student admitted plainly that if the whole project had been assigned on day one and due during finals week, “it is very likely I would not even look at it until the day before.”
Students also weighed AI against human feedback directly. One put it bluntly: “I definitely trust their advice more than I do AI services,” referring to a classmate’s review. Even the skeptics said the overall setup helped them get organized.
Researchers also scanned the reflections for first-person language signaling monitoring or revision of thinking, words like “realized,” “reconsidered,” or “revised.” Twenty-two of the 45 reflections, or 49%, contained that kind of language, suggesting nearly half the students explicitly described reconsidering, evaluating, or revising their thinking.
Structured AI Use, Not Free Access, Made the Difference
Researchers behind the study frame this as more than a report card on one ecology course. Their bigger point is about design: AI tools are fluent enough to sound convincing, and that’s exactly the problem. A chatbot can spit out a polished-sounding research question in seconds, but polish isn’t the same as evidence, feasibility, or originality. The value in this workflow appears to have come partly from refusing to let that fluency stand in for the real thing, since AI brainstorming kept getting checked against actual sources or actual people at multiple points along the way.
Students didn’t describe AI as replacing their judgment. They described it as pressure-testing their ideas, while human feedback and manual literature checks kept that pressure grounded in something real. If schools keep letting students use AI for research, this study suggests the smarter move isn’t blocking the tool or handing it free rein. It’s building a process where AI has to earn its keep by asking hard questions instead of answering them.
Paper Notes
Limitations
The authors are upfront that their findings rest on self-reported impact ratings, which measure how helpful a step felt rather than proving it actually improved the quality of students’ final research questions. Positively worded rating scales can also compress differences near the top end, and social desirability may shape how students describe their experience. The study also draws on a single undergraduate ecology course with 45 consenting students, so the results describe one classroom implementation rather than a broadly generalizable outcome. The authors note that future work should pair these perception-based findings with direct evaluation of research question quality, such as rubric-based scoring of early versus final drafts.
Funding and Disclosures
The study was supported by a Scholarship of Teaching and Learning (SoTL) Institute Mini-Grant from the NC State University Office for Faculty Excellence. The authors reported no commercial or financial conflicts of interest. One author, Aram Mikaelyan, disclosed serving as an editorial board member for Frontiers at the time of submission, which the authors state had no bearing on the peer review process or final decision. The authors also stated that generative AI was not used in writing the manuscript itself.
Publication Details
The paper is titled “The Socratic Challenger: a structured GenAI-assisted workflow for undergraduate research inquiry,” authored by Aram Mikaelyan, Erin A. McKenney, Olivia L. Mathieson, and Dhvani Toprani, affiliated with North Carolina State University and Elon University. It was published in Frontiers in Education on August 18, 2026 (Front. Educ. 11:1913451), DOI: 10.3389/feduc.2026.1913451. The study was conducted under exempt status granted by the North Carolina State University Institutional Review Board (IRB Protocol #28410).







