1. Conversational style – naturalness vs. verbosity/hedging
Many commenters praise Gemini’s smooth, agreeable tone for casual use, while others criticize excessive hedging or verbosity in models like Claude/Astra.
- “For heavyweight work I have been using Astra, but for rabbit holes and brain storming Gemini is far more enjoyable to interact with.” – WarmWash
- “Mostly because it answers quickly and is more agreeable (too agreeable at times).” – drivebyhooting
- “Sometimes that's what being smart sounds like.” – mapontosevenths (defending hedging)
- “I consider, on the contrary, caveating and hedging annoying 'typical redditor'/internet behaviors…” – JW_00000 (critiquing hedging)
2. Coding ability – Gemini seen as weak, Astra/Fable preferred
Several users note that Gemini lags behind dedicated coding models, with Astra or Fable often chosen for serious programming tasks.
- “Astra might as well be an alien super intelligence at code compared to Gemini.” – aventured
- “For coding models? I don't think Google is motivated to fight in that market. There's no incentive for them.” – cmrdporcupine
- “I think 3.7 Flash is very good at coding… Opus is only slightly better IMO and it's much more frustrating to read.” – asdfman123
3. Strength in language, translation, voice, and everyday tasks
Gemini is frequently highlighted for its fluency in multilingual conversation, live voice interaction, and utility in apps like Google Maps.
- “I've been using Gemini to live chat in Afrikaans… It is phenomenal at speaking the language.” – jeanbza
- “Try Gemini live in a multi‑lingual environment. It can pick out speakers and live translate to you.” – schainks
- “I love Gemini in Google Maps for long drives… I start grilling it with random questions I've always wondered about.” – asdfman123
- “Gemini is great for querying text in another language, it provides really cogent, useful responses…” – nutjob2
Commenters argue Google prioritizes broad, low‑cost AI for search and consumer products rather than chasing state‑of‑the‑art coding benchmarks.
- “Google is clearly motivated to make better search and information finding tools… Selling coding plans isn't something they are going to want to do.” – cmrdporcupine
- “As an 'everyday mans AI' I'd say 3.8 Flash definitely already has. Smart enough for the vast swath of people…” – WarmWash
- “Considering how far backwards Google has gone… I would say they're more likely to drop trying to compete at the frontier…” – smcleod
- “Even Googlers internally do not have access to top‑tier models…” – mlmonkey
🚀 Project Ideas
Generating project ideas…
We need to read discussion to identify pain points, frustrations, unmet needs.
Let's parse conversation: Many comments about Gemini models vs others. Pain points:
Gemini's speech-to-speech live mode: good for natural conversation, but lacking in reasoning, hallucinations, poor for coding, sometimes invents extra requirements, etc.
Users want better coding ability, but also enjoy Gemini's natural language, especially for translation and language learning.
Need for ability to use Gemini as a translator or language practice, especially for niche languages (Afrikaans, Catalan) and for live translation.
Need for integration with Google Maps/voice for Q&A while driving, but issues with mic cutoff, hallucinations.
Need for ability to use Gemini as a "rubber duck" for design but not get code suggestions; they'd like a mode that suppresses implementation advice.
Need for ability to configure verbosity, less caveats, less hedging.
Need for ability to disable permissions prompts (like antigravity's dangerous skip permissions). Many users complain about Claude's verbose caveats and need to filter.
Need for ability to use Gemini for live voice interaction with existing chess or other apps? Not strong.
Need for a tool that can let you use Gemini as a "translator" for other model outputs (like Claude) to make them more readable (someone mentioned using Gemini to explain what Claude is saying).
Need for ability to get Gemini to do research before answering (to avoid hallucinations). Some mention wanting Gemini to do research (like browsing) before answering.
Need for ability to use Gemini for language learning via voice chat.
Need for ability to use Gemini for live translation in multi-lingual environment (but current implementation picks up background speech, needs speaker separation).
Need for ability to have Gemini generate less verbose output, more concise.
Need for ability to have Gemini produce less hallucinations, more grounded.
Need for ability to have Gemini support extended thinking but currently insufficient.
Need for ability to have Gemini produce less "caveats and provisos" style (like Claude) which is annoying for some.
Need for ability to use Gemini as a "backend brain" for voice mode (like GPT Live with Claude backend) - some desire to combine models.
Thus we need to propose 4 concrete viable project ideas (software/tools/services) solving these pain points.
Let's brainstorm ideas:
"Gemini Voice Companion for Language Practice" - an app that uses Gemini Live mode to facilitate conversational practice in target languages, with features like speaker separation, context-aware correction, vocabulary hints, and ability to save sessions. It addresses need for language learning via voice, especially for niche languages (Afrikaans, Catalan). Use Gemini's live speech-to-speech, integrate with mic, add noise suppression, speaker diarization, and optional transcription.
"Gemini Research Assistant" - a tool that, before answering a user query, automatically performs a quick web search (using Google Custom Search or SerpAPI) and feeds the top results as context to Gemini to reduce hallucinations and improve grounding. It could be a browser extension or CLI that wraps Gemini API with a retrieval-augmented generation step. Addresses frustration about hallucinations and lack of up-to-date info.
"Gemini Output Clarifier" - a service that takes verbose/hedgy output from models like Claude (or even Gemini itself) and rewrites it to be concise, less caveat-heavy, more direct, using Gemini's strength in natural language. Could be a web UI or API where you paste Claude output and get a cleaned version. Addresses frustration with Claude's verbose caveats and desire for simpler explanations.
"Gemini Coding Rubber Duck Mode" - a VS Code extension or IDE plugin that when activated, sends prompts to Gemini with a system instruction to act as a rubber duck (ask clarifying questions, avoid giving implementation advice, focus on design discussion). It could also suppress code generation unless explicitly requested. Addresses need for rubber duck without unwanted code suggestions.
Alternatively, we could propose "Gemini Live Transcription & Translation Hub" for multi-lingual meetings with speaker separation and ability to translate to user's language, addressing need for live translation in noisy environments.
But we need exactly 4 ideas. Let's pick the most compelling based on discussion.
Let's extract quotes:
"I am an adult, and i lose to elemental school kids in chess." Not relevant.
"I am not so advanced enough as a human being." Not relevant.
"Gemini is underrated in that it produces the only prose that is somewhat bearable to read." Good.
"For heavyweight work I have been using Astra, but for rabbit holes and brain storming Gemini is far more enjoyable to interact with." So Gemini good for brainstorming.
"I'm worried in their push to catch up on the SOTA front, it's going to lose that natural sounding touch it currently has."
"Mostly because it answers quickly and is more agreeable (too agreeable at times)."
"Meanwhile Claude and Astra like to couch all their agreements with caveats and provisos."
"Sometimes that's what being smart sounds like."
"And sometimes that’s what trying to sound smart sounds like."
"Reality is full of special cases."
"This is commonly why, on Reddit in particular, you can get eaten alive."
"Someone confident but incorrect, can often sound more convincing than someone with actual expertise."
"The expert must add caveats/hedge, because those are the facts on the ground, whereas the person reciting google can be entirely confident."
"Of course the people judging aren't experts, so they side with confidence and simplicity."
"Heck, just writing shorter replies on Reddit is rewarded."
"Sometimes LLMs on high-thinking go off on full tangents based on little, and don't have the self-awareness to bring it back."
"In my experience actual experts don’t hedge because they have a perspective."
"To an expert communicating with a layperson is a form of compression."
"You must turn some very complex idea into one that you suppose the other person can grasp given their limited frame of reference."
"It's always lossy, and you have to guess how much you can remove without sounding patronizing or being inaccurate."
"That much knowledge can actually be detrimental to communication."
"I consider, on the contrary, caveating and hedging annoying 'typical redditor'/internet behaviors: they care more about being 'technically correct' than conveying the message."
"So I won't be addressing this, for those reasons and others."
"Yea. I have asked it to verify my ideas with experiments sometimes. And it cheats and warps the results so that the results are reached"
"caveats and provisos” makes me think of Robin Williams’ genie imitating William F. Buckley Jr."
"When you're a ChatGPT Projects or Claude Projects user, those caveats and provisos are your worst enemy because they'll change caveats into hard rules (either for the session or committed to memories) and you end up in absolute hell having to make it investigate to figure out why it can no longer produce anything but read-only pre-check code that never actually does anything but keeps performing stupid safety checks."
"Yep the only way out is hooks to forbid what can be detected by ast and second model to prune comments, flatten pyramids of fallback, and squash the test suite removing quirks maintaining wanted behaviors."
"Agreed. My impression is that the more verbose output of sol, astra etc is that it helps it steer itself on long running tasks (but is worse for the human user to read)"
"Yes, when post training models for long tasks this happens gradually. It is not easy to prevent it as such."
"Yes I've noticed there's also this drive to implement and start talking about how it would write specific portions of code in response to design/trade off questions. I have to prompt Sol/Astra almost every time with a note that I am not looking for implementation advice since I mostly use them as a rubber duck in the design phase"
"For rabbit holes, how do you get Gemini to do any research before answering? I've very recently had it hallucinate on me like it's 2023, and that was on Pro/Thinking, as far as I remember."
"Is Google still chasing frontier? Seems like they haven't had a "Pro" model in forever. I think a good niche for them would be right where they are now."
"They are. They were supposed to release 3.5 Pro over the summer but haven't because of persistent architectural/technical issues allegedly."
"Their AI leadership team has taken some hits recently too, in the form of departures."
"3.8 Flash has been a great model for me."
"How do you know you're using Astra?"
"My ChatGPT env only says "low", "medium", "high"."
"Is this a "pro" thing? I have totally no idea what I'm talking to, so actually I'm thinking of stopping my plan. Gemini and Claude are much more clear about it."
"Anyway, I like the speed at which Gemini responds so indeed for simple things it is preferable."
"For me (Plus plan, iOS app), Astra only shows up under the Work tab."
"I’ve been using Work for all my queries, since it seems to just be the same interface as Chat but with more features. I don’t understand why they’re two separate things."
"For the old school (I hate that that’s arguably applicable) AI dating types, and so on. The people using it not for productivity."
"It's entirely possible that in their testing of newer models, the whole problem is that even if it's doing better in benchmarks, maybe it's insufferable to work with, thus they're not releasing it."
"Wondering if people have managed to have Gemini in-front of other models like claude/codex models and only interact with that. Having Gemini act as a pure human/llm translator."
"Not in the principled sense you mean but I have in fact recently started having Gemini explain to me what Claude is talking to me about, lol."
"BTW, as kind of a follow-up to this, I think the most important finding to report--for those who only use one model, as many people seem to--is just how much more often Claude seems to be extremely confidently (and even insufferably) wrong than any of the other three big models (all of which I use quite often... yet I only pull out Claude when I've given up hope in a problem and are looking for out-of-the-box brainstorming)."
"And like, it does this despite it speaking in extremely dense math, which both makes it sound correct and requires a lot more effort to prove when it is wrong... yet, it isn't actually correct more often, and so that time sink just isn't worth the benefit. I then think many people--including people who can speak math (as can I)--just stop bothering to check everything, as if you come across a human who speaks like this it probably does correlate with slow and careful thought that helps prevent errors."
"Instead, Claude has the mistake rate of a somewhat accelerated beginner impossibly combined with the language of an expert professor; and we as humans just aren't good at that combination: it becomes very dangerous and makes it take longer to spot its egregious mistakes and trained-in biases. If you have to use Claude, I thereby claim you really need to have a team of not-Claudes to help insulate you from this, and Gemini (while being a bit senile) is a lot more collaborative and approaches problems in ways that makes it harder to get tricked."
"(To translate this into more of an engineering analogy: Claude always feels to me like the engineer who put more effort into learning how to program in functional languages than into how to actually develop working code, and then confidently presents you answers in Haskell or Lisp that never quite work. To find their errors is then very costly. In contrast, Gemini feels more like a Java or Go developer who knows they are a cog... that's helpful! <- Which maybe just goes to show that AI has finally turned me into a manager, omg.)"
"Yes I have noticed this. I frequently have stronger models review weaker models. It’s very instructive to see what they get wrong."
"Somebody shared this a few days ago: https://github.com/adnanakil/nobuzz"
"Does anyone know if there is a dedicated model which makes Claude output nore human readable and less slop?"
"Lately it became load-bearingly-reality-difficult to not only read, but to comprehend the Claude output"
"It's also the only model that generates accurate translation and localization. No other frontier model comes close. Although Gemini's coding capabilities are subpar, its natural language processing is top-tier."
"I’m curious how you guys keep track of each model’s coding capabilities. The landscape keeps changing. I don’t suppose you benchmark all frontier models every other month, right?"
"True but it has a niche in SQL reviews for me. Looks like Google has a lot of good sql in their corpus and in their RL digital lobotomy factory."
"I wonder if it's a harness thing or a model thing at this point. I feel all coding models are quite capable for most tasks I want them to do."
"Most of the time I don't need what the bench tests and I'm not really giving them completely ambiguous tasks without any refinement."
"I only find marginal differences between models at this point and it almost feels like personality quirks in each model than anything."
"When comparing OpenAI and Claude thats pretty much true, but not Gemini... And have you tried Antigravity? Yikes"
"The CLI version of agy is great. Have you tried it?"
"Do you dangerously allow permissions? I absolutely cannot use it until they ship an auto approver. As it is now I have it write one bash/python script to do everything it wants to, then I review that. Otherwise it is COMPLETELY unusable and it shocks me when I hear people are using it."
"Sounds like they shipped some changes today that might reduce approvals: "
"
alias agy="agy --dangerously-skip-permissions"
"
"is anti gravity open sourced just like codex or grok code?"
"Compared to gemini-cli that they took out behind the woodshed, I hate it."
"I've used Antigravity as my main coding agent on one of my biggest projects for about a year. It's been great for me. (and I use Claude, Codex, Grok and Muse for all the other projects)"
"I did a test involving implementing cobol control flow in Java for a source to source translation project. Gemini was the only model to get the edge cases. Cobol is very peculiar in this regard."
"It's very good at Elixir in my experience too. And it just does what I ask and doesn't wind me up like Opus. I don't think I've had to insult it more than once per day."
"I had a typical $20 Gemini plan that I just downgraded to their $5 plan (to keep access to some of the models). It had been so long since I let Gemini work on (or review) any code / design / html (anything) that I couldn't justify bothering to keep wasting money on it. It fell behind badly over the past year. Astra might as well be an alien super intelligence at code compared to Gemini. I enjoy talking to Gemini, it is very good at conversation, I get solid answers to everyday questions. I intend to keep the $5 plan indefinitely for basic use. I don't expect they'll ever resurface as a competitor in coding with Astra & Fable et al."
"I found that it's shockingly good with R. (the only language I know and can correct for)"
"I doubt they even intended it to be, but it seems like I kept going from resorting to 3.5-3.8 (over time) to realizing that Claude and GPT, while great at Python, will make rudimentary mistakes with R; even when they compose giant complicated R code."
"That might depend on whether you are translating fiction or nonfiction."
"Anecdotally I'd rate Gemini behind Claude and OpenAI models at fiction and I can't find any benchmarks showing Gemini is the clear winner at this task."
"We have an agentic system that produces insights for end users, and runs most of its work on DeepSeek v4.1 Flash but as an output stage transforms the resulting text through Gemini 3.8 Flash for readability, and it works."
"On my TODO is try and run all of the analysis pipeline in dense "machine speak" to save on tokens and just let Gemini sort it out at the end."
"In my experience, with minimum prompting, deepseek also generates very decent text."
"Not at all in mine. Deepseek has some of the worst prose of the close to frontier models in my opinion."
"I find its style the most sycophantic and annoying personally."
"Try Gemini live in a multi lingual environment. It can pick out speakers and live translate to you. Truly underrated for its capabilities."
"I actually did (was going to travel internationally), and it wasn't as useful as you'd think. I would be talking to someone, and in the background someone else would be talking, and it would translate both people.
Only worked in a 1:1 in a quiet place. Still, can't complain for free."
""Only worked in a 1:1 in a quiet place"
well then its not model problem
"It's not a model problem, but it is a SW problem. It should be able to distinguish nearby people from people farther away and give me options to set a threshold on who to include."
"As sister comments have pointed out, I find that it's very much dependent on the phone you're using. If it can't even show me a waveform then I don't think it's going to be able to tease out any sort of speech"
"Sorry I wasn't clear. It's about the hardware microphone. They're usually optimized to filter out far field sounds and only focus on near field sounds. So if you're trying to understand a conversation from even 10 ft away, it can be a problem. Not saying that's the situation that you encountered, but it does make sense that a $50 phone versus a $500 phone might have different capabilities when it comes to the microphone"