top of page

OpenAI o3 和 Gemini Sheets 引发新一轮工作流战争

OpenAI o3 和 Gemini Sheets 推动 AI 工作流工具超越原始基准分数。知识工作者现在通过电子表格和日常任务节省的分钟数来评判工具。

过去一个月的 Reddit 讨论显示了这种分歧。用户发布 o3 的高分成绩,却仍在询问日常 Excel 工作为何依然缓慢。实验室数据与办公时间之间的差距已成为新的焦点。

AI 工作流工具现在面临直接考验。 它们能否减少人们每天打开的应用程序中的摩擦。

从基准测试转向节省分钟数

OpenAI 发布 o3 时着眼于文档内的代理式操作。Gemini Sheets 则推出了在表格内通过自然语言命令生成公式和清理数据的功能。Google's official AI blog 强调,这些应用内功能正在重新定义企业用户的生产力预期。

知识工作者注意到了这一时机。Reddit 上的讨论充满了对实际节省小时数而非排行榜分数的疑问。人们希望证明新模型确实减少了格式化单元格或查找错误所花费的时间。

这一变化之所以重要,是因为大多数办公室已经在使用 Google Sheets 或 Excel。新功能必须在这些标签页内运行,否则用户就会放弃。实验室性能不再能解决争论。相反,决定性因素在于分析师能否在 45 分钟内完成每周对账报告,而不是四小时。

考虑 o3 如何与文档代理集成。用户可以高亮包含不一致日期格式的区域,并指示模型在保留原始源列的同时标准化所有内容。Gemini Sheets 通过其原生侧边栏执行类似操作,接受“删除重复供应商 ID 并标记超过预算阈值的金额”等提示。这两种方法都绕过了编写辅助列、应用筛选器和双重检查结果的传统流程。

这一演变对团队速度有直接影响。曾经需要安排全天清理会议的部门现在可以安排专注的审查会议。节省的时间在每月多份报告中累积。一个使用 Gemini Sheets 的财务团队报告称,在采用自然语言数据清理流程后,每月结账流程从六天缩短至三天半。这些收益转化为更早提交给领导层以及降低加班成本,正如 The Verge 最近的报道所述。

从基准测试到分钟数的转变也改变了供应商营销产品的方式。OpenAI 和 Google 不再发布新的 MMLU 或 GPQA 数字,而是展示量化消除击键次数和降低错误率的案例研究。竞争舞台已从学术测试集转向分析师屏幕上实际打开的标签页。

更多证据出现在采购决策中。企业买家越来越多地要求现场演示,衡量在现有工作簿内完成端到端任务的时间,而不是孤立的准确性指标。采购团队现在要求供应商计时执行相同任务,例如合并三个不同的 CRM 导出、核对预算差异并生成执行摘要幻灯片。当 o3 在九分钟内完成序列,而传统基于宏的工作簿需要四十分钟时,采购对话立即转向部署。

根据 Bloomberg 的报道,采购周期本身得以压缩,因为切实的时间节省比抽象的准确性百分比提供了更清晰的投资回报叙述。因此,审查工具预算的财务主管优先考虑提供相同任务有无新功能并排视频录制的供应商。

自然语言命令如何改变数据清理

传统的电子表格清理要求用户掌握 TRIM、SUBSTITUTE 和 QUERY 等函数。这些操作需要语法准确性,通常需要多次迭代。o3 和 Gemini Sheets 通过将普通英语指令翻译成可执行步骤来降低这一门槛。

一个示例工作流从一列大小写混合的产品名称开始。用户输入“标准化大小写并展开常见缩写”。几秒钟内,模型返回一个格式一致的新列,同时保留原始值以供审计。Gemini Sheets 在结果旁边显示生成的公式,允许偏好保留手动控制的用户将逻辑复制到现有模板中。

o3 将同一原则扩展到多工作表工作簿。请求“将 Q3 选项卡中的预算数据同步到预测模型,并高亮超过 8% 的差异”会触发一个代理,该代理定位正确范围、执行查找并应用条件格式。该代理还可以在侧边栏中记录每次转换,以便团队保持可审计的轨迹。

实际要点很明确:以前消耗分析师每周 30% 至 40% 时间的数据准备任务,现在变成了 5 至 10 分钟的提示和验证周期。量化这些节省的组织看到项目吞吐量的直接影响。一位分析师现在可以支持两位额外利益相关者,而无需延长工作时间。

除了基本清理外,这些模型还能处理曾经需要冗长 VBA 脚本的嵌套条件逻辑。营销分析师可以提示“标记支出超过上月平均值 25% 以上且属于 Q2 之前启动的活动的行”,并获得高亮数据集以及底层公式的解释。这种透明度减少了新团队成员的入职时间,他们以前必须解码不透明的遗留宏。

跟踪击键级指标的团队报告称,一旦提示代理处理重复的连接和格式转换,即使是中级 Excel 用户也能以少 60% 的点击完成相同工作量。这种物理交互的减少也降低了 HR 部门记录的重复性劳损投诉。

谁现在感受到压力

处理每周报告和数据清理的团队承担了重任。他们在实时文件上测试 o3 代理和 Gemini 公式,并将输出与常规流程进行比较。

Google 和 OpenAI 都面临对可衡量速度的相同需求。如果工具需要额外步骤来导出或重新格式化,声称的收益就会消失。产品团队面临在用户返回手动编辑之前缩小差距的压力。

同样的动态也影响按小时计费的独立分析师。更快的电子表格工作改变了他们的日常产能。以前花费两小时准备客户数据的顾问现在可以将时间分配给更高价值的解释,在不提高费率的情况下提高可计费利用率。

营销运营团队也经历了这一转变。曾经需要跨广告平台和 CRM 导出手动连接的活动绩效表现在接受“按渠道计算每合格线索成本并显示表现最差的三项活动”等提示。模型在内部处理连接,并返回指标表和建议的后续问题。

这些群体有一个共同约束:他们的数据存在于现有的电子表格应用程序中。任何强制迁移到新环境的 AI 解决方案都会立即遇到阻力。因此,采用取决于就地性能,而不是理论能力。

中型 SaaS 公司的主管描述说,季度董事会材料准备从专门的三天冲刺转变为单一半天审查,因为差异评论和瀑布图从用于每周运营审查的同一源选项卡自动生成。

直接比较 OpenAI o3 和 Gemini Sheets

o3 强调跨链接文档的代理式多步推理,适合跨越财务、运营和销售选项卡的跨工作簿分析。Gemini Sheets 专注于无摩擦的表内操作,在不要求用户离开浏览器标签页的情况下显示建议。

当两个工具在同一数据集上接收相同提示时,o3 往往生成更复杂的转换脚本,预测下游用途如数据透视表准备。Gemini Sheets 提供更快的单操作结果,并与 Google Workspace 共享权限更紧密集成。因此,团队必须评估哪种摩擦点——上下文转移或就地执行——主导他们的工作流。

真实世界的对比测试揭示了更多细微差别。当一位投资分析师向两个系统输入包含私募股权现金流预测的 12 选项卡投资组合模型时,o3 自动生成了差异评论和贴现率变化的敏感性表,而 Gemini Sheets 快速生成了干净的摘要选项卡,但需要额外提示才能显示相同的评论。因此,决策者需要权衡洞察深度或执行即时性对他们的交付物更重要。

跨工具比较还揭示了可审计性的差异。o3 维护可见的中间代理操作链,可导出为 JSON 日志,而 Gemini Sheets 将公式来源注释直接嵌入单元格。受 SOX 控制的财务团队在文档要求超过实时速度时往往更喜欢前者。

真正的对手:上下文重置与持久记忆

核心竞争在需要每次会话新上下文的通用代理与已经持有持续工作历史的工具之间展开。OpenAI o3 和 Gemini Sheets 在一个应用程序内有所改进,但当工作表更改所有者或任务跨越数周时,仍然要求用户重述项目细节。

这种模式在许多团队中重复出现。新员工需要重复相同的背景。后续分析会丢失之前的决定,除非有人手动复制笔记。

持久记忆消除了这个循环。像 remio 这样的工具将会议笔记、文件编辑和之前的聊天保存在一个地方,因此下一个电子表格任务从已加载的事实开始。差异体现在减少来回沟通,而不是更高的基准分数。当分析师打开预算差异模型时,remio 会显示之前的 Q2 假设和产生这些假设的决定,消除了搜索聊天记录或电子邮件线程的需要。有关更深入的工作流模式,请参阅此 practical AI workflows 指南。

这种优势随着团队规模和项目持续时间而扩大。运行季度规划周期的组织受益最大,因为累积的上下文否则会在同一文件的多个版本中碎片化。

跨职能的详细工作流示例

Finance teams handling accounts-receivable aging reports now prompt o3 to “ingest the latest bank statement CSV, match remittances against open invoices, and list unmatched items older than 60 days.” The agent performs fuzzy matching on customer names, flags potential duplicates, and adds a column of suggested write-off reasons drawn from prior months’ notes. The analyst then reviews only the exceptions list rather than every line item.

Supply-chain analysts use Gemini Sheets to reconcile inventory counts from warehouse handheld scanners. Typing “highlight SKUs where physical count differs from system quantity by more than 3 percent and suggest reorder quantities based on last 90-day velocity” triggers an in-place table that calculates days-of-supply metrics and links directly to the ERP export already stored in Drive.

These examples illustrate how prompt specificity determines output quality. Follow-up refinements such as “exclude consignment inventory from the variance calculation” are handled without rewriting formulas, preserving the original source data for compliance audits.

知识工作者的实际影响

The move toward embedded AI changes performance expectations. Analysts are now measured not only on accuracy but on cycle time. Managers ask how quickly a forecast can be updated after new actuals arrive rather than how many manual checks were performed.

Job roles also evolve. Junior staff spend less time on rote formatting and more time validating model outputs and interpreting anomalies. Senior analysts shift focus toward scenario modeling and stakeholder communication because data hygiene tasks consume fewer hours.

Compensation structures may follow. Freelance analysts who demonstrate consistent delivery speed can justify premium rates or higher volume. Internal teams that publish time-saved metrics gain leverage in budget discussions for additional AI tooling.

Training programs are adapting as well. Internal workshops now dedicate more time to prompt engineering for domain-specific calculations and less time to function syntax. Career ladders increasingly list “AI-augmented workflow fluency” alongside traditional Excel mastery as a promotion criterion.

依赖 AI 电子表格工具的局限与风险

Some sheets contain company-specific logic that models misread without extra guidance. A formula that worked last quarter may fail when labels shift or new data arrives mid-month. Users report cases where o3 or Gemini produced clean-looking output that still needed manual checks for accuracy.

These gaps keep teams from full handoff. The tools speed initial setup yet leave a verification step that adds time back in. That remaining check remains the clearest limit on claimed workflow gains.

Additional risks include data-privacy exposure when sensitive financials are sent to cloud models and potential error propagation when downstream decisions rely on undetected hallucinations. Teams mitigate these concerns by maintaining shadow manual calculations on high-stakes reports and by restricting model access to anonymized subsets during early testing phases.

Version-control conflicts can also arise when multiple collaborators accept or reject AI suggestions simultaneously. Without clear governance on which user’s edits win, teams risk silent overwrites that surface only during final executive reviews.

下一季度值得关注的事项

Teams will share concrete time logs from repeated tasks. Look for reports that measure hours spent on the same monthly report before and after switching to the new AI features.

Watch whether Google or OpenAI adds connectors that pull project history from outside the sheet. Such updates would show they recognize context loss as the current bottleneck.

Finally, note any shift in Reddit discussions toward practical examples rather than score comparisons. When the conversation moves to saved hours, the market has settled on the new standard.

Knowledge workers already vote with their daily habits. The tools that match those habits without extra context will hold the advantage.

常见问题

How accurate are the formulas generated by o3 and Gemini Sheets?

Both tools produce correct formulas for common patterns yet require verification on domain-specific logic. Always inspect generated calculations before relying on them for board-level numbers.

Can these tools work offline?

Current versions require cloud connectivity. Offline fallback remains limited to cached suggestions generated during the last online session.

Will AI replace spreadsheet analysts?

The more common outcome is role redefinition. Analysts who adopt the tools increase output and focus on interpretive work that commands higher strategic value.

What data should never be sent to these models?

Personally identifiable information, unreleased earnings figures, and proprietary pricing models should remain behind enterprise data-loss-prevention controls until vendors publish clearer retention policies.

未来展望与新兴竞争者

Microsoft’s forthcoming Copilot enhancements for Excel and new startups focused on persistent spreadsheet memory will extend the same battlefield. The winner will be the platform that minimizes both manual formatting and repeated context entry while preserving auditability. Organizations should therefore pilot multiple approaches rather than committing exclusively to any single vendor during this rapid iteration cycle.

Enterprise buyers are also tracking integration depth with ERP systems and BI platforms. Deeper bidirectional sync would allow an o3 agent or Gemini Sheet to pull live ledger balances directly, further compressing the monthly close. Early pilots already show that combining these sheet-level tools with lightweight orchestration layers can reduce the full financial-close timeline by an additional 30 percent, demonstrating that the workflow war continues beyond any single model release.

Vendors introducing desktop-only encryption layers and on-premise model hosting are gaining attention from heavily regulated industries where cloud transmission of transaction-level detail remains prohibited. These alternatives underscore that the decisive variable continues to be seamless workflow fit rather than raw model scale.

 
 

免费开始

一款本地优先的AI助手,具备个人知识管理功能

为了获得更好的人工智能体验,

remio 目前仅支持Windows 10+ (x64)M-Chip Mac

在你的大脑里添加一个搜索栏

Ask remio

记住一切

​无需整理

bottom of page