CTech AI Security Survey: Automation Raises the Bar for Human Researchers
CTech’s AI security survey found that automation now handles substantial research work, yet the shift is raising expectations for the humans supervising it. The September 18 survey covered 30 security researchers across Israeli cybersecurity companies. Participants described AI handling log parsing, documentation reviews, initial code scans, and parallel investigations.
That efficiency comes with a sharp reversal. Faster research does not make sound judgment less important. It gives weak assumptions, incomplete instructions, and excessive permissions more opportunities to cause damage at machine speed.
Security researchers are therefore moving from manual investigation toward orchestration, validation, and accountability. Their emerging opponent is not simply an AI model or an automated attacker. It is automation without reliable human judgment.
The change reaches beyond specialist research teams. Enterprises are filling their environments with agents, service accounts, application programming interface keys, and other nonhuman identities. Many can read sensitive data or take actions across connected systems.
Meanwhile, attackers can use similar automation to explore vulnerabilities and accelerate exploitation. Defenders must automate without trusting every generated conclusion. They must also secure the agents performing that work.
The result is a demanding new standard for the profession. Researchers need the technical depth to challenge a model, the operational discipline to control agents, and the communication skills to define precise objectives.
The CTech AI Security Survey Shows What Changed
AI has absorbed repetitive security tasks, moving the researcher’s value from manual output toward direction, verification, and technical judgment.
The security researcher survey collected responses from 30 professionals working across Israel’s cybersecurity sector. It was not a controlled productivity study. It was a structured view into how active research teams describe their changing work.
Every company in the survey said AI had taken over tasks such as log parsing, documentation review, and initial code scanning. One quarter of respondents explicitly described the technology as a force multiplier for individual output.
Yuval Barak, a founding research engineer at Astelia, offered the survey’s strongest productivity claim. He said agents had raised each researcher’s output to roughly ten times its previous level by handling laborious work.
That figure reflects one practitioner’s assessment, not an independently audited industry benchmark. Still, the underlying workflow change appeared consistently across the responses.
A researcher once had to collect logs, search documentation, compare code paths, and organize potential findings before testing a hypothesis. An agent can now perform several of those steps simultaneously.
The human role begins earlier and ends later. Researchers must define the question, select tools, constrain access, examine evidence, and decide whether a reported vulnerability is real.
Roy Itzhaky, a security researcher in Linx Security’s chief technology officer organization, described this role as a director rather than a digger. The agent explores several paths, while the researcher keeps those paths aligned and judges the results.
That distinction matters because research is not measured by how many leads a system produces. A credible finding must be reproducible, relevant, and supported by evidence.
Automated scanning can increase the number of hypotheses without improving their quality. It can also hide uncertainty behind confident prose or convincing technical artifacts.
A model might identify suspicious code and propose a plausible attack path. The researcher must still determine whether the path is reachable, whether existing controls block it, and whether exploitation creates meaningful impact.
The same principle applies to documentation work. AI can summarize previous investigations and surface related findings. It cannot guarantee that the context is current, complete, or applicable to a new environment.
This changes how teams manage institutional memory. Barak argued that knowledge can become a property of the team instead of remaining with individual researchers.
That outcome requires more than adding a chatbot to a document collection. Teams need organized evidence, traceable decisions, and access controls that distinguish trusted findings from unfinished speculation.
A searchable knowledge base can support that transition when it preserves source context. Without provenance, faster retrieval can circulate old assumptions as quickly as valid knowledge.
The event is therefore larger than a simple productivity story. AI is changing where researchers spend their effort and where security failures can enter the process.
The repetitive work is shrinking. The burden of specifying, supervising, and validating automated work is growing.
Faster Research Puts Security Teams Under Pressure
The productivity gain creates a new baseline, forcing researchers to deliver validated results faster while attackers gain access to similar automation.
Security teams face pressure from both sides of the workflow. Their own organizations expect more output, while adversaries can use AI to shorten reconnaissance and experimentation.
The CTech AI security survey shows how quickly internal expectations have adjusted. If an agent can review documentation or scan code in minutes, a long manual queue becomes harder to justify.
Ofri Ziv, co-founder and vice president of research at Tenzai, said work that previously stayed in research for months must now be evaluated, hardened, and delivered within weeks. That shift connects research performance directly to product delivery.
Faster delivery can improve defense when teams preserve validation. It can also encourage premature findings, noisy alerts, or inadequately tested controls.
The pressure is especially visible in vulnerability management. Organizations must find exploitable weaknesses, prioritize them, develop fixes, and verify those fixes across changing environments.
Google’s Mandiant team has described a similar operating model. Its agentic review framework combines multiple agents with structured validation and human expertise.
The framework is designed to help experts examine source code, find exploit paths, and test potential vulnerabilities. Its structure is important because unconstrained model exploration can waste time or produce unsupported conclusions.
Speed alone does not resolve the defender’s traditional disadvantage. An attacker needs one workable path, while a defender must reduce exposure across many assets, identities, and applications.
AI can help defenders parallelize that work. It also helps attackers automate searching, adapt public techniques, and examine exposed systems more frequently.
Google Cloud reported that the interval between vulnerability disclosure and active exploitation fell from weeks to days during the second half of 2025. That observation supports the case for continuous discovery and faster defensive validation.
It does not prove that AI caused every shorter exploitation window. It does show why periodic security work is becoming less compatible with current operating conditions.
A penetration test conducted twice each year captures a temporary view of an environment. New code, permissions, services, and integrations can make that view stale soon after delivery.
Four companies in the CTech survey said they had built proprietary AI hackers or internal training environments. These systems continuously simulate attacks and test how weaknesses could appear in production.
The survey participants described a zero-false-positive standard for those programs. That standard is an aspiration requiring evidence, not a general property of AI security tools.
False positives still carry real costs. They consume investigation time, interrupt engineering teams, and can train users to ignore future warnings.
False negatives are more dangerous. An agent might miss a novel technique because its training data or available tools do not represent the attack.
Security leaders must therefore judge automation through validated outcomes. Useful measures include reproducible findings, detection coverage, remediation time, and the rate of rejected agent conclusions.
The researcher’s forced response is clear. Teams must automate repetitive discovery while investing more heavily in review, testing, and evidence management.
This pressure is likely to persist. Once an organization integrates automation into its research pipeline, a return to slower manual throughput becomes difficult.
The harder question is whether teams can scale judgment at the same rate as generated output. That question defines the central conflict now facing security research.
AI Security Researchers Are Becoming Directors, Not Diggers
The profession’s main contest is human judgment against unsupervised automation, not human speed against machine speed.
AI security researchers no longer create value by personally completing every investigative step. Their value increasingly depends on deciding which steps should run, under what constraints, and with which evidence requirements.
That role resembles technical direction. A researcher can assign separate agents to review source code, search documentation, inspect logs, or test competing hypotheses.
Parallel work expands the search space. It also creates coordination problems that a single manual investigation might avoid.
Agents can repeat the same mistaken assumption across several branches. They can lose relevant context, pursue attractive but unproductive theories, or produce conflicting conclusions.
Itzhaky summarized the danger with a pointed contrast. Sloppy thinking once consumed an afternoon, but now it can send a fleet of agents in the wrong direction.
The statement captures the central reversal. Automation reduces the cost of action, yet it can multiply the cost of a poorly defined objective.
A human director must translate an investigative goal into bounded tasks. Each task needs a target, permitted tools, stopping conditions, and a standard for sufficient evidence.
The researcher must also decide how independent the agents should remain. Several agents repeating the same reasoning pattern do not provide meaningful confirmation.
A better design separates roles. One agent can generate hypotheses, another can challenge assumptions, and a controlled tool can reproduce the suspected behavior.
Even that process needs human review. Models share training patterns, and apparent disagreement does not guarantee independent reasoning.
Ofek Haviv, a cybersecurity researcher at Terra Security, argued that researchers are not competing with agents on speed. Their value lies in setting the standard that agent results must meet.
That standard includes technical truth, but it also includes organizational relevance. A theoretical weakness may have little impact when controls make the vulnerable path unreachable.
Conversely, a minor configuration error can create major exposure when an agent holds broad privileges. Context determines whether a finding deserves urgent action.
Researchers also need to understand the tools beneath generated answers. A person reviewing reverse-engineered code must recognize when a model invents behavior or misreads a control flow.
Idan Revivo, head of security research at Island, warned about newcomers who can prompt models but never learned to reverse engineer software. His concern is not resistance to automation.
It is a warning about verification capacity. When a model is confidently wrong, someone needs enough technical depth to challenge it.
That requirement raises the bar for junior researchers. Entry-level work traditionally offered repeated exposure to logs, code, systems, and common failure modes.
If agents absorb those tasks, new researchers might receive fewer opportunities to build intuition. Teams must deliberately replace that lost practice with supervised labs, reproduction exercises, and adversarial review.
The survey found that one fifth of respondents mentioned AI’s effect on the next generation of researchers. Roey Vilnai, director of cyber research at Axonius, said access to knowledge has raised expectations across experience levels.
Tamir Ishay Sharbat, director of security research at Zenity, offered a complementary view. He entered cybersecurity without a prior background and now considers AI use a requirement for researchers.
Together, those perspectives expose a training challenge. AI can widen access to knowledge while making superficial competence easier to mistake for expertise.
Hiring processes will need to test both tool use and independent reasoning. A candidate should be able to direct an agent, inspect its evidence, and continue when the model fails.
Teams should also preserve manual exercises for essential skills. Reverse engineering, exploit development, identity analysis, and network reasoning cannot survive as purely theoretical knowledge.
The best researchers will combine machine-scale exploration with grounded technical skepticism. They will know when to automate, when to narrow the scope, and when to stop the system.
That combination, not prompt fluency alone, is becoming the profession’s new baseline.
Nonhuman Identities Expand the Attack Surface
The agents performing security work also become privileged identities that organizations must discover, restrict, monitor, and revoke.
Eleven companies in the CTech survey identified AI agents as a major change to the security perimeter. Respondents grouped those agents with API keys, service accounts, and other nonhuman identities.
Tomer Bar, associate vice president of security research at Semperis, said nonhuman identities already outnumber people in some customer organizations. He projected that the ratio could reach ten to one in coming years.
That projection is a company executive’s forecast, not an independently verified market measurement. The underlying identity problem is nevertheless concrete.
An agent needs access to data, applications, and tools to perform useful work. Those connections give it an identity, a set of permissions, or borrowed access through another account.
Omer Nissim, a security researcher at Sweet Security, highlighted the danger of pairing broad permissions with an unpredictable decision process. He specifically pointed to agents connected through Model Context Protocol servers.
Model Context Protocol, or MCP, is a standard that lets AI systems connect to external tools and data sources. Its usefulness depends on how those connections are authorized and constrained.
An agent with read access to source code presents one risk level. An agent that can change cloud settings, send messages, or deploy code presents another.
OWASP describes excessive agency as a condition involving excessive functionality, permissions, or autonomy. Unexpected or manipulated model outputs can then trigger damaging actions.
The root problem is not necessarily malicious intent. An ambiguous instruction, hallucinated conclusion, compromised tool, or injected prompt can redirect an otherwise legitimate agent.
Least privilege remains the primary defense. An agent should receive only the tools and permissions required for its current task.
Read-only access is preferable when the work does not require changes. High-impact actions should use separate approval paths and narrowly scoped credentials.
Organizations also need individual agent identities. Sharing a human credential with an agent makes attribution difficult and can allow the system to impersonate its operator.
NIST’s recent agent identity guidance recommends treating agents as distinct entities. Each should have identifiers, credentials, and entitlements connected to its responsible user or system.
The guidance also highlights risks from static API keys and long-lived bearer tokens. Anyone obtaining those credentials can often use them without proving possession of a specific identity.
This makes credential storage and rotation part of AI security. Agent secrets can leak through configuration files, logs, memory stores, tool outputs, or debugging records.
Human approval does not solve every problem. Constant permission requests can create consent fatigue, encouraging users to approve actions without careful review.
A stronger design gives agents approved operating boundaries. Those boundaries should define accessible resources, permitted actions, time limits, and escalation conditions.
Monitoring must capture more than the final answer. Security teams need records of tool calls, identity use, retrieved data, changed resources, and approval decisions.
That evidence supports incident investigation and everyday quality control. It lets researchers reconstruct why an agent reached a conclusion or performed an action.
Ownership must also remain visible. Itzhaky noted that agents can be created by employees, automated pipelines, or other agents.
Without an accountable owner, a forgotten agent can retain permissions after its original task ends. That pattern resembles abandoned service accounts, but autonomous behavior makes the exposure harder to predict.
Discovery should therefore include agent inventories, credentials, tools, data access, and parent-child relationships. Revocation must terminate both the agent and any delegated access it created.
This is where the productivity story meets the security story. An organization can deploy agents faster than its identity program can govern them.
The resulting gap becomes a new research target. Security teams must investigate agent behavior while using agents to investigate everything else.
Automation Cannot Replace Skeptical Validation
The survey captures a real workflow shift, but its strongest productivity and accuracy claims still require independent measurement.
The CTech AI security survey offers valuable firsthand observations from working researchers. Its 30-person cohort remains concentrated within Israel’s cybersecurity industry.
The published article does not provide a randomized sample, standardized questionnaire, or independently audited performance data. Readers should not treat its ratios as universal workforce measurements.
A respondent’s tenfold productivity estimate can signal a meaningful change without proving a tenfold increase in validated discoveries. Output can mean reports, hypotheses, code reviews, experiments, or confirmed vulnerabilities.
Those categories have different value. Ten times more initial findings can create more work if most fail reproduction.
The same caution applies to zero-false-positive claims. A system can reduce false alarms by narrowing what it reports, but that choice might increase missed vulnerabilities.
Teams need to measure both sides. Precision describes how many reported findings are valid, while recall describes how many relevant weaknesses the system actually finds.
A tool can appear accurate when it reports only obvious issues. It may still miss subtle attack chains, unfamiliar software, or vulnerabilities requiring long contextual reasoning.
Google Cloud’s vulnerability guidance warns that agents can silently overlook novel techniques or zero-day vulnerabilities poorly represented in training data.
That limitation gives human expertise a specific role. Researchers must recognize when a scan’s coverage is too narrow and design tests outside the model’s familiar patterns.
Hallucination is another concern. A generated report can cite nonexistent functions, misunderstand a dependency, or infer exploitability from incomplete code.
Reproduction should therefore happen in controlled environments. A credible pipeline should preserve inputs, tool versions, prompts, system state, and observable results.
Agent outputs also require threat modeling. An attacker can plant text that manipulates a model, especially when the system reads untrusted websites, repositories, tickets, or documents.
That attack is called indirect prompt injection. Malicious instructions are hidden inside retrieved content and attempt to redirect the agent’s behavior.
The defense cannot rely entirely on telling the model to ignore bad instructions. Systems need isolated tools, scoped permissions, content boundaries, and explicit approval for consequential actions.
Researchers must also distinguish model confidence from evidence quality. Fluent explanations can make a weak finding look complete.
A strong review process asks several concrete questions. Can the behavior be reproduced? Is the vulnerable path reachable? What privilege does exploitation require? Which control blocks or detects it?
The process should also record rejected hypotheses. Those failures help teams improve benchmarks and prevent later agents from repeating the same unproductive path.
A security benchmark is a defined collection of tasks used to test system behavior. Internal benchmarks should include misleading evidence, partial access, tool failures, and unfamiliar vulnerability classes.
Production performance matters more than benchmark scores. Teams should compare agent findings with later incidents, manual reviews, and independent testing.
Human reviewers need protection from automation bias, which is the tendency to favor a machine recommendation simply because a system produced it.
Rotating reviewers, hiding an agent’s confidence score, or requesting an independent reproduction can reduce that bias. High-impact findings deserve stronger separation between discovery and approval.
The skeptic’s case does not imply that teams should abandon AI. Manual research also misses vulnerabilities, follows bad hypotheses, and suffers from inconsistent documentation.
The more defensible conclusion is narrower. AI expands research capacity, but capacity becomes useful only when teams maintain disciplined validation.
The new higher bar is not perfection. It is the ability to explain what the agent did, verify the result, and contain the consequences of error.
Continuous Security Becomes the Next Test
The next phase will be judged by continuous validation, measurable research quality, and enforceable control over agent identities.
The first signal to watch is whether continuous security assessment replaces periodic testing in production workflows. Several survey respondents argued that twice-yearly penetration tests no longer match the speed of software and attacker activity.
Continuous assessment means testing assets as code, configurations, identities, and external exposures change. It should complement deeper human reviews rather than turn security into an endless scanner feed.
Google described an agentic security pipeline that scans code, produces proofs, and constructs fixes. The company said it applies the approach across hundreds of millions of lines of internal code.
That system offers a useful indicator because it connects discovery with reproduction and remediation. The critical measure is whether other organizations can achieve similar discipline with smaller datasets, teams, and infrastructure.
Broad adoption would strengthen the survey’s central claim. Persistent alert noise or weak remediation rates would weaken it.
The second signal is how companies govern nonhuman identities. Agent inventories should become part of identity and access management, not an isolated AI policy document.
Useful progress would include unique agent credentials, short-lived access, delegated authorization, clear ownership, and complete action logs. Organizations should also be able to suspend an agent without disabling its human operator.
Watch for standards and products that connect an agent’s actions to a specific user, task, and approved operating boundary. That chain supports accountability without sharing personal credentials.
If deployments continue relying on broad API keys and generic service accounts, productivity will outpace governance. That outcome would reinforce concerns about shadow administrators and unpredictable access paths.
The third signal is how hiring and training change. Junior researchers still need direct practice with code, systems, exploitation, and evidence.
Teams should publish clearer competency expectations for AI-assisted roles. Interviews and training programs can test whether candidates detect fabricated findings, challenge model assumptions, and reproduce results manually.
A healthy transition will create structured apprenticeships around agent supervision. Experienced researchers can expose trainees to both successful runs and failures.
An unhealthy transition will remove foundational work without replacing its educational value. That would produce researchers who can operate interfaces but cannot verify the systems behind them.
These signals matter to enterprise buyers as much as security professionals. A vendor’s use of AI says little without details about permissions, evidence, review, and failure handling.
Buyers should ask who approves findings, how agents authenticate, what data they retain, and whether results can be reproduced. They should also ask how the system behaves when a tool fails or an input contains malicious instructions.
Developers should expect security reviews to move closer to each code change. They may receive faster feedback, but automated findings will still need enough context to support a fix.
Knowledge workers should care because agent identity problems extend beyond cybersecurity tools. Any assistant connected to email, documents, code, or business systems can become an overprivileged nonhuman identity.
The CTech AI security survey ultimately describes a redistribution of work, not the disappearance of researchers. Machines perform more collection and exploration. Humans carry more responsibility for objectives, constraints, and truth.
That arrangement can improve security when evidence remains central. It can increase risk when organizations confuse generated activity with verified progress.
The practical question is no longer whether AI belongs in security research. It is whether teams can make automation accountable before its scale exceeds their ability to supervise it.
Ask your security vendors and internal teams for one complete agent-assisted finding, including its permissions, evidence trail, human review, and final disposition. That record will reveal more than a broad productivity promise.



