For the last year and a half, I have been pushing LLM into its terminal chains - the result, to put it mildly, mixed. On one engagement, Ollama with mistral:7b generated a working awk parser to output nmap in 15 seconds - with my hands I would pick five minutes. On the next project ChatGPT confidently offered the flag --no-tls-validation for gobuster, which does not exist (correct flag - -k). The Pentester terminal is the point where AI experiments break down about reality: a properly tuned shell, alias for recon chains, and tmux sessions that experience a SSH cliff. Without this foundation, any AI assistant is a text generator, not a working tool.
No marketing tinsel: specific configs and teams that are copied in .zshrc and checked in five minutes. Everything is checked at Kali 2024.x/2025.x and Parrot OS 6.x with zsh 5.9+.
Kali has been delivered with 2020 zsh default, but it is set up weakly out of the box. Oh-My-Zsh with the right set of plugins turns the terminal into a full-fledged pentester tool, not just a window with a cursor.
Installation is standard: git clone each plugin in ~/.oh-my-zsh/custom/plugins/, then add names to the array plugins=() File ~/.zshrc. On the fresh Kali 2024.x everything is put in two minutes. Separately - topic powerlevel10k: it shows the current directory, git branch and the status of the last team right in the prompte. When ten tmux panels are open, it helps not to get confused, in which you are a directory.
The custom prompt with IP VPN session is not decoration, but navigation. Add in .zshrc Substitution $(ip -4 addr show tun0 2>/dev/null | grep -oP 'inet \K[\d.]+') the variable RPROMPT. When you work simultaneously with several client VPNs, see the address directly in the Prompte - must have. On the three engagements of the last five, I caught myself driving traffic through the wrong tunnel. After adding IP to the prompt, the problem disappeared.
The set that lives in mine .zshrc The last three years and are finalized after each project:
Bash:
# Recon
alias ns='nmap -sC -sV -oA nmap_default'
alias nf='nmap -sC -sV -p- -oA nmap_full'
alias nu='sudo nmap -sU --top-ports 50 -oA nmap_udp'
# Web
gob() { gobuster dir -u "$1" -w /usr/share/wordlists/dirb/common.txt -o gobuster.txt -t 50; }
alias ff='feroxbuster -w /usr/share/seclists/Discovery/Web-Content/raft-medium-directories.txt'
# Quick checks
alias myip='curl -s ifconfig.me && echo'
alias serve='python3 -m http.server 8888'
alias listener='rlwrap nc -lvnp'
A couple of comments. nmap_default and nmap_full - fixed output file names because the script wrapper (about it below) is looking for them for automatic parsing. serve - instant HTTP server for delivering tools to the target host. According to the classification MITRE ATT&CK this technique Ingress Tool Transfer (T1105, Command and Control) - the attacker lifts the web server on his machine and downloads utilities from it to the compromised host through wget or curl. In the context of the pentest - standard operation after initial access. listener with rlwrap adds a readline wrapper to the netcat: up-down arrows and history work in a reverse shell. Who at least once tried to press the arrow up in the bare nc - knows what pain it is.
Another function of practice is webcheck. Accepts URL, does curl -sI with the conclusion of the headlines, then whatweb for fingerprinting technologies, then nuclei -silent -t cves/ for quick inspection of known CVE. Three tools with a single call - the basic fingerprinting of a particular host in seconds.
Practical application - generation of single-liners directly from the terminal:
Model mistral:7b is responsible for 3-5 seconds on GPU for simple queries, for code generation - longer. For "generate regex for parsing" level tasks, "explain the output of this team," "suggest wordlist for API endpoint" - works well. To analyze the code for 200+ lines of the context window, the 7B model is not enough: it loses the thread after the third block. Needed here ollama pull llama3:70b, but it's already 48GB RAM minimum.
Hallucinations are the main problem. LLM periodically invents flags and tool options. Check each generated command through man or --help before starting for a real purpose - without exception. I got burned on this more than once.
When scope does not prohibit cloud services (lab-environment, CTF, own projects), ChatGPT API through curl closes tasks that local models pull badly: analyzing large pieces of code, generating complex payloads, writing Sigma rules for detection. But sending real customer data to the OpenAI API is an NDA breach in the vast majority of contracts. Only on their stands.
Nebula (BerylliumSec) is a CLI assistant that integrates with nmap, sqlmap, metasploit. Supports self-hosted LLM via Ollama - solves the problem of data breach. According to spark42.tech, one of the most mature projects with good documentation and a simple deploy.
CAI (Cybersecurity AI) is a framework with multi-agent architecture. Allows you to run several agents in parallel - one makes recon, the other analyzes the results. A number of observers consider CAI a more mature alternative to PentestGPT.
PentestGPT - the pioneer of the direction (2023), uses LLM through three modules: Reasoning (strategy), Generation (teams), Parsing (tool output analysis). In essence, proof-of-concept: for production tasks, CAI may be preferable.
All three work in the "human-in-the-loop" mode - the agent offers actions, the solution remains behind the pentester. This is a key limitation and at the same time advantage: AI accelerates the automation of the routine tasks of the pentester, but does not replace the expert assessment of the appetite risk and business context of the goal.
On his working car this is also relevant. File ~/.zsh_history contains a complete history for: IP purposes, test accounts, paths to exploits. If the laptop is compromised, it is a complete drain of the client's data.
Basic measures. Variable HISTCONTROL=ignorespace in .zshrc - commands starting with a spacer are not recorded in history. Recruited mysql -u root -p'пароль' (with a gap in front of the team) - history will not get. The second option is HISTIGNORE='password:secret:token' to automatically exclude commands with sensitive words.
After completion of the project - cleaning. Technique Clear Command History (T1070.003, Defense Evasion) used and attacking for the sweeping of traces, and pentesters for OPSEC. Team history -c && rm ~/.zsh_history on your car is a mandatory item in the click-through checklist of the engagement.
One more point is GTFOBins. Utilities awk, bash and even cat, according to the catalog gtfobins.github.io, are used for privesc and bypass restrictions in the presence of SUID or sudo access. The same binary we use for automation is the potential privilege effect vector on our working machine. It is necessary to check periodically find / -perm -4000 -type f 2>/dev/null and on your own Kali. (Yes, I also thought for a long time that "well, this is my car" - and then I found SUID on python3, which left the script there two years ago.)
No marketing tinsel: specific configs and teams that are copied in .zshrc and checked in five minutes. Everything is checked at Kali 2024.x/2025.x and Parrot OS 6.x with zsh 5.9+.
Zsh and Oh-My-Zsh - setting up a pentest terminal
Kali has been delivered with 2020 zsh default, but it is set up weakly out of the box. Oh-My-Zsh with the right set of plugins turns the terminal into a full-fledged pentester tool, not just a window with a cursor.
Plugins for daily work
Three plugins that are really worth the attention.- zsh-autosuggestionsprompts commands from the story with a grey text - you shake the arrow to the right and get the full command you recruited yesterday. For repetitive nmap scans with long flags - saving 30-40 taps on each team.
- zsh-syntax-highlightinghighlights valid commands green, non-valid in red: see the error before pressing Enter.
- fzf- fuzzy search by history through Ctrl+R, instead of a linear bulkhead, you get an interactive filter throughout history.
Installation is standard: git clone each plugin in ~/.oh-my-zsh/custom/plugins/, then add names to the array plugins=() File ~/.zshrc. On the fresh Kali 2024.x everything is put in two minutes. Separately - topic powerlevel10k: it shows the current directory, git branch and the status of the last team right in the prompte. When ten tmux panels are open, it helps not to get confused, in which you are a directory.
The custom prompt with IP VPN session is not decoration, but navigation. Add in .zshrc Substitution $(ip -4 addr show tun0 2>/dev/null | grep -oP 'inet \K[\d.]+') the variable RPROMPT. When you work simultaneously with several client VPNs, see the address directly in the Prompte - must have. On the three engagements of the last five, I caught myself driving traffic through the wrong tunnel. After adding IP to the prompt, the problem disappeared.
Alias Linux for Pentester - Recon by one team
Alias are not about laziness, but about speed. On a typical engagement, the same commands with the same flags are recruited dozens of times. Instead of nmap -sC -sV -oA Each time - Alias ns, and the fingers are free to analyze.The set that lives in mine .zshrc The last three years and are finalized after each project:
Bash:
# Recon
alias ns='nmap -sC -sV -oA nmap_default'
alias nf='nmap -sC -sV -p- -oA nmap_full'
alias nu='sudo nmap -sU --top-ports 50 -oA nmap_udp'
# Web
gob() { gobuster dir -u "$1" -w /usr/share/wordlists/dirb/common.txt -o gobuster.txt -t 50; }
alias ff='feroxbuster -w /usr/share/seclists/Discovery/Web-Content/raft-medium-directories.txt'
# Quick checks
alias myip='curl -s ifconfig.me && echo'
alias serve='python3 -m http.server 8888'
alias listener='rlwrap nc -lvnp'
A couple of comments. nmap_default and nmap_full - fixed output file names because the script wrapper (about it below) is looking for them for automatic parsing. serve - instant HTTP server for delivering tools to the target host. According to the classification MITRE ATT&CK this technique Ingress Tool Transfer (T1105, Command and Control) - the attacker lifts the web server on his machine and downloads utilities from it to the compromised host through wget or curl. In the context of the pentest - standard operation after initial access. listener with rlwrap adds a readline wrapper to the netcat: up-down arrows and history work in a reverse shell. Who at least once tried to press the arrow up in the bare nc - knows what pain it is.
Bash features for complex tasks
When you need to pass the argument in the middle of the team, the alias does not cope - you need a function. In .zshrc: quickscan() { nmap -sC -sV -oA scan_$1 $1 && grep "open" scan_$1.nmap | awk '{print $1}'; }. Challenge - quickscan 10.10.10.1. Result - files scan_10.10.10.1.* plus a list of open ports in stdout. Primitively, but saves two steps on each host.Another function of practice is webcheck. Accepts URL, does curl -sI with the conclusion of the headlines, then whatweb for fingerprinting technologies, then nuclei -silent -t cves/ for quick inspection of known CVE. Three tools with a single call - the basic fingerprinting of a particular host in seconds.
Ollama - local LLM without leakage scope
Ollama allows you to run LLM locally on a working machine. For a pentester, this is critical: scope client, IP addresses, credentials found - all this cannot be sent to the cloud API. Installation on Kali - curl -fsSL https://ollama.com/install.sh | sh, then ollama pull mistral to download the model. Precondition: GPU with 8 GB+ VRAM for the 7B model, or 16 GB of RAM for CPU-inference (will be slow, but working).Practical application - generation of single-liners directly from the terminal:
Model mistral:7b is responsible for 3-5 seconds on GPU for simple queries, for code generation - longer. For "generate regex for parsing" level tasks, "explain the output of this team," "suggest wordlist for API endpoint" - works well. To analyze the code for 200+ lines of the context window, the 7B model is not enough: it loses the thread after the third block. Needed here ollama pull llama3:70b, but it's already 48GB RAM minimum.
Hallucinations are the main problem. LLM periodically invents flags and tool options. Check each generated command through man or --help before starting for a real purpose - without exception. I got burned on this more than once.
When scope does not prohibit cloud services (lab-environment, CTF, own projects), ChatGPT API through curl closes tasks that local models pull badly: analyzing large pieces of code, generating complex payloads, writing Sigma rules for detection. But sending real customer data to the OpenAI API is an NDA breach in the vast majority of contracts. Only on their stands.
Open-source AI agents for pentest
In 2025-2026, there were several mature open-source projects that automate the pentest comprehensively rather than just answering questions. According to spark42.tech and ostorlab.co, three projects deserve attention.Nebula (BerylliumSec) is a CLI assistant that integrates with nmap, sqlmap, metasploit. Supports self-hosted LLM via Ollama - solves the problem of data breach. According to spark42.tech, one of the most mature projects with good documentation and a simple deploy.
CAI (Cybersecurity AI) is a framework with multi-agent architecture. Allows you to run several agents in parallel - one makes recon, the other analyzes the results. A number of observers consider CAI a more mature alternative to PentestGPT.
PentestGPT - the pioneer of the direction (2023), uses LLM through three modules: Reasoning (strategy), Generation (teams), Parsing (tool output analysis). In essence, proof-of-concept: for production tasks, CAI may be preferable.
All three work in the "human-in-the-loop" mode - the agent offers actions, the solution remains behind the pentester. This is a key limitation and at the same time advantage: AI accelerates the automation of the routine tasks of the pentester, but does not replace the expert assessment of the appetite risk and business context of the goal.
OPSEC terminal - when shell history becomes a vulnerability
The bash/zsh terminal is an attack tool, and you need to configure it with the OPSEC. According to the classification MITRE ATT&CK technique Shell History (T1552.003, Credential Access) describes a scenario in which an attacker or a defender extracts credentials from the history of the teams. Recruited mysql -u root -pSecretPass123 on the target host - the password remains in .bash_history. Classic.On his working car this is also relevant. File ~/.zsh_history contains a complete history for: IP purposes, test accounts, paths to exploits. If the laptop is compromised, it is a complete drain of the client's data.
Basic measures. Variable HISTCONTROL=ignorespace in .zshrc - commands starting with a spacer are not recorded in history. Recruited mysql -u root -p'пароль' (with a gap in front of the team) - history will not get. The second option is HISTIGNORE='password:secret:token' to automatically exclude commands with sensitive words.
After completion of the project - cleaning. Technique Clear Command History (T1070.003, Defense Evasion) used and attacking for the sweeping of traces, and pentesters for OPSEC. Team history -c && rm ~/.zsh_history on your car is a mandatory item in the click-through checklist of the engagement.
One more point is GTFOBins. Utilities awk, bash and even cat, according to the catalog gtfobins.github.io, are used for privesc and bypass restrictions in the presence of SUID or sudo access. The same binary we use for automation is the potential privilege effect vector on our working machine. It is necessary to check periodically find / -perm -4000 -type f 2>/dev/null and on your own Kali. (Yes, I also thought for a long time that "well, this is my car" - and then I found SUID on python3, which left the script there two years ago.)