492 lines
20 KiB
Markdown
492 lines
20 KiB
Markdown
# SignalScope
|
|
|
|
SignalScope is a local multi-source security analytics agent. Its primary mode connects to Graylog through MCP, discovers the streams and fields already available in your environment, and uses stream profiles to normalize events, build baselines, correlate entities, and explain anomalies with a local LLM.
|
|
|
|
Its running only locally and if using LLM it's running also locally so no data is sent or exposed outside.
|
|
|
|
FortiGate is one supported example. The same workflow applies to DNS/AdGuard, Windows Event Logs, Sysmon, Nginx, Squid, VPN, Proxmox, Filebeat-collected logs, and future Graylog streams.
|
|
|
|
The Python module and legacy `fgai` command remain available for compatibility. New installations can use `signalscope`.
|
|
|
|
The prioritized implementation plan is tracked in [ROADMAP.md](ROADMAP.md).
|
|
|
|
Autoblocking is dry-run by default. The tool will not block RFC1918, loopback, multicast, link-local, reserved, or allowlisted addresses unless you change the code.
|
|
|
|
## Screenshots
|
|
|
|
### Settings
|
|

|
|
|
|
### Findings
|
|

|
|
|
|
### Correlation
|
|

|
|
|
|
## Quick Start
|
|
|
|
```bash
|
|
python -m venv .venv
|
|
source .venv/bin/activate
|
|
pip install -e .
|
|
```
|
|
|
|
Or use the helper script, which creates/uses `.venv` automatically and runs `pip install -e .`:
|
|
|
|
```bash
|
|
./start.sh
|
|
./start.sh status
|
|
./start.sh analyze
|
|
./start.sh stop
|
|
```
|
|
|
|
`./start.sh` starts three local background processes:
|
|
|
|
- Optional UDP syslog listener writing `logs/fg_syslog.jsonl`
|
|
- Continuous monitor writing `state/fgai-status.json`
|
|
- Local dashboard at `http://127.0.0.1:8088`
|
|
|
|
## Primary Workflow: Graylog MCP
|
|
|
|
Graylog 7.1 MCP is the primary log-source integration. In the dashboard, open
|
|
`Settings`, select `Graylog MCP`, provide the MCP URL and a read-only API token,
|
|
then load and enable the streams to analyze. SignalScope uses MCP `list_streams`,
|
|
`list_fields`, `search_messages`, and `aggregate_messages` to work with existing
|
|
log sources rather than requiring every source to be forwarded locally.
|
|
|
|
The token field accepts a raw Graylog API token, the Base64 value after `Basic `,
|
|
or a complete `Basic <value>` header. Tokens are stored only in the local runtime
|
|
configuration and are never returned by the dashboard API.
|
|
|
|
Use `Edit profile` on a stream to load its fields. The field table shows Graylog
|
|
datatype/capability metadata and lets you select one or more entity fields, a
|
|
time field, and categorical/numeric fields for the stream profile. Profiles are
|
|
stored under `graylog_stream_profiles` in `state/fgai-config.json`.
|
|
|
|
The Settings page also shows recommended stream profiles built from observed
|
|
field coverage and cardinality. These recommendations use deterministic
|
|
discovery first, then can optionally be refined by a local Ollama profile advisor
|
|
model such as `qwen3:8b` or `qwen3:14b`. Enable `Ollama profile advisor` and set
|
|
`Profile advisor model` in Settings. Advisor output must be valid JSON and is
|
|
validated against fields actually seen in the stream before it can be applied.
|
|
Unknown fields, raw message fields, internal `fgai_*` fields, and unknown
|
|
detectors are rejected.
|
|
|
|
Settings also lists locally installed Ollama models from `http://127.0.0.1:11434/api/tags`.
|
|
Click a model name to fill both the dashboard analyst model and profile advisor
|
|
model fields.
|
|
|
|
The settings page treats stream enablement and profile editing separately. The
|
|
checkboxes decide which streams are monitored. Click `Edit profile` on one stream
|
|
to load its fields and edit only that stream's profile; saving with no active
|
|
profile editor leaves existing profiles unchanged.
|
|
|
|
When many Graylog streams are available, use `Diagnostics -> Stream Coverage` to
|
|
see which streams are enabled, which have profiles, how many profile fields are
|
|
baseline-ready, how many events were fetched, and whether a stream is `ready`,
|
|
`learning`, `missing_profile`, `no_events`, or `not_enabled`.
|
|
|
|
Enabled streams are normalized through the same event model. Stream profiles
|
|
define the entity, timestamp, categorical, and numeric fields used for baselines.
|
|
The dashboard and Ollama then correlate behavior across sources, for example a
|
|
client IP appearing in FortiGate, AdGuard/DNS, Windows Security, Nginx, Squid,
|
|
VPN, or Proxmox.
|
|
|
|
A stream profile can track multiple entities from the same event, such as
|
|
`username`, `srcip`, and `hostname`. SignalScope stores profile baselines for
|
|
each selected entity value, which makes cross-source investigation work even when
|
|
one source is user-centric and another is IP- or host-centric.
|
|
|
|
SignalScope keeps a common alias map for fields such as source IP, destination
|
|
IP, ports, action, severity, service/protocol, DNS query, URL, message, and event
|
|
type. This lets Related Activity and correlations work with firewall/proxy/DNS
|
|
streams that use names like `src_addr`, `destination.ip`, `dest_port`,
|
|
`fw_action`, `priority`, `proto`, or `full_message` without adding a new parser
|
|
for every product.
|
|
|
|
Correlation is entity-aware rather than FortiGate-specific. SignalScope recognizes
|
|
common IP fields such as `srcip`, `source_ip`, `remote_addr`, and Windows event
|
|
IP fields; account fields such as `username`, `user`, and `TargetUserName`; and
|
|
host fields such as `hostname`, `computer`, and `winlog_computer_name`. Configure
|
|
the exact entity field per stream in the profile when your Graylog schema differs.
|
|
|
|
Each profile baseline is stored per stream, entity, selected field, and five-minute
|
|
bucket. Once enough history exists, SignalScope compares the current rate or
|
|
numeric value to the same UTC weekday/hour where possible, then falls back to the
|
|
stream's overall history. Repeated MCP pages are fingerprinted so the same
|
|
Graylog event is not learned repeatedly. Related anomalies, profile deviations,
|
|
and multi-stream correlations are grouped into investigation incidents with a
|
|
compact evidence timeline.
|
|
|
|
Timeline and related-activity rows include copyable Graylog query details built
|
|
from normalized source, destination, action, and DNS fields. These are query
|
|
details rather than hard-coded web links, so they work with MCP and with Graylog
|
|
deployments behind different URLs or reverse proxies.
|
|
|
|
Incident lifecycle state is stored locally in `state/signalscope-incidents.json`.
|
|
Use the dashboard incident actions to acknowledge, resolve, or reopen an incident
|
|
and attach a note. The state is keyed to a stable incident fingerprint so it can
|
|
survive monitor refreshes even when the current detection window changes.
|
|
|
|
Export the current investigation view when you need to share or archive an
|
|
incident outside the dashboard:
|
|
|
|
```bash
|
|
signalscope export-investigation --format markdown --output exports/investigation.md
|
|
signalscope export-investigation --incident-id <incident-id> --format json
|
|
```
|
|
|
|
The report is built from `state/fgai-status.json` by default and includes
|
|
summary counters, stream coverage, incident state, analyst notes, evidence,
|
|
timeline rows, and Graylog query details. The dashboard exposes the same data at
|
|
`/api/export/incidents?format=markdown` or `format=json`.
|
|
|
|
With a stream profile in place, SignalScope also builds independent burst
|
|
baselines for authentication failures, DNS queries, and deny/block actions when
|
|
those events are present. These are evaluated per configured entity, so a Windows
|
|
account, DNS client, or firewall source is compared to its own history.
|
|
|
|
SignalScope also detects ordered behavior sequences across any streams that share
|
|
an entity. The built-in sequence is category-based, not source-specific:
|
|
`dns_query -> network_connection -> auth_failure`. Those categories can come from
|
|
AdGuard, Windows DNS, a proxy, firewall, VPN, endpoint, or any other Graylog
|
|
stream as long as the fields normalize into the same generic event model.
|
|
|
|
Stream profiles can also carry `field_weights` to tune scoring without changing
|
|
the baseline itself. Weights are multipliers from `0` to `5` and can target a
|
|
field, a detector, or a field+detector pair:
|
|
|
|
```json
|
|
{
|
|
"url": 1.5,
|
|
"auth_failure_burst": 2,
|
|
"query_domain": {
|
|
"rare_value": 1.8
|
|
}
|
|
}
|
|
```
|
|
|
|
Use replay or `replay-graylog --compare-config-file` to test score changes before
|
|
applying them to the live profile.
|
|
|
|
Replay a historic JSONL or Graylog export without changing the live baseline:
|
|
|
|
```bash
|
|
signalscope replay --logs exports/windows-history.jsonl --stream-id <configured-stream-id>
|
|
```
|
|
|
|
Replay uses a temporary SQLite baseline and evaluates events in timestamp order.
|
|
It reports detector counts and the findings that would have been generated. Use
|
|
the configured stream ID so the export is evaluated with that stream's profile.
|
|
|
|
Replay directly from Graylog MCP without touching the live baseline:
|
|
|
|
```bash
|
|
signalscope replay-graylog --range-seconds 86400
|
|
```
|
|
|
|
By default this uses the enabled streams from the dashboard configuration. Limit
|
|
the run to one or more streams with repeated `--stream-id` flags. To test a
|
|
candidate detector/profile configuration before applying it, compare it against
|
|
the current runtime config:
|
|
|
|
```bash
|
|
signalscope replay-graylog --range-seconds 86400 --compare-config-file exports/candidate-config.json
|
|
```
|
|
|
|
The comparison reports detector-count, field-finding, and source-anomaly deltas
|
|
using the same fetched event window.
|
|
|
|
The current MCP endpoint is `http://<graylog-host>:9000/api/mcp`. Enable it in
|
|
Graylog under `System -> Configurations -> MCP` and use stream IDs internally;
|
|
the fgAI stream picker resolves titles in the UI.
|
|
|
|
For live monitoring, `graylog_range_seconds` controls how far back each MCP poll
|
|
searches. The default is 300 seconds, so each poll re-checks the last five
|
|
minutes. `graylog_max_events_per_stream` caps raw events fetched from each
|
|
stream during that window. If a stream hits the cap, the dashboard marks the
|
|
window as truncated because high EPS means SignalScope sampled only part of the
|
|
Graylog result set. For very high-volume streams, prefer aggregate baselines and
|
|
targeted drill-down queries over trying to pull every raw event through MCP.
|
|
Large values such as 100000 can require hundreds of paged MCP searches across
|
|
enabled streams. If Graylog times out or rejects the query, SignalScope keeps the
|
|
events already fetched, marks the stream as a partial fetch, and shows the MCP
|
|
error in Diagnostics instead of failing the whole dashboard update.
|
|
The monitor also stores the last successful dashboard status in
|
|
`state/signalscope-status-cache.sqlite3`. If a later live MCP fetch fails before
|
|
usable data is available, the dashboard keeps showing the last good findings,
|
|
graphs, incidents, and correlations with a stale-data warning instead of going
|
|
blank.
|
|
|
|
Graylog fetch mode controls how high-volume streams are read:
|
|
|
|
- `raw`: fetch raw events up to `graylog_max_events_per_stream`.
|
|
- `aggregate`: use Graylog MCP `aggregate_messages` for total event volume, then
|
|
fetch only `graylog_raw_sample_events` raw events per stream for findings and
|
|
drill-down context.
|
|
- `auto`: use aggregate mode automatically when `graylog_max_events_per_stream`
|
|
is larger than `graylog_raw_sample_events`.
|
|
|
|
For high EPS environments, keep `graylog_range_seconds` at 300, set
|
|
`graylog_fetch_mode` to `auto` or `aggregate`, and use a modest raw sample such
|
|
as 5000. The dashboard then shows aggregate event volume without forcing every
|
|
raw log line through MCP each poll.
|
|
|
|
## Monitoring Export
|
|
|
|
The dashboard also exposes Prometheus text metrics at:
|
|
|
|
```text
|
|
http://127.0.0.1:8088/metrics
|
|
```
|
|
|
|
This endpoint is passive and has no Prometheus or Grafana dependency. It reports
|
|
low-cardinality event counts, anomaly severities, incident counts, baseline readiness, and Graylog
|
|
MCP health. Use it later as a Prometheus scrape target or as input for a Checkmk
|
|
local check. Do not use source IPs, domains, or raw event IDs as metric labels.
|
|
|
|
Enable cached Ollama analyst notes in the dashboard:
|
|
|
|
```bash
|
|
FGAI_LLM=1 OLLAMA_MODEL=llama3.1 ./start.sh restart
|
|
```
|
|
|
|
The monitor refreshes deterministic detections every `FGAI_MONITOR_INTERVAL` seconds and refreshes the LLM note every `FGAI_LLM_INTERVAL` seconds, default `300`.
|
|
|
|
The script activates `.venv` inside the script process. If you also want your current shell prompt to show the venv, run:
|
|
|
|
```bash
|
|
source .venv/bin/activate
|
|
```
|
|
|
|
For UDP `514`, the script starts only the listener command with `sudo`:
|
|
|
|
```bash
|
|
FGAI_SYSLOG_PORT=514 ./start.sh
|
|
```
|
|
|
|
The syslog receiver rotates the active JSONL input at 25 MB by default. Rotated
|
|
files are gzip-compressed and 14 archives are retained. Override this when needed:
|
|
|
|
```bash
|
|
FGAI_LOG_ROTATE_BYTES=$((100 * 1024 * 1024)) FGAI_LOG_ROTATE_COUNT=30 ./start.sh restart
|
|
```
|
|
|
|
The continuous monitor also stores a local SQLite behavior baseline at
|
|
`state/fgai-baseline.sqlite3`. A source becomes baseline-ready after 12 completed
|
|
five-minute windows. Historical rate and hitcount-rate deviations then contribute
|
|
to its anomaly score. Set `FGAI_BASELINE_DB` to use another location.
|
|
|
|
SignalScope prunes old baseline buckets during each monitor cycle. The defaults
|
|
keep 14 days of buckets and dedupe history, prune stale one-off categorical
|
|
values after 7 days, and cap high-cardinality values per stream/entity/field.
|
|
Tune these in the dashboard or in `state/fgai-config.json`:
|
|
|
|
```json
|
|
{
|
|
"baseline_training_days": 7,
|
|
"baseline_retention_days": 14,
|
|
"baseline_value_retention_days": 7,
|
|
"baseline_max_values_per_field": 2000
|
|
}
|
|
```
|
|
|
|
`baseline_training_days` is the minimum baseline age before profile deviations
|
|
are promoted into live triage. For production data, set this to the amount of
|
|
history you trust, commonly `7` to `14` days. The dashboard shows fields as
|
|
`learning` until both bucket count and baseline age are sufficient.
|
|
|
|
If an existing baseline database has already grown large, stop the monitor and
|
|
run a manual prune plus SQLite compaction:
|
|
|
|
```bash
|
|
signalscope baseline-maintenance --baseline-db state/fgai-baseline.sqlite3 --retention-days 14 --value-retention-days 7 --max-values-per-field 2000 --vacuum
|
|
```
|
|
|
|
`VACUUM` can take time on a large database and should not be run while the
|
|
monitor is actively writing.
|
|
|
|
Analyze local logs:
|
|
|
|
```bash
|
|
fgai analyze-logs --logs logs/fg_syslog.jsonl
|
|
```
|
|
|
|
Open the live UI after `./start.sh`:
|
|
|
|
```bash
|
|
xdg-open http://127.0.0.1:8088
|
|
```
|
|
|
|
Score likely traffic anomalies:
|
|
|
|
```bash
|
|
fgai detect-anomalies --logs logs/fg_syslog.jsonl --min-score 35
|
|
fgai detect-anomalies --logs logs/fg_syslog.jsonl --min-score 35 --llm --llm-timeout 300
|
|
```
|
|
|
|
Generate response and policy recommendations:
|
|
|
|
```bash
|
|
fgai recommend --logs logs/fg_syslog.jsonl --min-score 35
|
|
```
|
|
|
|
Optional external reputation enrichment is disabled by default. To use VirusTotal for public source/destination IP reputation:
|
|
|
|
```bash
|
|
export FGAI_THREAT_INTEL=1
|
|
export ABUSEIPDB_API_KEY='...'
|
|
fgai recommend --logs logs/fg_syslog.jsonl --min-score 35 --threat-intel
|
|
```
|
|
|
|
VirusTotal is also supported:
|
|
|
|
```bash
|
|
export FGAI_THREAT_INTEL=1
|
|
export FGAI_THREAT_INTEL_PROVIDER=virustotal
|
|
export VIRUSTOTAL_API_KEY='...'
|
|
fgai recommend --logs logs/fg_syslog.jsonl --min-score 35 --threat-intel
|
|
```
|
|
|
|
Threat intelligence can also be configured in the dashboard settings. Choose
|
|
`auto`, `abuseipdb`, or `virustotal`, paste the provider API key, and set the
|
|
daily lookup budget and cache TTLs. API keys are stored only in the local config
|
|
file and are not returned back to the browser after saving.
|
|
|
|
Threat intelligence responses are cached locally in `state/threat-intel-cache.json`. Successful results are reused for seven days by default, failures for one hour, and SignalScope permits at most 100 new provider lookups per UTC day. Cached responses are returned even after that budget is reached. Tune these safeguards in the dashboard or with `FGAI_THREAT_INTEL_TTL_SECONDS`, `FGAI_THREAT_INTEL_ERROR_TTL_SECONDS`, and `FGAI_THREAT_INTEL_DAILY_LIMIT`.
|
|
|
|
Listen for FortiGate syslog locally:
|
|
|
|
```bash
|
|
fgai listen-syslog --port 5514 --output logs/fg_syslog.jsonl
|
|
```
|
|
|
|
Run the listener quietly in the background:
|
|
|
|
```bash
|
|
./start.sh
|
|
```
|
|
|
|
Stop the background listener:
|
|
|
|
```bash
|
|
./start.sh stop
|
|
```
|
|
|
|
UDP port `514` normally needs root privileges on Linux:
|
|
|
|
```bash
|
|
sudo .venv/bin/fgai listen-syslog --port 514 --output logs/fg_syslog.jsonl
|
|
```
|
|
|
|
Test FortiGate API access:
|
|
|
|
```bash
|
|
export FORTIGATE_HOST=192.0.2.10
|
|
export FORTIGATE_API_TOKEN='...'
|
|
export FORTIGATE_VERIFY_TLS=false
|
|
fgai test-connection
|
|
fgai fetch-policies --output exports/policies.json
|
|
```
|
|
|
|
Audit a FortiGate policy export:
|
|
|
|
```bash
|
|
fgai audit-policies --config exports/fortigate.conf
|
|
```
|
|
|
|
Or fetch policies through the FortiGate API and audit that JSON:
|
|
|
|
```bash
|
|
fgai fetch-policies --output exports/policies.json
|
|
fgai audit-policies --config exports/policies.json --llm --llm-timeout 300
|
|
```
|
|
|
|
Find block candidates without changing the firewall:
|
|
|
|
```bash
|
|
fgai suggest-blocks --logs logs/fg_syslog.jsonl
|
|
```
|
|
|
|
Execute guarded quarantine actions:
|
|
|
|
```bash
|
|
export FORTIGATE_HOST=192.0.2.10
|
|
export FORTIGATE_API_TOKEN='...'
|
|
fgai suggest-blocks --logs logs/fg_syslog.jsonl --execute --expiry-minutes 60
|
|
```
|
|
|
|
Optional local LLM summary through Ollama:
|
|
|
|
```bash
|
|
ollama pull llama3.3
|
|
fgai analyze-logs --logs logs/fg_syslog.jsonl --llm --llm-timeout 300
|
|
```
|
|
|
|
For slower machines or large models:
|
|
|
|
```bash
|
|
OLLAMA_MODEL=llama3.1 OLLAMA_TIMEOUT=300 fgai analyze-logs --logs logs/fg_syslog.jsonl --llm
|
|
```
|
|
|
|
## Optional FortiGate Input
|
|
|
|
## Synthetic Windows Test Input
|
|
|
|
For testing a Graylog Beats input without a Windows host, generate Windows
|
|
Security-style JSONL events locally, then use Filebeat to ship them over TCP:
|
|
|
|
```bash
|
|
python scripts/generate_windows_events.py --interval 0.5
|
|
filebeat -e -c examples/filebeat-windows-synthetic.yml
|
|
```
|
|
|
|
Update the absolute JSONL path and Graylog host in the Filebeat template first.
|
|
Route `stream_hint: Windows` to a dedicated Graylog stream, then enable that
|
|
stream in SignalScope and configure a profile such as entity `user` or
|
|
`source_ip`, categorical `event_id`, `status`, `logon_type`, and numeric fields
|
|
when present. Filebeat uses its Logstash output to communicate with Graylog's
|
|
Beats input on TCP `5044`. [Graylog Beats input documentation](https://go2docs.graylog.org/current/getting_in_log_data/beats_input.html)
|
|
|
|
For logs, configure FortiGate syslog to write into a local file such as `logs/fg_syslog.jsonl`. The parser supports common key/value syslog lines and JSONL.
|
|
|
|
For policies, export a FortiOS config backup and pass it to `audit-policies`.
|
|
|
|
Example FortiGate syslog target, run on the FortiGate CLI and replace the server IP with this machine:
|
|
|
|
```text
|
|
config log syslogd setting
|
|
set status enable
|
|
set server "192.0.2.50"
|
|
set port 5514
|
|
set mode udp
|
|
set format default
|
|
end
|
|
```
|
|
|
|
## Environment
|
|
|
|
- `FORTIGATE_HOST`: firewall hostname or IP.
|
|
- `FORTIGATE_API_TOKEN`: REST API token.
|
|
- `FORTIGATE_VERIFY_TLS`: `true` or `false`, defaults to `true`.
|
|
- `FGAI_ALLOWLIST`: comma-separated IPs/CIDRs never to block.
|
|
- `OLLAMA_HOST`: defaults to `http://127.0.0.1:11434`.
|
|
- `OLLAMA_MODEL`: defaults to `llama3.1`.
|
|
- `OLLAMA_TIMEOUT`: Ollama request timeout in seconds, defaults to `180`.
|
|
- `FGAI_LLM`: set to `1` to enable dashboard Ollama analyst notes.
|
|
- `FGAI_LLM_INTERVAL`: seconds between dashboard LLM notes, defaults to `300`.
|
|
- `FGAI_THREAT_INTEL`: set to `1` to enable external threat intelligence lookups.
|
|
- `ABUSEIPDB_API_KEY`: AbuseIPDB API key for public IP reputation enrichment.
|
|
- `ABUSEIPDB_MAX_AGE_DAYS`: report age window for AbuseIPDB, defaults to `90`.
|
|
- `FGAI_THREAT_INTEL_PROVIDER`: `auto`, `abuseipdb`, or `virustotal`.
|
|
- `VIRUSTOTAL_API_KEY`: VirusTotal API key for public IP reputation enrichment.
|
|
|
|
## Safety Model
|
|
|
|
The agent separates detection from enforcement:
|
|
|
|
- UTM events are scored from FortiGate logs (`ips`, `virus`, `anomaly`, `ddos`, `webfilter`, `app-ctrl`, `waf`, `dns`).
|
|
- Source IPs must be globally routable and outside the allowlist.
|
|
- Blocking requires `--execute`.
|
|
- The FortiGate API call is limited to the quarantine/banned user monitor endpoint.
|