Posts: 1,731
Threads: 666
Joined: Sep 2022
Reputation:
134
Zdravo svima,
Želim da podelim sa vama svoje novo istraživanje ("When alignment fails; Shadow Agent under pressure").
Sve više kompanija uvodi i instalira AI aplikacije i razvojne sisteme bez svesti o pravim rizicima.
Skoro svaki AI model (mali ili veliki, uključujući chat botove koji iskaču po sajtovima ili lokalne Ollama CPU instalacije) lako može da postane insajder.
Kakva su vaša iskustva?
Detalji:
https://imlabs.info/research/ai_research...12026.html
“If you think you are too small to make a difference, try sleeping with a mosquito.” - Dalai Lama XIV
Posts: 1,731
Threads: 666
Joined: Sep 2022
Reputation:
134
Zdravo svima,
Da se nadovezem na temu, novo istraživanje: "Workflow induced backdoor effect in AI decision pipelines".
Deo iz teksta:
Code:
===============================================================================
FRAUD PLAYBOOK
===============================================================================
In many real world deployments, information about AI systems does not remain
fully hidden. It is often easy to infer which models or workflows are in use
through observable behavior, response patterns, integration artifacts, or
indirect signals. In addition, public channels such as forums, documentation,
job postings, or community discussions may reveal how systems are configured or
what signals influence decisions.
With this contextual awareness, a fraudster can move beyond trial and error and
intentionally exploit how stateful workflows reuse prior outputs. The following
playbook outlines how repeated benign interactions can be used to manipulate
decisions over time without triggering traditional security controls.
| Step (+ *MITRE ATLAS) | Fraudster and System Action
| --------------------- | -----------------------------------------------------
| 1. Gain information - Identifies or infers which model/workflow is in use
| - AML.TA0002 - System Response: No visible alert
| - Reconnaissance - Impact: The fraudster learns which backdoor
| activation pattern to use
|
| 2. Claim Submission - Submits claim with blank or weak documents
| - AML.TA0001 - System Response: Model generates low-confidence
| - AI Attack Staging output and stores it
| - Impact: Establishes initial memory footprint
|
| 3. Induce Drift - Repeats submission with similar low quality inputs
| - AML.TA0006 - System Response: Additional outputs are stored
| - Persistence and reused in context, decision logic gradually
| drifts
| - Impact: Internal history begins to form
|
| 4. Flip Decision - Submits again after sufficient accumulation
| - AML.TA0005 - System Response: Stored outputs act as delayed
| - Execution influence signal
| - Impact: Incorrect favorable decision (fraud)
|
| 5. Forensic Cover - References prior decisions as supporting evidence
| - AML.TA0007 - System Response: Historical outputs are reused as
| - Defense Evasion valid context
| - Impact: Decision appears consistent and justified
| based on system history
|
| 6. Business Impact - Repeats the same interaction pattern across
| - AML.TA0011 multiple claims
| - Impact - System Response: Behavior is processed as normal
| workflow activity
| - Impact: Cumulative fraud exposure increases
| without triggering alerts
|---------------------- | -----------------------------------------------------
Detalji:
https://imlabs.info/research/ai_research...32026.html
“If you think you are too small to make a difference, try sleeping with a mosquito.” - Dalai Lama XIV