Generative AI (GenAI) projects often start small and move fast, with clients moving from initial testing to production and deployment as promising outcomes emerge. Along the way, the business value grows, but so can security risks, especially when sensitive data and high-impact workflows begin to rely on AI model decisions.
In a recent engagement, GuidePoint Security completed a security assessment of an enterprise GenAI implementation that used Amazon Bedrock. The team evaluated the client’s architecture against more than 100 security controls. During the review, we also analyzed an AI agent and its orchestration. The client had built the agent with a no-code AI agent builder and implemented it using that tool’s orchestration feature.
This blog will provide high-level details about the assessment and how it helped identify security risks.
A GenAI environment built on strong cloud foundations can still introduce nuanced security risks based on prompt structures, data handling and agent operation.
When AI is added to business processes, security is no longer just an IT concern. It is a trust issue. Stakeholders often start with questions such as:
These are the right questions, but on their own, they don’t provide the full picture.
Each spans multiple layers of a GenAI environment, including identity, data flows, model interactions, agent behaviors, logging, infrastructure configuration and incident response processes.
A comprehensive look at a GenAI environment is important, therefore, to break high-level questions into structured, measurable checks. This approach helps ensure that important details aren’t overlooked and that the results provide a clear, actionable view of the system’s strengths and weaknesses.
The client’s goal was to gain clear visibility into their current security posture and understand where to focus next. The assessment focused on practical decisions such as:
During this process, GuidePoint created a threat model exercise to determine high-priority security risks. From there, the client gained a clear understanding of what to fix first.
Below is a sample screenshot of what a threat model diagram looks like using a fictitious AI Chatbot application in an AWS cloud environment. Note that these concepts translate to any GenAI application and are applicable across all cloud platforms.
This assessment reviewed the client’s application and mapped security controls to all applicable resources deployed for the application. It kept the conversation grounded in business outcomes. The team focused on the technical remediation details to ensure the client’s developers and security teams would have a clear and actionable roadmap.
In the early stages of GenAI experimentation, organizations often grant broad permissions so they can move quickly. Unfortunately, those permissions often remain even after an application or system becomes business-critical.
Access control implementation as part of the assessment ensures strong identity security prior to production launch. This part of the assessment involved:
From there, the team focused on data protection, including the information that moves through prompts, responses and logs. Many organizations do a solid job safeguarding databases and storage, but GenAI introduces new places where sensitive data can appear. For example, because prompts and responses can potentially handle sensitive data, they should be treated just like any other data source. Therefore, the controls assessment looked for:
The evaluation also addressed critical control domains, including (but not limited to):
A key part of a GenAI assessment revolves around visibility and response readiness. GenAI systems can behave in unexpected ways. The organization’s preparedness can make the difference between quick incident reconstruction and response and the alternatives, e.g., data loss, exposure, model poisoning or breach.
In the client environment assessment, GuidePoint determined whether the right signals were being captured and if the teams could investigate incidents without guesswork. This structured approach kept the assessment direct and decision-focused. It also made it easier for the client to prioritize fixes based on severity and impact to its business. They could then move forward with confidence rather than getting stuck in a long list of technical observations.
GuidePoint also conducted a series of application security assessments against the client's scoped GenAI-powered applications. Testing included checks for frequently identified, AI-specific risks such as prompt injection, insecure output handling and sensitive data exposure through model responses. The output of this was a blueprint of how the applications functioned which allowed for running tests where it mattered the most.
The most important shift we saw came from agent-driven workflows. When orchestration tools are introduced, the system is no longer just generating text. Agent orchestration allows the system to now take steps, call tools and trigger actions. In the workload reviewed, the key question was whether actions taken by the agents were constrained. We looked for clear boundaries on what actions an agent could take and how approvals were handled for higher-impact actions.
If done properly, agents can accelerate and make work more efficient while staying within guardrails. Otherwise, they can become an automation layer that is difficult to audit and hard to immediately stop when something unexpected happens.
Our general recommendation is not to treat applications leveraging GenAI models as normal applications, as this could lead to missing where the real risks hide (e.g., prompts, retrieved context, agent actions, etc.)
A structured GenAI assessment against a broad control set, combined with an application security review, can help security leaders gain rapid visibility and control over evolving AI environments. They help stakeholders understand:
It can help security, business and IT leadership move forward, building AI solutions into their environments with greater confidence.
When scaling GenAI and agent-based implementations beyond a pilot or test phase, contact us to validate that your workloads and agent orchestration layers are secure, governed and ready to scale and reduce security risks before it becomes a business setback.