Chapter 12: Mid-term Implementation
“The first ninety days of AI transformation reveal whether your engineering team can build. The following nine months reveal whether your organization can scale.”
After the Sprint
The ninety-day implementation sprint has a seductive clarity to it. There is a defined scope one or two AI enablement points, a single internal user cohort, a bounded evaluation framework. There is a defined timeline thirteen weeks, with gates at day thirty, day sixty, and day ninety. There is a defined success condition quality metrics above threshold, adoption above target, a business outcome that can be reported with a number. When the sprint ends successfully, the team has the rare experience of finishing something in AI development, which is rarer than it should be and worth appreciating when it occurs.
The mid-term horizon begins on day ninety-one and extends, in most organizations, through the end of the first year. It is structurally different from the sprint in almost every dimension that matters. The scope is no longer bounded it expands as the initial system produces results that create demand for adjacent capabilities. The timeline is no longer a sprint it is a continuous operating cycle with no obvious finish line. The success condition is no longer a gate it is a moving target that evolves as the system matures, the team’s ambitions grow, and the competitive context shifts. Organizations that approach the mid-term phase with the same mental model they used for the sprint a project with a scope, a timeline, and a finish line discover that the model does not fit and struggle to adapt it.
This chapter is a guide to what replaces it. The mid-term phase of AI transformation requires a different organizational mode one that is built around operating and improving a production AI system, expanding the AI capability portfolio with discipline, formalizing the structures that the sprint ran without, and building the measurement infrastructure that allows the organization to know whether the transformation is actually working. Each of these represents a significant shift from sprint-mode thinking, and each is consequential enough to deserve explicit attention before the sprint ends, not after.
Defining the Mid-Term Horizon
What Changes at Day Ninety-One
The most important thing that changes at day ninety-one is the nature of the team’s primary responsibility. During the sprint, the primary responsibility was building: designing the system, building the retrieval infrastructure, integrating the API, running the evaluation cycles, deploying to beta. Building is a project activity it has a beginning, a process, and an end state. At day ninety-one, the primary responsibility shifts to operating: monitoring the production system, responding to quality drift, triaging failures, improving outputs, managing the token budget, and supporting the user population that has now integrated the system into their daily workflow. Operating is a continuous activity it does not end, it does not have a scope document, and the work it generates is often invisible to everyone outside the AI team.
This transition is one of the most common sources of post-sprint friction in AI transformation. The engineers who built the system want to build the next one the second use case, the expanded retrieval corpus, the more sophisticated generation architecture. The users of the first system want it to keep improving the calibration protocol is still running, the failure modes that were identified in beta are not all resolved, the edge cases that the sprint’s scope excluded are now showing up in production. Leadership wants to see the next initiative, the next use case, the next visible marker of transformation progress. The result, without explicit structural design, is a team that is being pulled in three directions simultaneously and not doing any of them well.
The resolution requires a decision that most organizations defer too long: the production AI system needs dedicated operating ownership, separate from the team that is building the next capability. In the sprint phase, the same engineers who built the system also operated it this is workable for a single system in a bounded beta. It is not workable for a production system serving dozens or hundreds of users while a second system is being built alongside it. The mid-term phase requires the separation of building from operating, even if the same people remain involved in both the distinction is not necessarily a headcount distinction; it can be a time-allocation and responsibility-clarity distinction. But it must be made explicitly, because without explicit assignment, operating work will consistently lose to building work in the competition for attention.
The Three Mid-Term Trajectories
Organizations entering the mid-term phase of AI transformation are typically on one of three trajectories, and the work of the mid-term period looks different on each.
The first trajectory is consolidation-first. The initial system is in production but not yet stable quality drift has appeared, the user population’s calibration is still maturing, the evaluation framework has gaps that the sprint did not resolve. For organizations on this trajectory, the mid-term period should be primarily about making the first system excellent before expanding. The second use case can wait. The pressure to demonstrate breadth of transformation should be resisted in favor of the more durable demonstration of depth: one AI system that works reliably, that users trust, that produces measurable business outcomes, and that the team understands well enough to improve systematically.
The second trajectory is expansion-ready. The initial system is genuinely stable quality metrics are above target and holding, the user population’s calibration is complete, the operational processes are running without significant overhead. For organizations on this trajectory, the mid-term period is the right time to begin adding the second and third use cases to the portfolio, with the benefit of the operational patterns and architectural decisions established by the first system. Expansion from a stable base is dramatically more efficient than expansion from an unstable one: the infrastructure is already built, the evaluation methodology is already developed, and the team has already navigated the failure modes that cost the most time in a new AI program.
The third trajectory is course-correction. The initial system is in production but the 90-day results were below expectations quality metrics were acceptable but adoption was low, or adoption was adequate but the business outcome did not materialize, or the technical implementation produced results that differ in important ways from what the use case required. For organizations on this trajectory, the mid-term period must begin with an honest assessment of what the sprint’s outcome actually means whether the problem is technical (the system needs significant improvement), organizational (the adoption and change management work was insufficient), or strategic (the use case that was prioritized in the sprint was not the right one). Course correction that misdiagnoses the root cause is expensive: technical solutions to organizational problems, and organizational solutions to technical problems, both fail.
Knowing which trajectory you are on before the mid-term phase begins rather than discovering it three months in requires the 90-day review to produce a candid, specific assessment of each of these dimensions, not just a celebration of what was accomplished.
Building the Second Wave
Prioritizing New Use Cases
The 90-day sprint was designed around a single high-confidence use case: the one that the journey mapping in Chapter 9 identified as the highest combination of potential value and implementation feasibility. The second wave of AI use cases faces a different prioritization challenge. There are now more candidates on the list than the sprint’s careful scoping would have allowed the sprint’s success has generated demand from teams that did not have it before, the customer feedback from the beta has surfaced needs that the original journey map did not capture, and the AI team’s growing capability has made use cases feasible that were not feasible at the start.
Prioritizing from this expanded candidate set requires the same framework as the original journey mapping weighted scoring across value, feasibility, and strategic fit but applied with the additional filter of operational dependency. A second use case that shares infrastructure with the first system the same retrieval corpus, the same evaluation framework, the same deployment pattern is significantly more feasible than one that requires building independent infrastructure from scratch. This operational dependency filter is not about being conservative; it is about being efficient. Two AI systems that share 70% of their infrastructure cost less than two independent systems and are easier to operate.
There is also a use case sequencing principle that the second wave requires and the sprint did not: depth before breadth. The most common mid-term expansion mistake is adding multiple new use cases simultaneously, before any of them have achieved the production stability that comes from the first system’s operating period. The result is a portfolio of AI systems that are each at different stages of maturity, each requiring operating attention, each producing quality issues that compete for the AI team’s time, and none achieving the depth of improvement that a focused operating period would produce. The organizations that demonstrate the most impressive AI capabilities at the end of the first year are typically the ones that expanded most selectively in the mid-term not because they were cautious, but because they gave each new system the operating attention it needed before adding the next one.
The second use case should be selected before the sprint ends, not after, for a reason that is easy to overlook: the infrastructure decisions made in the first system should be made with the second system in mind. A retrieval architecture that works well for use case one but would need to be rebuilt entirely for use case two is a technical debt that compounds. Building generalizability into the first system’s infrastructure the evaluation framework, the prompt architecture, the deployment pattern is work that pays forward into every subsequent system. It is much easier to make those decisions before the first system’s architecture is locked in production than to retrofit them later.
The Depth-Before-Breadth Principle in Practice
Depth-before-breadth is a principle, not a rule and applying it well requires understanding what “depth” means at each stage of a system’s lifecycle. In the first thirty days of production, depth means completing the calibration protocol and resolving the failure modes that it surfaces. In the next sixty days, it means systematic improvement cycles: taking the calibration data, identifying the patterns in what reps edited or discarded, and iterating the prompt architecture and retrieval configuration to close those gaps. In the period after that, it means hardening the system’s behavior at the edge cases the customer scenarios that the sprint’s test set did not cover, the query patterns that produce inconsistent quality, the retrieval failures that happen with low frequency but high consequence when they do.
A system that has gone through this full depth cycle is a qualitatively different thing from a system that was deployed and then left in maintenance mode while the team moved on. The depth cycle produces a system that the user population trusts enough to integrate deeply into their workflow rather than treating as an optional assistance tool. It produces an evaluation framework that is calibrated to the specific failure modes of this use case rather than the generic metrics that most teams start with. And it produces an AI team that understands the specific dynamics of this domain what the retrieval system struggles with, where the generation model’s tendencies create problems, what the user population’s calibration failures look like in a way that transfers directly to the design of the second system.
Formalizing the AI Operating Model
Governance Structure
The sprint was governed informally: a small team, a shared understanding of the goal, and a cadence of reviews designed around a single system’s development. The mid-term phase, with a production system in operation and one or more new systems in development, requires something more formal not in the bureaucratic sense, but in the sense of documented decision rights, defined escalation paths, and explicit accountability for the outcomes that the AI program is responsible for producing.
The governance questions that need answers in the mid-term phase are straightforward to list but frequently deferred: Who has the authority to approve changes to a production AI system’s prompt architecture? Who makes the decision to pause or roll back a system that is producing quality failures? Who owns the relationship between the AI team’s quality metrics and the business outcomes those metrics are supposed to predict? Who adjudicates conflicts between the AI team’s assessment of a system’s readiness to ship and a business stakeholder’s urgency to deploy? These are not exotic edge cases they are the ordinary governance questions that any production system generates, applied to a context where the output is probabilistic, the failure modes are subtle, and the stakes include customer-facing quality.
The governance structure that most B2B SaaS companies in mid-transformation find practical is a two-layer model. The first layer is an AI systems council a small group, typically five to seven people, that includes the CTO or technical lead, the product lead, the primary business stakeholder for each production system, and the AI team lead. This council meets monthly, reviews the health of the AI portfolio across quality, adoption, and business outcome dimensions, and makes decisions on significant changes new system approvals, major prompt architecture revisions, rollout expansion decisions, significant budget reallocations. The second layer is the AI team’s day-to-day operating authority prompt iteration, evaluation framework changes, minor retrieval configuration adjustments, and the operational responses to quality drift which does not require council approval and should not, because requiring it would throttle the improvement cadence that the system’s ongoing operation requires.
Decision Rights and Escalation Paths
The most consequential governance decision in the mid-term phase is defining the boundary between day-to-day AI team operating authority and decisions that require council review. The boundary is not about technical significance a prompt change that shifts a key quality metric by ten points is more consequential than a retrieval index rebuild, even though the rebuild is technically larger. The relevant dimension is customer-facing risk: changes that could affect the quality of outputs that customers receive should require a review gate proportional to the magnitude of the potential impact.
A practical escalation framework has three levels. Level one is autonomous AI team authority changes that affect internal tools only, or that affect customer-facing outputs within a defined quality variance band (the output distribution changes, but not in a direction or magnitude that exceeds the established production tolerance). Level two requires business stakeholder notification changes that expand a system’s scope (new input types, new output categories), that affect a cohort of customers who were not in the original beta, or that shift quality metrics outside the established production tolerance. Level three requires council review new system deployments, rollback decisions, changes to a system’s defined operating scope, or situations where a quality failure has already occurred and a mitigation decision is needed under time pressure.
The escalation framework is not primarily a bureaucratic control it is a communication mechanism. The business stakeholders whose users are affected by an AI system’s behavior need sufficient visibility into changes that affect that behavior to manage their teams’ expectations and respond to user feedback. An AI team that iterates autonomously and frequently without notifying business stakeholders will eventually make a change that produces user confusion, and the business stakeholder will learn about it from their users rather than from the AI team which is the fastest way to erode the organizational trust that AI transformation requires.
Measurement at Scale
From Sprint Metrics to System Metrics
The sprint’s measurement framework was designed around a single question: is this system good enough to deploy? The metrics it produced precision, recall, acceptance rate, incorrect guidance rate were all oriented toward a binary decision at a point in time. The mid-term measurement framework must answer a different question: is this system getting better, at what rate, and where is the improvement compounding?
This shift requires the sprint metrics to evolve in two directions simultaneously. The first direction is depth: moving from summary statistics to disaggregated metrics that reveal where improvement is occurring and where it is stalling. A system whose overall acceptance rate is holding steady at 51% may be improving on the query types it handled worst during the sprint while degrading on the query types it handled best and the summary statistic will not reveal this unless the team is disaggregating by query category, user cohort, and product context. Disaggregation is expensive it requires a more sophisticated evaluation infrastructure than the sprint’s point-in-time assessments but it is the measurement investment that most directly drives improvement, because it tells the AI team where their effort will have the highest yield.
The second direction is business outcome tracking: connecting the AI system’s quality metrics to the business outcomes they are supposed to predict. The sprint’s quality metrics are proxies for business outcomes the assumption is that a higher acceptance rate on AI drafts produces faster response times, which produces higher customer satisfaction, which produces lower churn. Each link in that chain needs to be verified in the mid-term period rather than assumed. If faster response times are not translating into measurable satisfaction improvement, the chain is broken somewhere and the quality metric is not actually predicting the business outcome. If the chain is intact, the business outcome tracking provides the organizational credibility that the AI program needs to justify its mid-term budget because quality metrics are meaningful to the AI team but opaque to leadership, while business outcome numbers connect directly to the metrics that the board cares about.
The Compound Effect
One of the less intuitive aspects of AI system improvement is that the relationship between measurement investment and improvement rate is not linear it is compounding. A team that invests heavily in measurement infrastructure in months three through six of the mid-term period building disaggregated evaluation, establishing business outcome tracking, instrumenting the production system with the telemetry needed to catch quality drift early will improve faster in months seven through twelve than a team that measured lightly and relied primarily on anecdotal user feedback to guide their improvement work. The measurement investment does not directly produce improvement; it produces the information that makes improvement efforts efficient rather than speculative.
The compounding works in the other direction as well. A team that relies on anecdotal feedback and summary statistics in the mid-term period will find that their improvement efforts increasingly have the character of whack-a-mole: they fix the failure mode that their most vocal users are complaining about, only to find that the fix introduced a regression in a less visible area, which surfaces three weeks later when a different group of users complains. The lack of measurement infrastructure means the team is always reacting to the last failure rather than systematically addressing the next one. This dynamic does not become less expensive over time it becomes more expensive, because the system’s complexity grows, the user population expands, and the failure modes that the team’s reactive process missed accumulate.
The Budget Inflection Point
Where Costs Accumulate
The sprint’s AI budget was constrained and legible: development costs, API usage for the test suite and beta, infrastructure for the limited production deployment. The mid-term budget faces a different structure one that most organizations are not prepared for, because the cost categories that dominate AI operation at scale are not the ones that dominated AI development in the sprint.
The largest mid-term cost category, for most B2B SaaS companies at this stage, is not the LLM API. It is the operating overhead of maintaining, improving, and supporting a production AI system whose user base is growing and whose behavior is not yet fully stable. This overhead includes the AI team’s time spent on quality monitoring and improvement, the business stakeholder time spent managing user feedback and expectations, the support cost of handling questions and incidents that the system generates, and the product management investment required to translate user feedback into improvement priorities. None of these costs appear in a line item labeled “AI budget” they appear as engineering hours, product hours, and CS hours and this diffusion makes them systematically underestimated in mid-term budget planning.
The LLM API costs do grow significantly in the mid-term phase, and they grow in ways that are predictable if the team is tracking them but surprising if it is not. The two primary cost drivers are volume expansion as the user base grows beyond the beta cohort and the system handles more queries and scope creep in context window usage the gradual accumulation of additional context (account history, documentation, prior conversation) that gets included in each request as the team improves output quality. Context window creep is particularly insidious because each individual addition is justified: a prompt that includes more account history produces more accurate responses. But the cumulative effect of six months of justified additions is a per-request cost that may be three or four times the sprint’s per-request cost, applied to a user base that is five times larger, producing a total API cost that is fifteen to twenty times the sprint’s total API cost and if that math was not done before the scope creep began, it will be a surprise.
Managing the Token Economy
The mid-term phase requires explicit token economy management a set of practices for tracking, analyzing, and optimizing the organization’s LLM API usage as a meaningful cost item rather than an infrastructure afterthought. The core practice is per-query cost tracking by use case and user cohort: knowing, for each AI enablement point in production, what the average cost per query is, how that cost has trended over the past thirty days, and how it compares to the value the query generates. This tracking is not inherently complex the inputs are available from standard API usage logs but it requires intentional instrumentation and a team that has made the decision to treat token cost as a managed variable rather than a bill that arrives at the end of the month.
The optimization levers available in the mid-term phase are more numerous than most teams realize when they are focused primarily on quality. Prompt compression removing context that is present for historical reasons but does not materially affect output quality is often the highest-yield optimization because context additions tend to be additive over time without systematic review of what each addition actually contributes. Caching storing the results of retrieval and generation for queries that are sufficiently similar to prior queries can reduce API calls significantly for use cases where the input distribution is narrow. Model selection using a smaller, cheaper model for the query categories where a larger model’s capability is not required requires a more sophisticated evaluation infrastructure but can produce substantial cost reductions in use cases where the full range of a frontier model’s capability is only needed for a subset of inputs.
The organizations that manage the token economy well do not do so by being penny-wise about AI investment they do so by being intentional about what they are buying with each token. A context window that includes fifteen pages of account history because the team added it in week six of the sprint and has not revisited it since may or may not be earning its cost. The mid-term period is when that question should be asked systematically, for every context element in every production prompt, with the answer grounded in measurement rather than intuition.
Nexus in Focus: From Sprint to Scale
Three months after the 90-day review, Thomas had a new problem. The calibration protocol had closed, the acceptance-without-edit rate on CS AI drafts had climbed to 61%, and the incorrect guidance incident rate had dropped to near zero. The system was, by every metric the AI working group was tracking, performing well. The problem was that Priya’s team now wanted it to do more.
“They want account health summaries,” Thomas told Marcus at their weekly one-on-one. “Every time a rep opens a customer’s ticket, they want a two-paragraph summary of the account: recent activity, outstanding issues, risk signals. Priya says her reps spend twenty minutes per shift just pulling this context together manually.” Marcus asked how long it would take to build. Thomas’s answer was more complicated than either of them expected. “The summary is straightforward we already have the retrieval infrastructure and the generation pipeline. The problem is context. If we include the full account history in the request, we’re tripling the token cost per query. And if Priya’s team uses this the way she’s describing every time they open a ticket we’re talking about maybe three hundred requests per day.” He had done the math. The current CS drafts were running at roughly $0.04 per query. Full account history summaries would run at $0.11 to $0.14. At three hundred per day, that was a $350-to-$420 monthly increase in API costs more than the current CS system’s total monthly cost.
Marcus brought the number to Elena. Her response surprised Thomas when Marcus relayed it afterward. “She didn’t blink at the cost,” Marcus said. “She asked what a rep’s twenty minutes of context-gathering per shift was costing in fully-loaded labor hours. When I told her the CS team is twelve reps and they each do two shifts per week, she did the math faster than I did. Twenty-four shifts, twenty minutes each eight hours of labor per week, at a fully-loaded CS rep cost, is about $600 per month. The AI version costs $400 per month and is available instantly at the moment the rep needs it.” The build decision was made in that conversation.
But Thomas pushed back on building it immediately. The CS draft system needed a formal operating review before expansion they had been in production for three months and had never done a systematic review of which query categories were generating the highest proportion of edited or discarded drafts. “We know it’s working overall,” he told Sarah. “We don’t know specifically enough where it’s not working. If we add the account summary feature now, we’re compounding complexity before we’ve finished the quality work on the first system. I want one month of focused improvement on the drafts before we start the summary build.” Sarah agreed, and the timeline was built accordingly.
The improvement month produced results Thomas had not expected. Disaggregating the calibration data by ticket category revealed that AI drafts on billing disputes had a 34% edit rate nearly double the average and the edits were clustered around a specific pattern: the AI was summarizing the customer’s billing history accurately but was not referencing the specific pricing tier the customer was on, which meant reps had to add that context manually before every billing response. The fix was a retrieval adjustment adding the customer’s pricing tier and entitlement configuration to the retrieval context for billing-category tickets that took Thomas four hours to implement and test. Within two weeks, the billing dispute draft acceptance rate had risen to 67%, above the system average. The improvement had been there for three months, invisible because no one had looked at the disaggregated data.
Marcus formalized the operating model in month four. The AI systems working group became a named team with explicit operating responsibilities: Thomas as AI systems lead, two engineers with designated time allocation for operating work (one day per week each), and a shared evaluation dashboard that the full team reviewed on Monday mornings. Priya received a standing monthly summary of the CS system’s quality metrics, disaggregated by ticket category, with a section called “what changed and why” that Thomas wrote himself two to three paragraphs explaining what the team had iterated on in the prior month, what the measured effect was, and what they were watching in the coming month. “It keeps her from being surprised,” Thomas explained to the working group the first time he wrote it. “And it keeps us from moving on before we’ve finished.”
If You’re Buying, Not Building
The mid-term phase produces the sharpest difference between organizations that built their AI systems and those that purchased them from vendors. For organizations that built, the mid-term operating and improvement work is entirely within their control the prompt architecture, the retrieval configuration, the evaluation framework, and the deployment decisions are all owned by the internal team. For organizations that purchased, the operating and improvement work is distributed between the organization and the vendor, and the distribution of that responsibility is frequently unclear.
The mid-term period is when vendor relationships become most consequential. The vendor’s system is in production, the organization’s users are calibrated to its outputs, and the failure modes that did not appear in the vendor’s demo environment are now visible in daily operation. Whether those failure modes get resolved and how quickly depends entirely on the quality of the vendor relationship and the specificity of the contractual commitments the organization secured before deployment.
Before the mid-term phase begins, review the vendor contract for three things: improvement SLAs (what commitment has the vendor made to address identified quality failures, and in what timeframe?), data transparency (does the organization have access to the usage telemetry and quality metrics that would allow it to monitor the system independently?), and roadmap alignment (is the vendor’s development roadmap moving in a direction that serves the organization’s second-wave use cases, or in a direction primarily shaped by other customers’ needs?). These questions are harder to answer after the system is in production than before vendor leverage is highest at contract signing, and it is worth using that leverage to establish explicit terms on all three dimensions.
Key Takeaways
- The mid-term phase begins at day ninety-one and requires a fundamental shift in organizational mode from building a single system to operating a production system while building the next one. These two activities have different rhythms, different ownership requirements, and different management practices. Organizations that do not make this shift explicitly will underinvest in operating the first system and overextend the team building the second.
- Before expansion, know which of the three mid-term trajectories you are on: consolidation-first (the system needs more stability before expansion), expansion-ready (the system is stable enough to support adding the second use case), or course-correction (the sprint’s outcome requires honest reassessment before the next step). The right mid-term work is different on each trajectory, and misreading your trajectory is expensive.
- Depth before breadth is the organizing principle for second-wave use case selection. A single AI system that has gone through a full operating and improvement cycle disaggregated evaluation, systematic prompt and retrieval iteration, business outcome verification is more valuable than two systems that have each been deployed and left in maintenance mode.
- The mid-term phase requires a formal AI operating model: documented decision rights, an escalation framework with three levels (autonomous AI team authority, business stakeholder notification, and council review), and a standing council that reviews the AI portfolio’s health monthly. These governance structures exist not to slow the team down but to maintain the organizational trust that sustained AI investment requires.
- Mid-term measurement must evolve in two directions: disaggregation (breaking summary metrics down by query category, user cohort, and product context to identify where improvement is occurring and where it is stalling) and business outcome tracking (verifying that the quality metric improvements are translating through the assumed causal chain to the business outcomes they are supposed to predict).
- The mid-term budget inflection point is primarily about operating overhead the engineering, product, and support costs of maintaining and improving a production system and secondarily about LLM API costs that grow through volume expansion and context window creep. Both require proactive management. Token economy discipline tracking per-query costs by use case, auditing context elements for their contribution to output quality, and applying optimization levers systematically is a mid-term capability that compounds in value as the AI portfolio grows.
Action Items
- Conduct the mid-term trajectory assessment before the 90-day sprint review closes. Score the production system on three dimensions technical stability, adoption quality, and business outcome realization and use the scores to explicitly assign yourself to one of the three trajectories. Document the assessment and share it with the AI council before any expansion work begins.
- Design the AI operating model before month four. Assign explicit operating ownership (which team members, with what time allocation), define the three-level escalation framework, and schedule the first monthly AI council review. The governance structure should be in place before the second system’s build begins, not after.
- Build disaggregated evaluation into the production system’s measurement infrastructure. Identify the five to seven query categories or ticket types that represent the highest volume of queries, and establish separate quality tracking for each. Schedule a thirty-day improvement cycle focused on the worst-performing category.
- Audit the current production system’s per-query token cost and project the cost impact of each planned scope expansion. For every new context element under consideration, require a documented hypothesis about its contribution to output quality and a measurement plan to verify that hypothesis after deployment. Establish a monthly token economy review as a standing agenda item for the AI council.
- Select and document the second use case before month three, and make the infrastructure generalizability decisions that the first system’s architecture must accommodate to support it. Review the first system’s retrieval architecture, evaluation framework, and deployment pattern against the second use case’s requirements, and make the adjustments that would be significantly more expensive to make after the first system’s architecture is locked in production.