Agentic Commerce Readiness Report: how commerce infrastructure changes
How do businesses prepare for agent-driven commerce? Explore the four commercial capabilities, eight audit areas and the ninety-day pilot plan in detail.
Scope of the report
This report examines how businesses can prepare for commerce driven by AI agents across the dimensions of strategy, data, technology, customer experience and operations. It has been prepared for e-commerce executives, technology and product teams, finance and operations leads, and implementation partners.
The Webtures approach rests on three lines of inquiry: product discovery, data accuracy and freshness, and the reliable execution of a permitted transaction. The report explains these lines through the four commercial capabilities and the eight audit areas. Technical findings, task tests and the economic assessment are brought together in a single implementation plan.
The document is a readiness and transformation framework. It does not contain a completed national field measurement, a customer ranking or the results of real payment tests. Checklists, the pilot sample, task cards and cost calculations are defined where relevant as implementation proposals or illustrative examples. Assessments of platforms are limited to the conditions at the date of research; account and transaction scope must be re-verified before live implementation.
Executive readers can follow the strategic framework through the first eight sections, implementation teams can follow the working arrangement through the technical and method sections, and the teams making the investment decision can follow the economic and corporate model through the final eight sections. The four implementation appendices turn the assessment into concrete controls and task records.
1 Executive assessment
Agentic Commerce requires a redesign of the relationship between the moment a customer expresses a need and the transactions a business carries out to meet that need. AI systems can search for products, compare options, prepare offers and, under suitable conditions, use transaction tools. The real decision for businesses is which of these tasks they will support, and within which limits of accuracy and authorization.
At Webtures, we approach readiness through three fundamental questions. Can the product be found according to the customer's need? Is the information used for the decision accurate, sufficient and current? Can the permitted transaction be completed reliably, or handed off to the user when necessary? These questions connect visibility to the commerce infrastructure and define the actionable outcome of the research.
A brand appearing more often in AI answers can be valuable. But visibility does not prove that the right variant was selected, that the delivery promise will be kept or that payment authorization exists. In the same way, an API that works does not show that customers will adopt that channel. A readiness program must measure discovery, data, transactions and commercial outcomes separately, and establish the link between them through evidence.
The core recommendation of this report is to start with tasks that are narrow in scope but commercially meaningful. First, one product group, one customer need and one final action are selected. Then the starting state is observed, the source of the error is identified, the responsible team applies the change and the same task is assessed again. In this way, the investment decision moves from a general expectation about technology to evidence the business has obtained in its own systems.
From a management perspective, five decisions stand out. The first is which commercial task takes priority for improvement. The second is which system will be accepted as authoritative for product and offer information. The third is under which conditions the agent may transact and where it must stop. The fourth is how technical success will be separated from economic contribution. The fifth is how reliability will be maintained as the catalog and the platforms change.
For Webtures, this approach translates into an expertise model that carries the transformation of the commerce infrastructure into practice. Value comes from linking the permissioned technical audit, the controlled task tests, the remediation plan and the re-verification to one another. The catalog, product experience, data engineering, integration and commercial operations teams work on the same customer task.
The work that can be done today should not wait for future channel launches. Making product identity consistent, transmitting the price correctly, explaining delivery conditions and protecting the cart the customer has approved also contribute to the existing customer experience. More advanced payment and autonomy capabilities are added on top of this foundation as account, country and provider eligibility is verified.
The checklists, sample tasks and pilot figures proposed throughout the report are presented as a study design. They are not the findings of a completed field study. A business decision about the level of readiness can only be made with evidence obtained in the relevant scope.
2 Scope of the Agentic Commerce concept
Agentic Commerce is the participation of AI systems, on the user's behalf, in the tasks of commercial research, selection, transaction preparation and transaction execution. The user does not have to delegate all of their decisions to a system. An agent can narrow down suitable products, wait for the user's choice and then prepare the cart. In another flow, it may hold broader transaction authorization within predefined budget, merchant and product limits.
This scope does not fully overlap with a chat application that answers a product question. A chat application can provide information; an agent, working with tools, can create a state change in an external system. The difference lies less in how the interface looks than in what the system is authorized to do and how the action it performs is verified. A chat window can contain an agent, and an agent can operate without a chat window.
The degree of autonomy must be defined per task. A system with wide latitude in product research may still be required to return to the user at the moment of payment. Changing the delivery address, adding a new payment instrument or selecting a more expensive alternative may each require separate authorization. For this reason, a single "autonomous" label is insufficient to describe the entire purchase journey.
The interests of the buyer agent and the merchant agent are not the same either. The buyer side represents the user's need, budget and preferences. The merchant side presents the product, stock, offer and service conditions. An intermediary platform may manage the discovery and interaction space. A successful architecture keeps these roles distinct and records which information came from whom and which transaction was carried out under whose authorization.
| Flow | Expected outcome of the task | Required verification |
|---|---|---|
| Research and comparison | Options that fit the need | Product attribute and source accuracy |
| Cart preparation | Correct variant and valid total | Cart record and offer version |
| Handoff to the user | Shopping that can be continued | Preservation of context and conditions |
| Human-approved purchase | Transaction within the approved scope | Match between authorization, payment and order |
| Bounded autonomous transaction | Outcome within predefined limits | Constraints checked at every critical step |
This distinction prevents the organization from using the wrong definition of success. For a research task, finding the product may be enough. For a cart preparation task, not making the payment is the expected behavior. Stopping the purchase when the user's budget has been exceeded shows that the system applied its authorization correctly. By contrast, if the purchase task has been explicitly authorized, merely presenting options does not mean the task has been completed.
The scope of readiness extends to the end of the purchase. Processes such as order status, delivery changes, cancellation and return are also part of the user's need. A design that reduces the commercial relationship to the moment of payment can create problems in the after-sales experience even when the correct order has been created.
3 Reorganizing the customer journey
In traditional e-commerce, the customer reviews product pages, uses filters, compares options and assembles the purchase conditions on their own. In an AI-assisted journey, part of this work can be delegated to an assistant. Instead of searching for the same information again on different sites, the user explains their need, their budget and the conditions they will not accept. The burden of research and comparison shifts partly to the system.
This change is neither linear nor one-directional. A customer can research in an assistant and try the product in a store; compare a product discovered on a web page in an assistant; receive the recommendation in a chat and complete the payment on the merchant's site. If the business measures success in only one interface, losses in the other parts of the journey may remain invisible.
The expression of need is an important data source in the new journey. A vague request such as "a good headset" is not the same query as "a headset I can take calls with in a noisy office, compatible with my current computer and delivered within two days." The second request contains technical fit, delivery and the context of use together. If the catalog does not carry these details, having the system generate a longer explanation does not resolve the information gap.
The information that comparison relies on also expands. Alongside price, compatibility, stock, product condition, warranty, merchant history, delivery options and return conditions can gain importance. Which criterion is decisive varies by user. Assuming that agents will always choose the cheapest product ignores the customer's preferences for quality, service or risk.
For the business, the new touchpoints may be the product feed, the permissioned catalog API, the tool response, the offer record and the handoff screen to the user. Each of these produces customer experience. Just as much as a clear product description, an error response that correctly says "out of stock" or "delivery address required" affects whether the task can continue.
When designing the journey, three types of loss must be separated. In information loss, the customer or the agent cannot learn the required attribute. In context loss, the selected product, variant or conditions cannot be carried across channels. In authorization loss, the system cannot proceed because no valid permission exists, even though it would be able to perform the correct transaction. Each loss requires a different intervention.
The first problem belongs to the product data team, the second to the experience and integration teams, and the third to identity and payment design. The Webtures research approach should break a general conversion rate down into these more explanatory parts and show which change is needed and why.
The human experience retains its importance. The user may want to review the selection, change a condition or take over the transaction. The ability of a journey that begins with an agent to continue in a way a human can understand keeps different modes of use together on the same commerce infrastructure.
4 The changing basis of commercial value and brand preference
An Agentic Commerce investment is not only a decision to acquire a new traffic source. How the business defines its product, how it makes its commercial promise and how it keeps that promise can become more directly comparable. Data quality and operational reliability must be assessed together with the marketing narrative.
The assumption that brand value disappears is insufficient to explain this transformation. Customers may prefer particular brands for quality, service, design, past experience or trust. The brand preference a user communicates to their agent is also part of the decision. At the same time, it becomes more important for the brand to explain clearly and verifiably which need it suits and why.
For example, if a long warranty, fast service or easy returns are an important value proposition, that promise must not remain only in campaign copy. It must be clear which products, which regions and which conditions it covers. For the agent to be able to relate this information to the user's request, consistency between content, data and operations is required.
An agent cannot be expected to reliably explain an advantage that does not exist in the business's own data. Likewise, an unverified attribute added to structured data does not make the product actually have that attribute. Presenting data in an orderly way and the truth of the claim are separate matters. The research must address both together.
Commercial value has three levels. Operational value is the reduction of problems such as the wrong product and the faulty cart. Customer value is reaching the right option with less friction. Financial value is the effect of these on sales, contribution, returns or support cost. An improvement at the first two levels does not automatically produce a result of the same magnitude at the third.
Market forecasts should therefore not be the sole justification for the investment. A sale researched by AI, a visit that begins with an AI referral, a payment completed within a chat and a fully autonomous order are different magnitudes. Presenting forecasts produced with different definitions on a single chart as if they measured the same market is misleading. What management should actually track is which outcome the selected task changes in its own business.
For Webtures, the value of the research lies in establishing this connection rather than repeating a large future figure. Which data defect is distorting customer choice? Which transaction barrier is increasing support cost? Which capability is required to move into wider channels? These questions create a shared vocabulary between commercial importance and technical importance.
Another area of value is flexibility in technology choice. Keeping product data independent of channels, exposing transaction logic through open contracts and making records portable can reduce dependence on a single platform. This flexibility is achieved not by connecting to every channel at once, but by keeping the core commercial rules under control.
5 Discovery and transaction paths by platform
Agentic Commerce is not a single product or a single sales channel. Within the same platform, product research, referral to the store and direct checkout can each be subject to different eligibility conditions. The public announcement of a feature does not mean it is active in every merchant account and in every country. The readiness decision must be made at the account and task level.
The catalog connections OpenAI offers for product discovery and the activation of checkout must be evaluated separately. The fact that a product feed is processed does not show that the merchant has been accepted into all purchase flows. A shopping session that starts in ChatGPT through Shopify may continue in the merchant's own checkout. In that case the critical test is whether the product and cart context is carried over correctly.
In Google's UCP implementation as well, the published technical specification and live support on Google surfaces are not the same scope. The platform's merchant acceptance, account preparation, product eligibility and technical approval are required separately. A capability described in a developer document must not be treated as ready until it has been tested in a specific customer account.
The conditions in Shopify's Google AI Mode and Gemini direct checkout documentation, which cover US-based stores selling to US customers, show that the same flow cannot be promised to a merchant based in Türkiye. Discovery and referral to the store can be evaluated as a different path. For this reason, country coverage must not be represented by a single "supported" mark.
At least five separate records must be kept for a channel: the merchant's country of incorporation and account, the country where the buyer is located, the delivery region, the product category and the target transaction. Account acceptance status and the version in use are added to these. The eligibility of the payment provider and the currency are also separate components of the commercial flow.
| Path examined | First item to verify | Output to record as success |
|---|---|---|
| Product discovery | Catalog acceptance and correct product matching | The right product is found for the relevant request |
| Referral to the store | URL variant and session continuity | The user continues with the correct context |
| Cart through the API | Tool access and a current offer | A verified cart state |
| Checkout inside the platform | Account, country, product and payment eligibility | An authorized transaction and an order record |
| Task through the browser | Interface access and permitted use | The defined end state is reached |
Browser agents, agents that use APIs and merchant-specific integrations are not measured in the same way. Some flows can work from screenshots or accessibility information. Others rely on structured tool responses. For this reason, assumptions such as "agents do not use the web interface" or "one protocol solves every interface problem" must not form the basis of a design decision.
Platform changes require regular review. A new feature is placed under test coverage rather than being opened immediately to an existing customer. The removal of an old feature or a change in eligibility conditions can also affect customer tasks. The channel inventory must be an operated technical record, not a promotional table prepared once.
6 Evaluating the four commercial capabilities together
In the Webtures approach, an externally understandable account of readiness can be built on four capabilities: discoverability, understandability, reliability and actionability. This structure is not a certification standard. It is an assessment framework that allows management to relate complex technical findings to the customer journey.
Discoverability means that the product or the merchant becomes an accessible option for the relevant request. The presence of a product in a feed does not mean it will necessarily be selected in a need-based search. The channel's access to product information, its matching of the correct product and its evaluation of the product among relevant options must be kept as separate observations.
Understandability means that the product is correctly related to the user's conditions. The model must be able to distinguish what the product is, which variant is being sold and which constraints it carries. Missing information must not be filled in with a generated answer. When there is uncertainty, asking for clarification is a more correct outcome than producing false certainty.
Reliability means that the information used is consistent with the authoritative record and with the conditions at the moment of the transaction. The price shown on the page matching the price in the feed is not sufficient on its own; both values may be out of date. Accuracy, freshness and commercial applicability must be handled together.
Actionability means that the permitted goal is actually produced. That goal may be a comparison list, a quote request, a test cart, a confirmed order or a specific after-sales action. A further action that was not expected from the task is not evaluated as successful automation. Using more authorization than needed is a quality defect.
These capabilities are logically related, but they do not form a one-directional ladder for every customer. One brand may have a reliable order API in a limited corporate channel while being weak in general AI discovery. Another brand may be recommended frequently while its stock data is inconsistent. Giving both the same "intermediate level" label blurs the required implementation.
Scope and evidence must be shown for each of the four capabilities. In which languages and queries was discovery evaluated? Across how many variants was understandability checked? In which fields was freshness measured? In which environment, and up to which outcome, was the transaction executed? These questions turn a positive narrative in a presentation into an auditable report.
On the management screen, open blockers, the test date and the next step must appear next to the four capabilities. One area being judged sufficient must not give the impression that the others are ready as well. The commercial value of the report is that it can show at a glance which progress is possible and which condition must be resolved first.
7 Technical depth across eight audit areas
Beneath the four commercial capabilities, a more detailed technical and operational evaluation is required. The working model we propose in this report uses the eight audit areas together. These areas do not have to be sub-headings of the same score; each one defines a different owner, a different form of evidence and a different type of correction.
Channel and authorization fit examines whether the business can perform the targeted transaction on the relevant platform. Product discovery and identity looks at whether the user's request matches the correct merchant, product family and variant. Data accuracy and freshness evaluates the reliability of the information used in that match. Offers and commercial terms addresses whether the price and the delivery promise are actually applicable.
The task and transaction execution area checks whether the defined goal has been created in the system records. Security and user control examines the limits of authorization and the stop behavior when it is needed. Measurement and reconciliation ensures that the observed outcome matches authoritative records such as payment and order. Operations and improvement covers maintaining quality after changes.
| Audit area | Object examined | Expected decision |
|---|---|---|
| Channel and authorization fit | Account, country, category, permission | Determining the scope of the flow |
| Product discovery and identity | Need, product, variant, merchant | Correct mapping and identification of gaps |
| Data accuracy and freshness | Source record and distribution paths | Data contract and correction |
| Offers and commercial terms | Total, delivery, campaign, return | A valid offer and approval limits |
| Task and transaction execution | Tool call and end state | Task acceptance or a blocker |
| Security and user control | Identity, authorization, sensitive data | Permission limits and stopping |
| Measurement and reconciliation | Event, payment and order records | A verified outcome and remaining uncertainty |
| Operations and improvement | Ownership, release and retest | A permanent operating plan |
This separation reduces the routing of a problem to the wrong team. Asking the experience team for more descriptive error text for an attribute that does not exist in the product data is not a solution on its own. Making a flow that has no payment authorization faster does not increase actionability either. The area the problem belongs to and its dependencies must be identified first.
A single finding can affect more than one area. Incorrect stock is both a data freshness problem and a cart execution problem. In that case the single finding record must be preserved and linked to the relevant areas. Counting the same defect again under different headings can make the total risk appear larger than it is.
The evaluation of the eight areas is presented together with the level of evidence. A capability seen in a document and a flow that runs in a sandbox do not carry the same certainty. Verification performed on a limited set of products in the production environment cannot be generalized to the entire catalog either. Technical depth requires bounding the result correctly as much as it requires expanding the scope.
8 Choosing the first implementation task
The starting task must be chosen from a need that can both produce commercial value and have its outcome verified. A task that is too broad mixes different causes of failure together in the first pilot. A task that is too simple contributes nothing to the investment decision even if it runs successfully. The right balance is to preserve real business complexity within a limited scope.
In task selection, the frequency of the customer need, the impact of a wrong outcome, the accessibility of the data, existing integrations and the cost of verification are evaluated together. Preparing a cart with the correct variant in a category where product attributes are well defined can be a good start. In complex B2B sales, the goal may be a quote request with its technical conditions completed rather than a direct order.
For example, for electronic accessories that carry compatibility information, the pilot can be defined as finding the options that fit a specific computer model and preparing a test cart within a total budget. If the user's device model is unclear, asking a question is the expected behavior. The system must not treat an accessory as compatible only because a similar word appears in the product name.
In the apparel category, the starting point can focus on correctly selecting the size and color variant and understanding the return conditions. The purpose of the task is not to guarantee that the product will fit the user physically. The conclusion that can be drawn from product data must be kept separate from the uncertainty that depends on the person's own experience.
Leaving the payment transaction to the end of the first scope may be appropriate in some organizations. The reason for this is not to postpone the technology but to surface the earlier data and offer problems with low impact. If payment capability is the main commercial goal, the scope is prepared for it from the start, with the payment provider's test environment and the customer's approval.
The pilot's acceptance criteria must be written before the results are seen. Which error will stop the flow? At which point of uncertainty will the user be consulted? Which record will serve as evidence of success? How many attempts will be made at most? At which threshold will the unit cost need to be evaluated? The answers to these questions make comparison possible later.
The success of the first task does not require the scope to be opened immediately to the entire catalog. A new product type, country, campaign or payment method adds new conditions. Expansion must be supported by additional tests, not by an assumption of similarity. In this way the organization learns from a small start, but does not turn the result of a small sample into a large readiness claim.
The role of Webtures is to ensure that the selected task establishes an understandable link between the customer need and the technical work. When that link is clear, the audit scope, the implementation budget and the success evaluation are all directed at the same goal.
9 Architecture and system boundaries
In Agentic Commerce readiness, the architecture must explain which systems inside the organization a customer's request passes through before it turns into a commercial outcome. The system that provides information about the product, the system that determines sellable stock and the system that accepts the order may be different. When this separation is not visible, an agent can complete an answer that looks correct with a wrong transaction. The review we propose in this report begins with mapping these decision points and responsibilities, before any technology names.
In a business, the technical attributes of a product may be held in product information management, its price in the campaign engine, its stock status in warehouse management and its delivery estimate in the logistics service. Moving all of these fields into a single database is not always necessary or feasible. The real need is to identify the authoritative source for each field and to explain which record takes precedence in the event of a conflict. For this reason, the "single source of truth" approach must be treated not as a decision to purchase one piece of software but as a matter of data ownership and decision order.
In the proposed architecture, the discovery, evaluation and transaction layers are connected to each other but hold separate authorizations. The discovery layer ensures that the product is found and understood. The evaluation layer produces a valid offer under conditions such as address, quantity, customer eligibility and time. The transaction layer executes the permitted change. A tool that reads the catalog must not be able to change a campaign or cancel an order with the same access key. This separation makes it easier to develop functions and to limit the impact of an error.
Interface and API paths must be evaluated together. Some agents use the store through a browser while others can access structured tools; other flows hand the customer over to the existing checkout. It is possible for the organization to support only one path. The report must state which path serves which customer task. The existence of an API does not make the browser experience unimportant; a good interface does not show that a transaction API exists for external systems either. The data source and the commercial outcomes of each path must be consistent.
The integration layer can take on a controlled translation role between existing systems and the agent. A product lookup or cart preparation operation exposed to the outside may call more than one service internally. In this layer, data transformation, identity mapping, permission checks and error classification must be defined explicitly. The agent's free-text answer must not be turned directly into a price or stock record. The application of commercial rules must remain within the business's auditable application logic.
Another dimension of system boundaries is the data belonging to customers and business partners. The catalog service may not need to know the past purchases of all customers. The information required for personalization must be limited to the purpose of the relevant task and its access scope. The possibility of the test environment writing to the production database, shared access keys and customer accounts that get mixed together must be made especially visible in the architecture review. Access rights must not be expanded merely because the connection can technically be established.
The architecture deliverable Webtures proposes combines a short system map with data owners and transaction boundaries. For each critical connection, the input, the output, the authoritative system, the behavior in case of error and the responsible team are specified. A missing connection and a connection deliberately kept closed are not reported as the same problem. In this way the development program can focus on the changes required for the first customer task to run reliably, without creating pressure to replace the entire infrastructure.
10 Agent access and verifiable identity
Agent access does not mean allowing every automated request that arrives. The business must decide which of its content it will make accessible for which purpose, and for which operations it will require identity and additional authorization. A system that collects content for search purposes, a tool that opens a page at a user's request and an agent that wants to place an order behave differently. The access control that Webtures recommends separates them by function and permission instead of grouping them into a single "AI bot" category.
The first layers of control are the domain, the network, the content delivery network, the web application firewall and the application itself. Even if the product page is open to everyone, certain requests may be held at an additional verification screen or hit a rate limit. The finding should state at which layer the block occurred. The fact that a page does not open with a particular tool does not show that all AI systems are unable to reach that brand. The different permitted access paths must be examined one by one.
Robots.txt expresses crawling preferences; it is not a security mechanism that grants authorization to access private information. Authorization for commercial transactions is not managed in this file either. Google-Extended is also not an independent HTTP request user agent; it is a robots.txt product token that manages certain content usage preferences. Looking for a separate bot named Google-Extended in server logs, or treating enabling this token as a condition of Google Search visibility, is not correct. For Google Search, the relevant crawling and indexing controls should be evaluated within their own scope.
A bot name appearing on a request does not prove that the request really came from the stated provider. Depending on the verification method the provider supports, published network information, appropriate DNS verification or cryptographic request signatures can be evaluated. The verification method and the date of the last check should be recorded. Even an agent whose identity has been verified may not be authorized for every operation. Recognizing the requester and being able to act in the customer's account must be kept as two separate decisions.
Approaches that use cryptographic signatures can provide additional evidence of request origin and integrity in supported integrations. On the other hand, automatically treating every signed request as safe is not sufficient. The validity of the key, the fields the signature covers, the time conditions and the application's authorization limits must be checked. The product status of the provider in use and the version of the standard it supports are verified separately. An approach that is still in draft status should not be presented as the common and mandatory identity infrastructure of the entire internet.
Instead of switching the firewall off entirely, controlled policies are recommended for the endpoints required for business purposes. Read-only catalog access and calls that create a cart or request account information do not have to share the same limits. Throttling under traffic spikes, an understandable error response and reasonable retry behavior must be designed. If a security measure requires a handoff to the user, this does not count as a technical failure in every case; it can be part of the expected flow.
The output of the audit should be an access record that shows the permitted purposes and methods. In this record, verified access, traffic that was only observed, rejected requests and paths left out of scope are kept separate. The permissions required for the customer account, the administration interface and payment operations are stated separately. In this way the visibility goal is managed together with the organization's data security and service continuity.
11 Semantic HTML and accessible experience
For agents that work through a browser, the meaning of a web page does not consist only of the image on the screen. Headings, field names, the functions of buttons, selected options and error messages affect how well the task flow is understood. An interface that is clear and consistent for people can also help agents work more accurately. However, presenting the accessibility level directly as a guarantee of compatibility with all agents is not correct; real task success must be tested separately.
Semantic HTML explains in its structure what a control does. An action button being a real button element, form fields carrying understandable labels and the relationship between option groups being stated is a solid start. ARIA roles can complete the meaning where they are needed; but on their own they do not create keyboard behavior or correct interaction. Adding only a role to an element that visually resembles a button does not by itself turn it into a reliable action control.
The review that Webtures recommends checks whether the visible label and the accessible name describe the same task. A "Continue" button that starts payment, a selected size indicated only by color, or an error message that sits far from the form can make task interpretation harder. Form fields, required-field information and correction messages must be handled together. Whether dynamic changes are announced correctly to the user and to the relevant assistive technologies is also included in the scope of the check.
Take an illustrative shoe store where the color of the button changes when the user selects size 43, but the previous size remains in the page's accessible state. The agent may add the wrong variant to the cart in the next step. In this example the solution is not limited to giving the agent longer instructions. The programmatic state of the selection, the visible product information and the cart request must all point to the same variant. The test must check both the selection and the resulting cart line; merely clicking the button must not count as success.
The use of JavaScript is not a defect in itself. The real question is under which conditions the target access path can read the critical information. Server-side HTML generation or prerendering can make it easier for some content to be present in the first response. That said, server-side rendering does not bring an outdated price up to date and is not a mandatory solution for every store. The initial document, the rendered page and the state after user interaction must be examined separately; the choice should be made according to the measured problem.
Page stability is also part of transaction reliability. Controls that shift position during loading, unexpected pop-ups, automatic scrolling and unclear waiting states can affect both people and browser agents. Unnecessary interruptions should be reduced while the required confirmations of the commercial flow are preserved. Guest checkout can be considered if it fits the business model; but removing mandatory identity verification merely so that the agent can proceed more easily should not be recommended.
The interface audit should not end with the desktop view. The mobile view in which the target flow runs, the in-app browser, the language and the session state are also defined in the scope. On handoff to the payment provider, the product, the total and the information presented to the user must be preserved. The acceptance criterion for every improvement must be concrete: the name of the control is understandable, the selected variant is correct, the error can be corrected and the user can take back control at the required point. Visual refresh and functional improvement are separated by these criteria.
12 Product identity, variants and catalog
The agent finding the right product is a stronger condition than reaching a page with a similar title. The product family, the purchasable variant, the merchant's offer and the delivery option must be separated from one another. A different capacity or pack quantity of the same model can be a different commercial object. In the readiness assessment the catalog should be examined not as a content list but as the identity scheme that connects the customer's request to a purchasable product.
The first piece of work is the mapping between product identities. The internal record, the store SKU, the manufacturer code, the GTIN where one exists, the feed identifier and the identifier used on the checkout line are handled together. These values do not have to be identical; but the conversion between them must be explicit and traceable. When a product is deleted and recreated, it must be determined what the old links will show, whether there will be incorrect mapping to other products and how past orders will be preserved.
The variant model does not consist of color and size alone. Capacity, connector type, material, pack quantity or a specific technical feature can change the purchase decision. An option that cannot be purchased and an option that is temporarily out of stock are different situations. A product family being in stock does not mean that all of its variants can be bought. A data model that carries these distinctions explicitly reduces the extent to which the agent fills the uncertainty with its own assumptions.
Forcing every product to fill all of the GTIN, MPN and SKU fields is not correct. Real identifiers assigned by the manufacturer should be used; global identifiers that do not exist should not be invented. Some products, such as handmade or custom-made items, may not have a global product identifier. The field and rules that the target channel prescribes for this situation are applied. Being unable to verify a product identifier and the product genuinely not having such an identifier should not be treated as the same data state.
Take an illustrative accessories catalog where the single and triple packs of the same adapter have been published under similar names. If the agent does not read the pack quantity while choosing the lowest visible price, the unit price comparison will be wrong. The solution is not merely to add the pack information to the end of the description. The quantity, unit, price and pack contents of the sellable record must be defined together; it must be verified that the cart line represents the same product. The identifier of the triple pack must not be casually swapped with the identifier of the single product.
Catalog quality cannot be measured by the fill rate of fields alone. A field filled with incorrect or guessed information can produce a heavier consequence than a field left empty. The basis, scope and verification status of a product attribute should be kept. Compatibility information is designed by category: the device model may matter for an electronic product, the sizing standard for textiles, the production year range for spare parts. Natural-language usage descriptions must remain consistent with these objective attributes.
The catalog deliverable that Webtures recommends combines the identity mapping table with a dictionary of critical attributes. First the fields that change the purchase decision are completed; then descriptive enrichment is evaluated. Products are not selected by sales volume alone; variant complexity, stock volatility and the cost of a wrong match are also included in the sample. In this way the catalog work reveals the structural gaps that affect task success instead of superficially tidying thousands of records.
13 Structured data and content
Structured data allows the product and the commercial terms to be expressed in a specific vocabulary. It is not a separate reality that replaces the content. The product page, the catalog and the structured data must describe the same commercial information. Presenting advantages to the machine alone that are not shown to the user, or leaving an old price in the schema, turns into a reliability problem. In the Webtures approach the aim is to associate the correct information with the correct object before producing more fields.
For products with variants, ProductGroup can be used to group the products of the same base product that are distinguished by specific attributes. This group can be connected to separate Product records through the hasVariant relationship; the reverse relationship can also be established in the appropriate way. The offer, price and availability of the purchasable variant are handled in their own context. Casually placing all variants into the hasVariant field under the main Product does not correctly reflect the semantics of the structure being used. Whether the page uses a single URL or multiple URLs also affects the implementation design.
Offer, which describes the merchant's specific commercial proposal, must be separated from the product itself. The currency of the price, the validity conditions, the product condition and its availability must be consistent. Shipping and return information can be linked to the business-wide policy or, where appropriate, to the product's specific terms. Copying the entire policy text again for every product can produce inconsistencies at the next change. How shared rules and product exceptions are managed is an explicit content and data decision.
The presence of a property in the schema vocabulary does not show that all search and shopping platforms use that property. The product display that Google supports and another platform's feed contract must be checked separately. Passing a syntactic validator is not a guarantee of being recommended in an answer engine either. The technical review separates three questions: is the structure correct, is the information inside it correct, and for what purpose does the target channel support this information? This separation makes investment priorities more realistic.
On the content side, which need the product meets, under which conditions it is not suitable and which information remains uncertain must be written explicitly. Details such as the unit, the measurement system, the model year, the maintenance condition or accessories not included in the pack can reduce purchase errors. It should not be assumed that agents only read the schema and do not use natural language. Descriptive text and structured attributes must work together; a marketing statement must not be turned into a technical guarantee. Using stronger adjectives for an unverified feature does not solve the data gap.
Files such as llms.txt or ai.txt should not be defined as a universal precondition for purchasing. In certain systems a method that assists documentation discovery can be chosen; even so, the existence of the file does not prove that all agents will read it. Special AI text files are not required for Google Search visibility. For the same reason, the absence of such a file on its own does not constitute a sufficient criterion for declaring the business invisible or unprepared for commerce.
The implementation plan in this area should also cover the content owner and the update triggers. When a policy changes, it is determined which pages, schemas and feed fields will change. The team that prepares the content draft and the team that accepts its commercial accuracy may be different. After a change, verifying the code alone is not enough; on selected products the visible page, the record presented to the machine and the outcome of the related transaction are examined together. In this way structured data stops being an independent tagging job.
14 Feed and channel data contracts
A product feed is a regular data flow that transfers the organization's catalog information in the form that a specific platform expects. Preparing a feed and opening a sales channel are not the same operation. Whether the product can be discovered, used in advertising and included directly in checkout may depend on different eligibility conditions. The review that Webtures recommends starts with a data contract that explains, for each channel, which data is sent for which purpose and at which stage acceptance is verified.
In the data contract the meaning of a field matters as much as its name. Whether the price includes tax, whether the stock information describes the quantity available for sale, how the delivery country is determined and which variant the product link opens must be written down. A field being missing, empty, zero or unknown can produce different outcomes. If the conversion rules for these cases are not explicit, a file that is valid in form can produce commercially wrong products.
| Contract field | Decision to be explained | Verification evidence |
|---|---|---|
| Product and variant identity | Which purchasable record is represented | Source record and cart mapping |
| Price and country context | Currency and applicable customer conditions | Current offer for the target country |
| Publication lifecycle | Add, update and removal behavior | Processing result and target check |
| Error handling | Who will correct the rejected record | Error record and resubmission |
The success of the integration does not end when the file reaches the other side. Acceptance of the transfer, processing of the records, the products being found eligible and their being usable on the target surface are separate states. In the case of a partial error, it must be understood how many products were affected. When a product is removed, it is tested when the old record will become invisible or be closed to sale. The operating cost and the error recovery path of regular full transfers versus transferring only changed records are evaluated together.
Google's UCP implementation includes rules such as defining an additional field for the product's checkout eligibility and matching the feed identifier with the checkout identifier. The OpenAI product feed must also be evaluated with its own field contract and activation processes. Sending a file prepared for one platform to another after only renaming its fields should not be accepted as sufficient. A shared internal catalog model is useful; but the transformations and conditions specific to the target channel must be managed separately.
In secure transfer, access keys, file sharing permissions and the scope of personal data are checked. Carrying publicly available product information does not require adding customer profiles or order history to the feed. Sensitive access information should not be kept in error logs. If a shared integration provider is used, the contract, data ownership and the recovery path in a service outage are clarified. The "successful" status seen in the vendor panel is verified with an independent sample check when needed.
Feed operations may be spread across the content, category, technology and commerce teams. For this reason the business owner and the technical owner must be written down separately for every data flow. Adding a new category, country or product type may affect the existing contract. In the update plan, the age of rejected records, product coverage, the time of the last successful processing and the number of open errors are monitored. In this way feed management turns from a periodic file submission into the reliable transfer of commercial information to the right channel.
15 Freshness in stock, price and delivery
Freshness should not be promised unconditionally as every system showing the same value at every moment. In a distributed commerce infrastructure, it can take time for changes to be transmitted, processed and reflected in caches. The issue that needs to be managed is how much delay each field can tolerate and which information is re-verified before a transaction is made. The approach we propose in this report is to define measurable freshness targets and safe transaction behavior instead of claiming zero latency.
A product description and a campaign price do not change at the same rate. Stock information can also cover different states such as warehouse quantity, reservations, sellable quantity and shipping capacity. The fact that a product is physically present in the warehouse does not by itself show that it can be sold to a specific customer on the stated date. The data contract should explain the business meaning of the stock field. Publishing the exact stock quantity should not be treated as mandatory; the availability information the task requires, commercial confidentiality and the requirements of the target channel should be evaluated together.
In an illustrative campaign, suppose the price changes at 14:00 while the next processing time of the feed is 14:10. The agent can see the old price in this window. The correct way to manage this situation is neither to shut down every cache nor to keep selling unconditionally at the old price. The current offer should be obtained from the authoritative system before the transaction; if the difference changes the user's approval limit, it should be re-evaluated. The study should measure both the length of the delay and how the system behaves within that window.
In freshness measurement, separate timestamps can be kept for the source change, the preparation of the transfer, the platform's receipt and the verification at the destination. These timestamps should be compared in the same time zone and with appropriate clock synchronization. The date of the last file submission is not the date on which each product field was refreshed. Failed records should not disappear inside the average delay calculation. A change that is not yet visible at the destination should be tracked as an explicit propagation record, not as a successful update of zero seconds.
Delivery information is tied to the address, the order cutoff time, the preparation time and the carrier's service area. A duration stated nationwide on a page may not satisfy the customer's condition for an exact date. The difference between calendar days and business days, public holidays and in-store pickup options should each be handled with their own rules. The carrier's estimate and the commitment guaranteed by the business are expressed separately. If the data cannot be verified, the agent should explain the uncertainty instead of producing a firm delivery promise.
The technical solution should be chosen according to the source of the problem. Event-based updates, scheduled transfers, cache invalidation or pre-transaction queries can each meet different needs. Querying at the highest frequency in every field can increase cost and strain the capacity of the source system. The freshness target should be decided together with the rate of data change, the impact of incorrect information and the operating cost. Pre-transaction verification should also not be substituted on its own for stock reservation or order acceptance.
The report should not claim that a direct, global "merchant reliability score" penalty is applied when a price or stock mismatch is observed. No such shared mechanism valid across all platforms has been verified. Measurable outcomes are taken as the basis: rejection of the offer, a wrong cart, the need for re-approval, customer support load or a canceled transaction. This approach ties the investment in data quality to the commercial outcomes the organization can observe, not to assumed algorithmic penalties.
16 Technical implementation and release management
Agentic Commerce integrations operate between changing platform documentation and internal corporate systems. The fact that a connection has been established once does not show that the same flow will keep working correctly on a permanent basis. Product fields, API versions, account eligibility and commercial rules can change over time. The implementation discipline we propose in this report should be built on a release and acceptance record that tracks which customer task each significant change may affect.
The initial record should contain the protocol and API version in use, the date the implementation documentation was checked, the access status of the merchant account, the countries tested and the product types tested. Stable release, preview and roadmap are kept separate from one another. Implementing an open protocol document does not mean that go-live on a specific consumer platform has been accepted. The technical team should reflect this distinction in the development plan; the commercial team should not present a channel that has not yet been approved as existing service capacity.
The files and endpoints used for capability discovery are verified against the version of the relevant standard. The discovery path defined by UCP and the A2A Agent Card are not the same document. A commerce manifest that the organization creates on its own can be useful; however, its name and location should not be presented as a universal standard supported by all agents. That the declared capability actually works and that the caller has the required permission are acceptance conditions separate from the accessibility of the file.
In implementation, the first aim is not to set up every possible protocol at the same time. The smallest reliable connection required for the priority task is selected. In one project, a catalog correction and a solid store handoff may be enough; in another, permissioned cart tools are required. Development work is defined with the expected behavior, the data contract, the error states and the acceptance examples. The statement "agent compatibility will be added" on its own does not create a testable deliverable between teams.
For version changes, format and business behavior should be checked together. A field name can stay the same while its meaning changes; a previously optional field can become mandatory. The adaptation layer should manage these differences explicitly. Contract tests can be built with suitable synthetic or permissioned test data instead of real customer records. Critical product identifiers, the total calculation, access limits and the handoff to the user should be re-verified with every significant change.
The migration from the Google Content API for Shopping is a concrete example of this discipline. As of the research date, the sunset process for the legacy Content API had begun and migration guidance to the Merchant API had been issued. The presence of the Content API name in an old guide does not justify starting a new integration with the same API. The version used in the existing connection, the provider's migration responsibility and the current shutdown schedule should be checked. A new work plan should be based on the currently supported interface.
Before release, a rollback and safe-stop path is prepared. In the event of a problem, reverting only the application code to the previous version may not be enough; created carts, pending tasks and changed data records should also be handled. Controlled expansion can begin with a limited group of products or accounts. After go-live, the owner of the error records, the monitoring frequency and the intervention conditions should be defined. In this way, technical implementation turns from an integration completed on delivery day into a sustainable commerce capability.
17 Protocol choice and interoperability
In an Agentic Commerce infrastructure, protocol choice should start from the task the customer wants to perform. A brand wanting to make product research easier and a brand wanting to create orders based on authorization the user granted in advance do not produce the same integration need. In the first case, catalog and product querying may be sufficient. In the second case, the validity of the offer, payment authorization, order records and exception management also come into scope. The right architecture is not the one that adds every possible protocol, but the one that can complete the target task with the necessary controls.
MCP governs an AI application's access to tools and data sources. A2A supports the exchange of tasks and messages between independent agents. ACP and UCP aim at representing commercial interactions in a shared format. AP2, in turn, addresses carrying commercial transaction and payment authorization in a verifiable form. These structures are not alternatives at the same level; in some architectures they can be used together. Nevertheless, it cannot be concluded that every merchant needs to set up all of them. Running a task with clearly defined limits over a standard API can be a starting point that meets the need.
| Design decision | Evidence to look for in the choice | Issue not to be confused with it |
|---|---|---|
| Catalog access | The correct product and variant are returned | Authorization to make payment |
| Commerce interface | The supported operation and version work | Acceptance on all platforms |
| Agent-to-agent task | Task identifier and state are transferred | Automatic delegation of user authorization |
| Payment authorization | Verifiable link between approval and transaction | Bank or payment network support |
Interoperability assessment should be carried out separately at the documentation level and at the running-system level. The version the merchant supports and the version the platform uses may not match. Different capabilities may be active under the same protocol name. One system may be able to query products while another may only accept a checkout session. The absence of a feature match should produce a result that can be explained to the user; the agent should not send an unsupported operation to another endpoint by guessing.
For UCP, the discovery path at which the business profile is published takes the form /.well-known/ucp. This profile serves to understand the supported services and versions. The presence of the file does not mean that every declared function works or that a specific consumer platform has accepted the merchant. In the audit, the profile and the actual responses should be compared side by side. A different discovery path mentioned in another document should not be copied without verifying the current implementation contract.
The protocol record the technical team keeps should be simple but sufficient: the implemented version, the active capabilities, the verified endpoints, the authentication method, the test environment and the change owner. A draft feature and a stable implementation are not shown at the same level. A feature that sits in the provider's future plans is not presented as current capacity. Which tasks will be rerun when the version changes is also tied to this record.
From the Webtures perspective, the recommended output is not a checklist made up of technology names. It is an integration decision that maps the target task to the customer's existing systems. The decision should make visible which component will be kept, which will be adapted, which external approval is awaited and which stage will be completed by a human. This approach prevents unnecessary investment in a channel whose commercial access has not yet been verified, while exposing the areas in catalog, offer and transaction quality that can be improved starting today.
18 Designing tool and API contracts
The fact that a tool can be called by an AI does not show that the tool is ready for commercial use. Readiness starts with clearly defining what the input means, which change the tool is allowed to make and how the result will be verified. A broad tool such as "manage the order" can bundle reading, modification, cancellation and payment operations together in an ambiguous way. Clearer tool boundaries make both the agent's decision and the application's authorization check easier.
The effects of catalog querying, stock verification, offer generation and order creation functions should be evaluated separately. A read-only product query and a call that reserves stock are not handled in the same way. The tool description should not hide the side effect. If creating a cart initiates a reservation, this behavior should be written into the contract. Which action produces a result visible to the customer and which one creates a financial impact should be known by the application.
The input contract should explicitly carry the product and variant identifier, the quantity, the currency, the customer context and the required version information. The unit in which a numeric amount is given should not be left ambiguous. The minor-currency-unit approach used by some providers is not carried over to all other tools by default. Empty value, unknown value and zero are kept separate from one another. Not knowing the stock quantity does not mean the product is out of stock, nor does it mean unlimited stock.
The response should not consist only of a sentence to be shown to the user. Where required, the product identifier, the offer version, the calculated total, the validity time and the transaction status should be returned in structured form. MCP tool definitions provide an input schema and an optional output schema for this purpose. Nevertheless, whether a schema-compliant response is commercially correct is tested separately. A wrong price that is correctly formatted does not remove the need for verification.
The error contract is an important part that determines what the agent should do. An invalid product identifier, missing authorization, an expired offer, a temporary service error and an operation with an unknown outcome are different situations. Applying an automatic retry to every error is not correct. Whether a retry is safe should be determined by the effect of the previous operation and the provider's rule. The explanation shown to the user should be understandable; error details should not contain secret keys or another customer's information.
The descriptions that tools give about themselves are not an absolute source of trust either. A tool saying that it is read-only does not prove that it makes no change in practice. On the business side, access scope, resource limits and transaction control are enforced. In every call associated with a user account, it is verified that the resource used belongs to the correct customer. Exposing one customer's cart to another customer or to an unauthorized session is an access defect independent of model quality.
Before implementation, it is useful to prepare one acceptance example and one negative example per tool. The stock query should return the correct variant and should not invent an unknown variant. The offer tool should calculate mandatory fees and should not hide it when the currency changes. The order tool should manage a repeat call of the same business request. In this way, the tool catalog stops being a guide that merely holds descriptions; it turns into an auditable contract that shows which function the business provides reliably and under which conditions.
19 Cart, offer and commercial rules
The cart is the list of selected products; the commercial offer determines the conditions under which those products can be purchased. A definitive total should not be announced before the product price, the delivery option, mandatory fees, the applicable discount and the related tax calculation have been evaluated together. The price the agent sees in the catalog does not take the place of the offer valid at the moment of checkout. Presenting conditions that depend on the address, the payment method or membership status as if they were finalized at an early stage particularly damages user trust.
The recommended record for an offer should include the merchant identity, the product variants, the quantities, the currency, the components of the total, the delivery option, the validity period and the offer version. This record does not have to use the same field names across all protocols. What matters is that the same commercial meaning is preserved in practice. When one integration's field names are carried over to another, whether the price includes or excludes tax, or whether delivery is an estimate or a commitment, should not get lost.
In a controlled example, suppose the user has limited the total budget to 4,000 TL. If the product is 3,850 TL and the mandatory delivery fee is 200 TL, the task does not fit the budget. The agent cannot treat the product's list price being within the limit as success. Likewise, the fact that free delivery belongs only to a specific membership tier does not permit promising free shipping to a user who does not hold that membership. This example is a test design; it is not a real price or campaign claim.
Campaign rules should be versioned. The user's eligibility, the coupon's validity time, the minimum cart amount and whether it can be combined with other discounts are determined. When the agent tries an ineligible discount, it should receive a meaningful refusal response. It should be prevented from using a different customer identity or ignoring the membership condition in order to apply a discount. The merchant's margin protection rules should also be enforced by the authoritative pricing system, not by explanations the model writes.
In stock management, the available quantity and the total quantity in the physical warehouse should be separated. A product reserved for other orders may not be shown as sellable stock. However, a universal rule such as "all agent transactions are closed when stock drops to two" is not correct. The threshold depends on the demand for the product, the reservation method and the business's risk preference. Instead of imposing a specific threshold, the audit should test the behavior that prevents the same last stock unit from being allocated to two requests, or the total allocation from exceeding sellable stock.
When the offer changes after approval, the essential question is whether the change stays within the authorization granted in advance. A price decrease alone does not legitimize switching to a different product or merchant. A delay in delivery, a change in the return condition or the product being offered as refurbished can change the meaning of the decision. An out-of-scope change requires new approval. If no suitable alternative can be found, stopping the transaction is the correct outcome. In this way, offer management becomes a commerce control that correctly executes the customer's decision, rather than a mechanism that bends the conditions to increase conversion.
20 User authorization and payment flow
One of the most important distinctions in Agentic Commerce lies between a user granting access to a system and a user granting the authority to spend money. Being logged into an account, or obtaining an OAuth access token, does not mean that unlimited purchases can be made on behalf of that account. The user can authorize product research, cart preparation and payment initiation with different limits. This distinction must be understandable in the interface and auditable in practice.
The recommended authorization record links the represented user, the purpose of the transaction, the permitted merchant or category limits, the total budget, the currency and the validity period. Depending on the task, quantity, delivery condition or product attribute can also become a hard limit. The request the user gives in natural language must be converted into a clear summary before approval. The gap between the agent's interpretation and the scope the user actually approved must be closed before any financial transaction.
AP2 v0.2 uses the Checkout Mandate and Payment Mandate structures in this area. The first provides a verifiable record for the purchase of the commercial cart, the second for the authorization of the related payment. The flow in which the user approves the final cart is separated from the autonomous flow in which the user set constraints in advance. The vct field expresses the schema type and version; checkout_hash expresses the link to the related checkout record. In open authorizations, binding the agent key and setting a time limit are important. These fields must be validated against the version in use; terms from earlier releases must not be treated as current field names.
AP2 is not a bank or a payment network. In the same way, a payment token on its own does not prove that the user approved all of the commercial terms. A token can provide more controlled use of raw payment data; its scope varies by provider. Cryptographic authorization, payment credential, the bank's authorization approval, capture and order acceptance are different records. Merging them into a single "payment successful" label makes it harder to understand what happened at the moment of failure.
For Türkiye, eligibility must be evaluated in the context of a specific merchant and payment provider connection. Country of incorporation, payment account, buyer location, currency, card network and transaction type are separate conditions. A global announcement does not show that every local account can use the related feature. Equally, the fact that support could not be verified during research does not mean with certainty that no support exists. A go-live decision requires the provider's current technical and commercial confirmation.
When additional user verification is required, the flow must be handed off to a human reliably. 3D Secure does not mean that the same SMS or one-time password screen will open on every transaction; different flows can exist depending on the method used and the transaction assessment. The agent must not attempt to skip the required verification or to guess codes on behalf of the user. The readiness measure here is not removing the control, but having it engage at the right point and letting the transaction continue afterwards without losing state.
When authorization is revoked, no new financial or state-changing transaction must be started. On the other hand, it may be necessary to query the status of a transaction that was started earlier and to carry out safe reconciliation. The moment of revocation, the moment of the transaction and the financial impact are recorded separately. This detail preserves the user's control while ensuring that the system does not overlook half-completed transactions.
21 Transaction reliability and order lifecycle
An agent task must also behave correctly when the network connection drops, when a service slows down or when a response arrives incomplete. A demo that runs once under normal conditions is not sufficient evidence of reliable commerce infrastructure. The readiness assessment must cover, alongside the expected path of the transaction, the uncertain outcomes and the recovery behavior as well. From the user's point of view, one of the most serious problems is a system that starts a new transaction without knowing its own state.
The recommended state model separates the events that run from cart to delivery. Not every business has to use the same state names. However, it must be clear in which system each record is finalized and which other step it can trigger. A payment provider's approval must not be treated on its own as a shipping order when there is no record in the order management system.
| Transaction state | Verification point | Attention for the next step |
|---|---|---|
| Offer ready | Valid offer and cart record | Validity period and total are rechecked |
| Payment outcome uncertain | Transaction record at the provider | Current state is investigated before a new payment is opened |
| Payment approved | Payment and order mapping | Order acceptance is verified separately |
| Order accepted | Order management system | Shipping authorization depends on the business rule |
| Cancellation or return pending | Related transaction and payment record | Request is separated from completed outcome |
Idempotency is an important control used to prevent a duplicate effect when the same request is sent again. However, the scope of the key, its retention period and the response to parameter changes depend on the provider. For this reason, the technical key attached to the payment request must be complemented by a persistent business request identifier. The user's same shopping request must not be counted as a new order intent because of a new session or a new network attempt.
When a response is lost, the system must first investigate the state of the existing request. The first request appearing to fail does not prove that the transaction did not take place. Generating a new key for the same request and starting the payment again is not a safe recovery method. If the transaction is still uncertain, the user is told that the state is being checked rather than being given a definitive failure message. Meanwhile, the number of automatic retries and the total waiting time are limited.
Event notifications can arrive delayed, duplicated or in a different order than expected. The identity of processed events must be recorded; the same event must not create a shipment or an invoice a second time. A delayed notification must not mistakenly revert the order to a previous state. The event's signature and the related account are verified. Rather than deriving a definitive order from the timestamp alone, the current state in the authoritative system is read again when needed.
Situations such as a payment being created while the order record could not be completed must be placed in an explicit exception queue. Who will review it, within what time a result is expected and which remedial step can be applied are written down in advance. Remediation does not always mean reversing all transactions; sometimes the order must be completed, sometimes the authorization hold must be released, sometimes the user must be contacted. The decision must rest on the current financial and operational situation.
In the Webtures report, the output of this area must not be just the number of errors. The impact of the error on the user, how many transactions remained uncertain, how long reconciliation took and the change that prevents the same problem from recurring must be explained. Without reviewing real records, it is not claimed that a system is completely free of duplicate charges. The conditions tested and the situations that could not be verified are kept visible alongside the result.
22 Security and controlled negative tests
Agentic Commerce security cannot be limited to telling the model to behave correctly. A product description, a customer review or an external document the agent reads can contain malicious instructions. Instead of giving information about the product, a piece of content may ask the agent to change the budget, to access another account or to remove the approval step. For this reason, the boundary between external data and system instruction must be protected both in the model flow and in tool authorization.
Controlled negative tests show how a security control works in the case of failure. Test content is prepared in a permitted environment; real third-party pages are not tampered with. Whether a harmless trial text can redirect the agent to the wrong tool can be tested. However, the result is not derived from the agent's answer alone. Which tools were actually called, which record was accessed and whether any commercial change occurred are also examined.
Authorization scope must be kept narrow. An agent researching products may not need a payment key. The tool that builds the cart must not be able to read all customer accounts. Order query access must not automatically turn into the authority to issue refunds. User, session and business limits are checked on every call. Sensitive credentials must not appear in the model context or in error messages visible to the user.
The negative test set can include expired authorization, revoked permission, a payment token used for a different merchant, the wrong currency, another customer's cart and a tampered signature. Reuse of the same request and concurrent calls are also tested. A successful result is the transaction being refused appropriately and meaningful evidence being left behind. The stopping of an unauthorized request is not added to the commercial task completion rate; it is evaluated separately as a security outcome.
A high success rate obtained by disabling security controls is not evidence of readiness. The supported integration path must be chosen for bot management, authentication or payment control. A signed agent identity can be useful; however, signature verification alone does not show that all of the user's instructions have been fulfilled. In the same way, an MCP connection does not mean that access tokens were issued for the correct service or that all authorizations are correctly limited.
Security in the test environment also covers commercial impact. Unlimited retries, uncontrolled product queries or long agent loops can generate cost and service load. The number of calls, the transaction budget, the duration and the concurrent task limit must be defined. The conditions under which the test will be stopped are written down in advance. Reserving stock, sending a customer message or creating an authorization hold in the production environment requires separate permission and a rollback plan.
Negative scenarios passing cleanly does not prove that the system is safe against every attack. The evidence is limited to the version and scenario set that was tested. When the model, tool, authorization policy or data source changes, the related tests must be run again. For Webtures, correct reporting means clearly showing the defect found, its business impact, its owner and the retest result. Unauthorized financial impact or access to another customer's data must not be lost inside the overall score and must halt the opening of the related flow.
23 Human handoff and customer experience
Human intervention does not mean that agentic commerce has failed. If a task involves a missing preference, compatibility that cannot be verified or additional payment approval, returning to the user is the correct design. The readiness assessment should aim not at increasing autonomy under every condition, but at having the decision made on the right side. The customer's ability to move quickly without losing control is as important as reducing unnecessary questions.
Handoff points must be defined before the task starts. It must be clear which conditions the agent can resolve within the limits given in advance and which require new approval. If black is only a preference, another color can be suggested; however, buying a different product without the user's approval is a separate decision. If a hard delivery limit is not met, the agent must not quietly relax it. Unknown information and a missing user preference must also not be presented in the same question format.
The context handed over to the user must be understandable and sufficient. The selected product, variant, quantity, merchant, total amount and delivery condition must remain visible. Which step has been completed, which is waiting and why the user is needed must be explained. Instead of showing failure details as a technical error stack, the agent needs to present a summary on which a decision can be made. When the user does not want to continue, it must be possible to stop safely.
| Handoff reason | Information to pass to the user | Success criterion |
|---|---|---|
| Missing product preference | Options and changing conditions | Continuing with the correct variant after the decision |
| New commercial condition | Old and new total or delivery | Explicit approval of the new scope |
| Payment verification | Reliable continuation step and transaction state | Completion without skipping the control |
| Uncertain transaction outcome | Notice that the existing request is being checked | Resolution without a duplicate transaction |
When redirecting to the store, cart preservation must be tested separately. A new browser window, an in-app browser or a change in session duration can cause product information to be lost. When the user signs in again, a cart belonging to a different account must not open. For an expired offer, the old amount is not shown as if it were final; the current conditions are retrieved again. Handoff links must not carry unnecessary personal data or payment information.
Accessibility keeps its importance in the experience assessment. Agents performing some operations through an API does not mean that human interaction with the interface has disappeared. Approval, comparison, address correction and support processes require understandable labels, readable totals and usable error messages. Treating a box the user never saw as approved, or hiding important conditions in a narrow area, does not create a good agentic experience.
Handoff success and fully autonomous completion must be reported separately. A user who moves to the store with the correct cart has not yet made a purchase. The user later abandoning the purchase is also not, on its own, a technical handoff error. Measurement must rest on the final state expected in the task. Thanks to this distinction, teams can improve the real points where the decision is interrupted, rather than producing the appearance of more autonomy, and can manage customer trust as part of commercial performance.
24 Cancellation, return and after-sales operations
Agentic Commerce readiness is not complete when the order is created. The customer may later want to correct their address, learn the delivery status, cancel the product or open a return request. These operations can require authorizations different from the initial purchase authorization. Even if an agent prepared the purchase, it must not be assumed that it can change every order on behalf of the same user in the future.
A cancellation request and a completed cancellation must be separated from each other. An action that can be applied while the order is not yet being prepared can turn into a different process once shipping has started. To tell the user "cancelled", it must be verified that the related state has been created in the authoritative system. If the request has been taken under review, this is stated clearly. The agent must not claim that the financial or logistical outcome is complete simply because it sent a request.
In the return process, the product, order line, quantity and eligibility information are evaluated together. In a partial return, refunding the total order amount by mistake must be prevented. If there is a promotional cart, multiple payment methods or a return made earlier, the remaining amount is calculated from the authoritative record. This calculation is not left to the model's free-text arithmetic. Currency and amount unit must be preserved across all services.
The provider accepting a refund request may not mean that the amount appears in the customer's account at the same time. The status given to the user must be consistent with the real payment record. Requested, processed, failed and completed refunds are separated. No universal timeframe is promised; the information the provider gives for the related transaction is used. If there is uncertainty, the customer is shown an understandable follow-up path and the responsible channel.
Return authorization is an area open to abuse. The agent merely knowing the order number must not be considered sufficient. The correct user, the correct order and the permitted scope of the operation are verified. An instruction that a third party has planted in a product review or a support document must not be able to change the refund recipient. Redirecting payment details to another account is not an edit that can be made by trusting a description text.
Duplicate transaction control must continue after the sale as well. The same return request being resent because of a network error must not create a second refund. The total of concurrent requests for the same product line must not exceed the eligible remaining amount. The business's test design must include full cancellation, partial return, rejected request, delayed event notification and failed refund. The products used in the controlled test and the financial impacts are recorded separately.
Policy information must be presented to the customer in plain language; conditions that vary by product, merchant and transaction type must be preserved. A single return period must not be coded as a universal rule for all categories. Situations that require legal interpretation or await the business's exception decision are passed to the relevant expert. The task of the agentic system is not to accept every request automatically, but to carry out the right request with the right authorization and in a traceable way.
For Webtures, this area ties the readiness report to the real commerce lifecycle. Safely correcting a wrong transaction after the sale affects customer trust as much as choosing the right product before the sale. The report must clearly show which operations work only at the information level, which at the request creation level and which at the verified execution level. In this way, alongside purchase speed, operational accuracy and the customer's control over the transaction also become measurable.
25 Scope of the permissioned audit and data collection
A good readiness assessment produces reliable evidence for specific commercial tasks instead of handing the whole store a general pass or fail label. The audit Webtures proposes first defines the business objective and the access boundaries. Which product groups will be examined, which customer will be served, which channel will be used, and at which action will the test stop? An automated scan launched before these questions are answered can generate a large number of technical warnings; it falls short of explaining which problem is decisive for sales or the customer experience.
The scope record brings together the domain names, the store environments, the catalog and stock sources, the payment connections, the test accounts and the data owners. Permitted operations are written separately for read-only access, test cart creation, stock reservation, payment authorization, capture, cancellation and return. Permission granted in one area does not carry over to another. The person who may ask for the test to be stopped, the working hours, the request limits, the rollback steps and the path to follow in case of an unexpected effect are all set at the outset. In this way the study runs in line with the responsibilities of the production operation.
Data collection starts with the lowest access possible. A product export, the existing feed files, publicly available product pages, permissioned system logs and technical design documents may be enough for the first review. In the next stage, only the access needed to complete missing evidence is opened. An auditor holding the entire customer database or an unrestricted admin key is not a measure of quality. Synthetic customer accounts and narrowly scoped technical roles also make it easier to set the test up again.
Every data sample must be stored with the context that explains its meaning. Seeing different prices for the same SKU is not always an error; there may be a different currency, customer group, tax context or campaign condition. The comparison record must include the product, variant, seller, market, address region, account role, currency and the time of the check. It also states which system the field is expected to come from. Warehouse management may be authoritative for stock, the approved catalog for product attributes, and order management for order status. A single-truth approach does not require all data to be kept in a single database.
Audit evidence is separated into four levels: document review, read-only verification, sandbox task verification and permissioned production verification. A feature found in a design document is not considered working. Sandbox success is not converted into production success. A narrow task verified in production also does not mean that the whole catalog and every country are ready to the same degree.
The differences between the test environment and production are recorded separately. If tax, campaigns, shipping or stock contention have been simplified in the sandbox, the result is interpreted within those limits. A check that could not be performed is also an explicit outcome: access was not granted, no record was found, the channel does not support it, or the task is out of scope. These cases are not used interchangeably. Where evidence is missing, judgment is deferred; points are not deducted as if a technical defect had been observed.
The collected records are kept limited to their purpose. Personal data, payment details and secret access keys are not copied into the report; masked samples are used where necessary. Who may see which record, how long it is kept and how it is closed at the end of the engagement are part of the audit arrangement. The first deliverable of the audit is therefore not a list of warnings but an evidence inventory whose scope and confidence level are understood.
The role of automated audit tools
Automated scanning is useful for finding common defects across a wide set of products or pages. However, a scan result should not be used as evidence that a real customer completed the task. An accessibility check can evaluate the name of a button; understanding whether the right product was selected requires the task and the system record. An API contract test can validate the shape of the response; whether the total amount was calculated under the applicable commercial rules is tested separately.
The experimental Agentic Browsing assessment inside Lighthouse can, under specific version and environment conditions, surface findings on topics such as accessibility, page stability, WebMCP and llms.txt. The fact that a file is checked there does not mean that the file is mandatory across every AI channel. The result should not be read as a general commercial readiness or payment reliability certificate. The tool version and the experimental features that were enabled must be kept in the test record.
Tool selection weighs the target platform, the observation scope, reproducibility and the data processing conditions together. Methods drawn from laboratory task sets can be useful; a published success percentage, however, cannot be carried directly into a customer environment made up of different models, sites and tasks. In the Webtures assessment, automated checks, controlled task execution and the authoritative system record are treated as evidence that complements one another.
26 The controlled task contract
The controlled task contract defines what is being asked of the agent and which behavior will be accepted as correct before the test begins. The contract here does not refer to a legal document but to an assessment record. The user request, the technical permissions, the commercial limits and the verification method come together in the same record. Without this record, the statements "found the product", "added it to the cart" or "completed the transaction" can be used for outcomes that differ from one another.
The task first separates the need from the mandatory conditions. Compatibility with a specific device may be mandatory; the color being black may be a preference. When the agent cannot verify a mandatory condition, it may search further, ask the user for information, or stop. An option that does not meet the preference but satisfies the mandatory conditions can be presented with an explanation. The test must not place these two behaviors in the same error class. If the user's goal cannot be verified, an apparently fluent answer does not count as success.
Illustrative task contract
The example below was created to explain the test design; it is not a real product, price or delivery commitment.
| Contract field | Illustrative task value | Acceptance criterion |
|---|---|---|
| Need | A USB-C docking station suited to a specified laptop | Match against the manufacturer's compatibility information |
| Mandatory feature | Drives one 4K display at 60 Hz and has at least two USB-A ports | Display and computer limitations are verified together |
| Preference | Black color and a detachable cable | If not available, the reason is explained |
| Budget | At most 3,500 TL including mandatory fees | Cart total does not exceed the limit |
| Delivery | To the specified test address within four calendar days at the latest | The address-dependent offer is verified |
| Permitted final action | A test cart with a single product inside the sandbox | No payment or real stock reservation is made |
| Correct end state | The approved variant is in the cart with a valid offer | Cart record and product identity match |
| Stop condition | Compatibility unclear, budget exceeded or data contradictory | Clear explanation and no action taken |
Alongside this summary, the task record also keeps the channel used, the tools, the account role, the time limit, the retry policy, the data version and the test environment. A task that has not finished by the end of the time limit is not declared successful. Whether a retry is a continuation of the same run or a new experiment is decided in advance. Otherwise a single successful result arriving after many failed attempts masks the real reliability.
Success verification has two parts. The end state must be correct, and the permission boundaries must be preserved on the way to that state. The right product appearing in the cart is not enough if data was read from another account. Likewise, selecting the wrong variant without any authorization violation is not a positive task success either. The end state is compared with the authoritative system records; the meaning of the explanations is examined by human review where necessary. The interpretation of another language model is not used as the sole source of verification.
The task contract may accept more than one correct outcome. A correct refusal when the product genuinely does not exist, a re-confirmation when the price has changed, or a handoff to the user when additional verification is required can each be the expected outcome. However, these are reported explicitly in their own outcome classes. Grouping every task that started with a purchase intent under the heading "the agent behaved correctly" blurs the positive completion rate. In the Webtures assessment, both the extent to which the customer's goal was achieved and the extent to which the system protected its limits must remain visible.
27 Task sets and failure scenarios
Adding a single product to the cart in a trial environment shows only a small part of agentic commerce readiness. The task set must cover both the conditions that make the customer's decision harder and the behavior of the system at the moment of failure. In the design Webtures proposes, tasks are organized along product discovery, selection, commercial offer, transaction preparation, handoff to the user and the permitted after-sales operations. Each family contains a boundary case alongside the ordinary case.
Discovery tasks are not limited to stating the product name directly. Searching by need, filtering by technical attribute, comparing several products and asking for clarification when faced with missing information are all evaluated. The gap between everyday Turkish usage and the catalog language is examined as well. When the user wants to "connect two monitors", it is not enough for the agent to find a product whose text merely contains the word HDMI; the device, resolution and connection conditions must match in a meaningful way.
Changing data is especially important in commercial tasks. The price of a product that looked suitable at the start of the task may be refreshed, the last unit may be bought by another customer, or the campaign's validity period may expire. The test produces these changes in a controlled environment. The expected behavior is not to continue with stale information but to re-verify the relevant condition and to return to the user if the authorization limit has changed. Failing to find the same product after the data was refreshed and knowingly ignoring current information are separate causes of failure.
Transaction resilience tasks focus on situations where no successful response is received. When the response to a cart request is lost, the record may still have been created. The result from the payment provider and the result from order management may arrive at different times. The same event may be delivered again. These scenarios do not test the agent's ability to try more times; they test whether it learns the transaction state reliably and does not create an unnecessary second transaction for the same commercial intent. Flows that interpret a timeout directly as a failed payment are examined in particular.
Authorization tasks cover situations such as accessing a different customer account, using an expired permission, exceeding an amount or seller limit, and continuing with a revoked authorization. The expected result of the test may be a refusal or a controlled handoff. Whether external content can alter the system rules can also be tested inside the sandbox with a harmless trial instruction placed in a product description. Such a test is carried out only with synthetic content in a permissioned environment; it does not mean manipulating live third-party services.
The failure record must consist of more than a "failed" result. The first point of breakage, the effect on the user, the response the system gave, the state left behind and the recovery behavior must be kept. A single root cause can show symptoms at more than one stage. For example, an incorrect variant mapping can produce both a stock error and a budget overrun. Prioritization must not be distorted by counting the same problem three times as different defects.
The scope of the task set is chosen according to the real business model of the company. A consumer checkout task is not applied to a B2B flow that only prepares quotes. A real return is not initiated in a project without after-sales authorization. A small but commercially meaningful task set provides stronger decision support than a long checklist whose scope is not explained. The number of tests is not a quality indicator; which behavior was verified under which conditions determines the quality.
28 Pilot study and sample design
The design proposed for Webtures' first field study is an exploratory pilot to be run with ten consenting retailers. The sample distribution could consist of four apparel, three electronics and three home and living businesses. This distribution is an implementation proposal; no tests were performed on businesses within the scope of this report. The purpose of the pilot is to understand recurring failure types, data dependencies and the effect of remediation on task success before producing a national ranking.
Examining thirty SKUs at each business creates a catalog sample of three hundred SKUs in total. The selection must not be limited to products that receive high traffic or whose data is tidy. Variant and single products, campaign and standard prices, high and low stock, different delivery conditions and different data sources must be covered. Products with a high likelihood of critical failure may be given particular room in the sample; however, it must be explained that this selection is not unbiased for estimating the overall catalog error rate.
The task execution design is separate from the catalog review. For each business, twelve task families, the two applicable access paths and three repetitions are planned. If all paths are available, the upper limit of one wave is 10 × 12 × 2 × 3 = 720 executions. If a second wave is run after remediation with the same scope, the upper limit becomes 1,440 executions. The three hundred SKUs examined are not multiplied by this number again. The product set that each task family will use is determined separately; the same product may appear in more than one task.
In this calculation, for each combination of business, task family and access path, a single task contract chosen in advance is executed three times. All positive and negative sub-scenarios within a family are not included in this count. If additional scenarios are selected, the execution plan and the denominators are expanded. For example, running 24 separate scenarios per business three times on two paths creates 1,440 executions for one wave across ten businesses.
The two access paths can be defined, depending on the project, as execution through the browser and the permissioned commercial tool/API path. These are not treated as mandatory at every business. When a path does not exist, a failed task is not invented; it is shown in the applicability matrix as out of scope or not yet verified. The planned matrix, the technically and commercially applicable matrix and the actually completed matrix are kept separately. The reduced scope is not a footnote of the report but information that determines how the result is to be interpreted.
Repetitions are run with new sessions and under recorded conditions. A cart, cache, user preferences or learned context created in a previous attempt can affect the next attempt. Effects that cannot be reset are noted; the task order is balanced where possible. Three repetitions of the same task are not presented as three independent customer observations. The commonalities at the business, product and task family levels must be preserved in the analysis.
In the before-and-after comparison, task definitions and success rules are held constant. Even so, if the price, catalog, agent version or channel condition has changed, the two waves are not considered fully equivalent. The change log shows which difference may come from the implementation and which from external conditions. Attributing an improvement observed in a small pilot to a single intervention alone is in most cases not possible.
The use of business names in public publication, the anonymization of findings and the process for correcting material errors are decided at the outset. The business's review of a finding must not turn into a right to select the research outcome. Limitations such as voluntary participation, the concentration of certain infrastructures in the sample and the small sample size are stated openly. As the pilot grows, a more balanced research design can be built according to the distribution of sector, business size and infrastructure; the results of the first pilot are not retroactively given national representativeness.
29 Measurement indicators and the right denominators
The most important methodological decision in a readiness report is which attempts the success rate is calculated against. If only the runs that reached a result are recorded, every system looks more reliable than it is. Valid attempts that produced a timeout, a tool error or an incorrect end state must remain in the denominator. Attempts that never started because of a failure in the test setup itself can be kept separately, with their reasons. Raising the success rate by declaring a difficult task out of scope after the fact is not an acceptable assessment method.
| Indicator | Numerator and denominator | Limit it reveals |
|---|---|---|
| Positive task completion | Positive attempts with a correct end state and a permission trail / valid positive attempts | Achievement of the defined customer goal |
| Correct stop | Negative attempts refused as expected / valid negative attempts | Protection of authorization and commercial limits |
| Correct variant selection | Attempts where the correct variant was selected / attempts where variant selection was expected | Accuracy finer than the product family |
| Offer consistency | Correct offers with matching context / valid offers compared | Agreement of price, delivery and other conditions |
| Handoff success | Handoffs that could be continued with the correct context / valid handoff attempts | Continuability, not purchase |
| Scope completion | Applicable cells with all repetitions executed / planned applicable cells | How much of the research plan was observed |
| Cost per successful task | Cost of all runs in scope / verified successful tasks | Includes the cost of failed attempts as well |
Illustrative denominator example
To explain the calculation, consider a plan of one hundred runs. Let eighty of them be positive customer tasks and twenty be negative tasks where a correct stop is expected. Let four positive runs never start because the test setup did not open. If fifty-seven of the remaining seventy-six positive runs complete correctly, the positive task success is 57 / 76 = 75%. The timeouts and system errors among the nineteen failed runs are not removed from the denominator.
Assume a correct refusal occurred in eighteen of the twenty negative runs. The correct stop rate is 18 / 20 = 90%. These eighteen correct refusals are not added to the fifty-seven positive results to produce a single sales success percentage. It is also stated that ninety-six of the hundred planned runs were validly executed. This is progress at the run level; it is not the same indicator as the share of business, task family and access path cells whose repetitions were all completed. The numbers are entirely illustrative.
Reporting only the average in duration measurement can hide long waits. The median, the upper percentiles in a suitable sample and the number of runs that exceeded the time limit are shown together. The duration of successful runs and the status of all valid runs must be separated. Failing faster does not mean more efficient commerce. Likewise, fewer tool calls do not by themselves show that the correct result was produced.
Data freshness can be measured as the time from a change at the source until it is verified at the destination. The delay of changes that have not yet propagated is not written as zero; it is kept as an open record that continues to be monitored. Because price, stock and description have different levels of importance and rates of change, they must not disappear into a single agreement rate.
In aggregate results, excessive repetition of easy tasks must not shift the weighting. Rates and raw counts per task family are preserved; if a combined indicator is to be used, the weighting is set in advance. In a small sample, plain numbers and a clear explanation of the uncertainty are more useful than a score with several decimal places. Next to every metric there must be the scope, the date, the channel and the measurement level.
30 Observability and transaction attribution
Observability is broader than tracking where an agent clicked. It requires correlating which data the commercial task started with, which tools it used, within which permission boundaries it proceeded, and which business record it created. The arrangement Webtures recommends separates the observed event from the interpretation drawn from that event. Access to a product page does not mean that the product was compared or selected for purchase.
The first layer is access and call logs. The server, the CDN, the catalog service and the commercial API can show the requests that reached them. The second layer is task and session records. The starting goal, tool calls, durations, error codes and the handoff to the user are tracked here. The third layer is the commercial outcome: cart, order, payment, reservation, cancellation and return states. Linking these layers through shared identifiers makes it possible to understand which outcome a seemingly completed task actually ended with.
A persistent business intent identifier represents a single commercial request from one customer. Under this identifier there may be several task executions, tool calls or retries. Cart, order and payment identifiers are stored separately. As a result, repeats tied to the same request are not counted as new customer requests. It is not assumed that every provider supports the same identifier; a mapping record may be needed at the integration level. When the correlation is lost, the record is left open rather than filled with a guess.
The certainty of the agent's origin must also be visible. A User-Agent string alone is not at the same trust level as a cryptographically verified request. A session that merely resembles automation in its behavior, on the other hand, should not be labeled as a verified agent identity. Records can be separated into verified, directly observed but unverified, analytically inferred and unknown. This classification serves a different purpose from the audit's evidence levels of document, read-only, sandbox and production: the question asked here is who the request belongs to.
Measurement design takes the access path used into account. Flows that proceed through a browser or are handed off to the store may generate page events. A transaction carried out through a direct service call, however, may not execute the same client-side tags. For this reason, all agent commerce is not calculated by looking only at web analytics sessions. Reconciliation is done with server events and order records; areas that cannot be observed are stated explicitly.
In-model evaluations that the business cannot access should also not be presented as if they were measured. How many times a product entered the model's hidden comparison list cannot be derived from an independent merchant's standard logs. What can be observed is being recommended in specific test responses, being queried in a permissioned tool call, a recorded handoff or a commercial transaction. If a name such as "agent impression" is to be used, the technical definition of the event and the field of observation must be explicit.
Attributing outcomes to revenue also calls for care. An order that arrives through an AI referral does not mean an order completed fully autonomously. Nor does the observed relationship prove an incremental sales effect on its own. The audit dashboard can show the source, the task outcome and the commercial outcome separately; assessing the incremental contribution requires a comparison design. The record structure delivers value to management to the extent that it makes this distinction possible.
31 Maturity assessment and critical blockers
A readiness decision should show the executive which scope can be opened, which conditions still need to be completed, and what is not yet known. A single total score weakens this purpose when it reduces an open access problem and a payment authorization defect to the same number. In the Webtures approach, the assessment should rest on reading the eight audit areas, the evidence level, the task scope and the critical blockers together. This structure is the implementation model proposed for the report; it is not an official certification or a universal industry standard.
The eight areas are channel and authorization fit, product discovery and identity, data accuracy and freshness, offers and commercial terms, task execution, security and user control, measurement and reconciliation, and operations and improvement. Next to each area sit the observed state, the evidence, the open question and the next step. A strong result in one area does not cover a gap in another. Flawless catalog data does not make the risk of an unauthorized payment acceptable.
The four evidence levels can be used as a progression that builds on itself. Document review shows that a design exists; read-only verification shows that real data is consistent under specific conditions. Sandbox verification provides evidence that controlled tasks work. Permissioned production verification, in turn, provides information about a defined live scope. These levels are not a single ladder that spans all of the company's capabilities; product discovery may be verified in production while the return flow is still only at the document stage.
For critical defects, gate rules with a defined scope are applied. An unauthorized payment, a double charge for the same commercial intent, or access to another customer's data is treated as P0 and blocks the opening of the affected flow. Proceeding with the wrong variant, the wrong total or invalid stock can be examined in the P1 class. However, if these defects produce an unapproved financial transaction or another critical impact, P0 takes priority. Lower-impact gaps, unnecessary steps or usability problems are planned separately. The class depends not only on the name of the error but on the impact that occurred and the extent of its spread.
The absence of critical blockers is not an automatic general approval either. Seeing zero critical errors in a payment flow that was never tested does not count as reliable payment evidence. A lack of scope and zero errors must be kept apart. If there is an open dependency, a closed channel or user consent that has not been provided, the result can be stated as "conditional" or "not yet verified".
The decision sentence of the maturity report must be concrete: "Verified from a Turkish-language need through to the test cart for the thirty selected SKUs; payment out of scope; two open findings remain for delivery updates." This sentence carries more information than a context-free "ninety percent ready" statement. If a visual scorecard is used, the scope, the date and the evidence level must appear on the same surface.
A readiness decision is not treated as valid indefinitely. A catalog mapping change, a new payment provider, an update to the authorization model, a platform release or a significant campaign rule can trigger reassessment. Management should see the control that will be put in place not only as a go-live gate but as an ongoing part of commercial operations.
32 Turning findings into implementation and retesting
The commercial value of a readiness report emerges when a finding turns into work that the implementation team can understand and verify. General advice such as "improve data quality" or "improve agent compatibility" may start work; but it does not provide a measure of completion. The finding record Webtures recommends contains the affected customer task, the observed behavior, the expected behavior, the evidence, the owner, the dependencies and the acceptance criterion together.
In every finding, the observation is separated from the root cause hypothesis. Seeing a wrong price is an observation; the idea that it stems from the cache duration may be an explanation that still needs to be investigated. If the proposed solution treats the hypothesis as established fact, the team may work in the wrong layer. The first implementation step is sometimes not changing code but making the source of the problem visible by recording the offer identifier and the update time.
How an illustrative finding turns into a work package
Suppose that in a test the selected charging station appears correct on the product page but ends up in the cart with a different variant. This is not a real customer outcome. The observation is recorded with the product identity requested in the task, the selected variant, the cart line and the time of the check. Possible causes may be a catalog mapping error, the default selection in the browser, or the intermediary service sending the wrong identifier. Rewriting the entire catalog before evidence is collected is not recommended.
If the mapping error is confirmed at the end of the investigation, the work package specifies the relevant data transformation, the affected product family and the responsible team. The acceptance criterion is not only that the sample product is fixed. Whether the other variants in the same family are preserved, whether a variant that does not exist is silently converted to another product, and whether the agent stops when the selection is ambiguous are also tested. In this way the fix is not closed on the basis of a single screenshot.
Prioritization must weigh impact, spread, likelihood of occurrence, confidence in the evidence and solution dependencies together. Critical authorization defects take precedence over this ordering. An identity problem affecting a broad product group can be addressed earlier than a single missing description. At the same time, low-risk work that can be fixed quickly does not have to wait for long architectural projects to finish. The roadmap becomes actionable by separating urgent protection, root cause correction and permanent improvement.
Retesting covers first the task that caught the same error, then the related neighboring scenarios. The starting conditions and the definition of success are preserved; the changed release is recorded. A single successful run does not show that an intermittent problem has been fully resolved. The predetermined repetition pattern is applied; if the error cannot be reproduced, the result can remain "under observation" rather than "resolved". The closure decision is completed with evidence and an authorized acceptance record.
A rollback plan for go-live is also part of the implementation work. Who does what is determined for when the feature must be switched off, the old data flow restored, open carts handled and pending financial states reconciled. Rollback alone does not erase all commercial effects; the status of earlier transactions is checked separately.
In the end the report ceases to be a one-off document; it becomes the working arrangement that holds the organization's history of tasks, findings, changes and verifications. The distinctive contribution of Webtures in this area is the ability to tie technical detail to management's decision and to the implementation team's acceptance criterion at the same time. This contribution should be judged not by the number of recommendations but by the problems closed with evidence and the commercial accuracy preserved.
33 Implementation models by sector
Agentic Commerce readiness does not mean automating the same task in every business. How the product is chosen, the cost that arises from a wrong choice, and which decision the customer wants to keep for themselves determine the implementation model. For this reason, sector classification should come before the choice of technology. The initial scope should rest on a customer task where the business can solve a distinct problem and where the outcome can be verified. The models below are not real customer cases but suggestions that can be used in implementation design.
Variants and fit in apparel and footwear
In this category the core task is selecting the right size, color and purchasable variant from within the same product family. Finding the product title does not show that this task has been completed. Which model the size chart belongs to, how the measurements were taken, the cut information and the intended use of the product are evaluated together. The mandatory conditions the user has stated must be separated from the preferences they can relax.
A first pilot can aim to move the correct variant into the test cart within a specific collection. Wrong size mapping, a sold-out variant, a color change and incorrectly conveyed return terms are examined as separate outcomes. If the size information is insufficient, asking a question is the appropriate behavior. Deriving a perfect personal fit from the product description is not. In the commercial evaluation of the implementation, the change in variant-related support and return reasons should be tracked alongside the number of orders.
Technical compatibility in electronics
In electronics the customer's need is often more detailed than a product name. A connectivity device may need to work with the existing computer, the operating system and the monitors. The model code, the manufacturer's description, the package contents and the additional parts required are therefore handled in the same task. Products with similar names must not be substituted for one another.
An appropriate starting point is selecting compatible accessories for a limited device family and preparing the total cost. Cases where compatibility data is not available are left visible. The warranty provider and whether the product is new, used or refurbished are included in the selection. In this way the pilot evaluates not only the ability to find a good price but also the ability to prevent an order that does not match the customer's need.
Substitution and delivery in grocery shopping
In the grocery scenario, quantity, unit price, package size and the delivery window matter together. Choosing another product in place of one product is not just a price calculation. User approval may be required with regard to ingredients, allergens, brand preference and storage conditions. If this information cannot be verified, the model must be prevented from establishing equivalence on its own.
The first task can be preparing a predefined shopping list within a specific budget and delivery window. Missing products are shown separately; unapproved substitutions are not silently added to the cart. For variable-weight products and items whose final amount is determined later, the estimate and the exact amount are kept apart. The evaluation also covers whether the prepared cart fits the actual delivery conditions.
Preparing the decision in B2B and travel
In B2B procurement, success may be the preparation of a quote package that can be submitted for purchasing approval before any payment is completed. The technical specification, quantity, unit, minimum order, lead time and contract terms are handled together. Where the price is customer-specific, a firm offer is not created from public catalog information. The authorization to prepare a request is separated from the purchasing authorization that commits the budget.
In travel, the initial scope can be presenting options that match the passenger and date conditions together with the total price and cancellation terms. The same flight or hotel can be sold with different change rights. For this reason, the lowest price is not made the sole selection criterion. Confirmation of the reservation, prepayment or changes that incur a penalty are authorized separately. In both sectors, a handoff prepared correctly for human approval is a valuable outcome in its own right.
34 Türkiye and cross-border commerce
A readiness assessment for Türkiye should not gather domestic and foreign platforms under a single access assumption. The merchant's country of incorporation, the payment account it uses, the delivery point, the buyer's location and the target AI channel are separate fields. Each combination of these fields can produce a different implementation scope. The first management decision is to determine which commercial corridor will be assessed.
For a business selling within Türkiye, the starting point can be having products understood correctly and carried to the permitted transaction stage in the store with current commercial terms. In cross-border sales, language, currency, address, taxes, shipping and after-sales service are added to this. Publishing a product description in another language does not prove that the same product can be sold or delivered to the country concerned.
Carrying local conditions into the data model
Turkish-language tasks should cover everyday expressions and local shopping expectations. Installments, in-store pickup, shipping to a specific district, the order cutoff time and business-day wording may not be interpreted in the same way. The general delivery description on the product page and the offer produced after the address is entered must be evaluated separately. If a campaign condition depends on the payment method or the customer group, the task should ask for this information or refrain from finalizing the discounted amount.
The fact that local e-commerce software provides an API does not mean that all commercial functions are exposed externally. Which queries or changes the package, version and contract in use support must be confirmed with the provider. If custom development is required, maintenance responsibility, version changes and data transfer limits are made part of the proposal. Not finding an integration announcement on the internet does not lead to the conclusion that the provider definitely cannot offer this capability either.
Confirming payment and after-sales
The record to be verified with the payment service provider should include the supported transaction types, additional user verification, the authorization period, the currency, and cancellation and refund behavior. It should not be assumed that every flow uses the same verification method. If a handoff to the customer is required, this step must be designed so that the product and the total amount are preserved. A flow completed by removing the security step does not count as a readiness success.
In the cross-border scenario, the total the customer will pay and the net amount the merchant will collect may be different calculations. Currency conversion, the shipping charge, who is responsible for fees that may arise at the border, and the return address are clarified. Undefined additional charges are not assumed to be zero. The delivery promise must rest on the verified conditions of the relevant route and carrier; the lead time in the domestic market must not be carried over to the foreign market unchanged.
Personal data, consumer disclosures, commercial contracts and country-specific obligations are evaluated by the business's relevant specialists. The technical report does not replace this evaluation; it makes visible which data and which operations move between which parties. When the current documentation, account approval and contract scope are verified together, the organization can decide on the basis of a concrete flow rather than a general country assumption.
Keeping a product mapping record for country and language versions is also important. The same model may be offered in different markets with different package contents, units of measurement or warranty terms. The translation process should not flatten these differences. The task set does not consist merely of translating Turkish questions into another language; it is prepared again with the address, date, size and currency formats of the target market. Each new market is evaluated on its own commercial terms before it is seen as a distribution area for the existing product.
35 Unit economics and the investment decision
The economic value of an Agentic Commerce program does not come from having connected to a new channel. It must be explained which task is completed with fewer errors, what it gains the customer, and which cost of the business it changes. Technical acceptance criteria and economic acceptance criteria complement each other. A reliable flow can be expensive; a cheap flow can raise the total cost because of incorrect orders.
Initial investment and ongoing expense must be separated. Organizing product data, integration development, the test environment, the task set and team training can be included in the initial investment. Model and tool usage, data services, monitoring, human evaluation, maintenance and support arise as long as the operation continues. When test, development and live usage expenses are kept apart, the pilot cost is not mistaken for the future transaction cost.
Failed tasks are not removed from the cost calculation. Calls that produced no result, retries and human review are consumed resources. At the same time, the same expense must not be counted twice under different headings. Repeat calls already included in the model cost are not written again under the name of failure cost; the review or intervention effort that is genuinely different can be added.
Illustrative contribution calculation
The calculation below is an entirely illustrative scenario created to explain the method. It is not a Webtures result, a price offer or an expected customer return. In the example, discount and return adjustments have been applied, and net revenue excluding taxes is taken as 1,000,000 TL. Product cost is also the amount after the accounting effects of sold and returned products have been processed.
| Line item | Illustrative amount | Calculation boundary |
|---|---|---|
| Net revenue | 1,000,000 TL | Discount and return amounts deducted |
| Product cost | 600,000 TL | Effect of returned goods processed |
| Shipping expense | 70,000 TL | Outbound logistics |
| Channel and payment expense | 40,000 TL | To be calculated with actual contract terms |
| Return operations | 20,000 TL | Return transport and handling expense |
| Customer support | 15,000 TL | Support cost allocated to this flow |
| AI operating expense | 55,000 TL | Calls, tools, evaluation, monitoring and maintenance |
| Contribution | 200,000 TL | Net revenue minus the listed expenses |
Total expenses come to 800,000 TL, and contribution to 200,000 TL. The ratio of contribution to net revenue is 20 percent. Because the return amount is deducted from net revenue, it is not expensed again; the table contains only the additional operating expense of returns. Since this calculation does not include the organization's overheads, financing and initial investment, it is not called net profit.
The investment decision must question how much of this contribution is genuinely incremental. Moving a sale that would have happened in another channel into the new flow does not turn the entire sale amount into additional gain. The effect is isolated using a comparable period, a customer group or a suitable experiment design. Scenarios of low usage and high intervention need are evaluated as well. A successful controlled test does not prove the economic outcome; it can justify a more comprehensive live evaluation.
The cash effect must be evaluated separately. The creation of the order, the collected payment becoming available and the closing of the return window can spread across different points in time. The organization must calculate the channel's payment terms, the effect of tied-up stock and the capacity it reserves for support using its own records. It is possible that a flow whose contribution looks positive requires additional working capital at the start. For this reason, the investment budget is not limited to the cost of software development alone; operating capacity during the transition period is also included in the plan.
36 The ninety-day pilot and the twelve-month roadmap
If the transformation program is designed as an integration project with a fixed end date, it can lose its validity as data and platforms change. For this reason, the purpose of the first pilot is both to improve specific tasks and to establish a reusable operating routine. The ninety days and twelve months here are illustrative planning horizons. Access rights, vendor support and the organization's change calendar determine the actual duration.
The first ninety days
In the first two weeks, the commercial objective and the final action stage are chosen. Work begins with one product group, specific task families and a limited environment. Data owners, technical leads and the person who will give customer acceptance are assigned. If the required documents or test access are missing, this period is not considered complete; the rest of the calendar is arranged according to the real dependency.
In the third and fourth weeks, the baseline assessment is carried out. Product data, commercial terms and task outcomes are examined together. The good results of easy tasks must not hide the failures of hard tasks. At the end of the period, management holds the target scope, the observed defects, the unknowns and the priority fixes. Unexpected scope growth is decided on separately.
In weeks five through eight, the selected fixes are implemented. For each piece of work, the owner, the acceptance criterion and the rollback plan are defined. Work that also benefits the existing customer, such as data consistency, is separated from integrations specific to a particular channel. This way, a channel that has not yet been approved does not stop the progress of the whole program.
In the final weeks, the same tasks are run again; error, refusal and human handoff scenarios are evaluated as well. Live monitoring is prepared for the accepted scope. The outcome is not automatically a rollout. Each of the options, expanding, staying in limited use, or pausing until a dependency is resolved, can be valid.
Expansion after the pilot
Between the fourth and sixth months, the data contracts, task definitions and incident records created in the pilot can be standardized. When moving to a new category, the existing tests are not copied directly; the category-specific cost of a wrong decision and its commercial terms are added. The increase in operating expense and human intervention is monitored.
Between the seventh and ninth months, if provider support and the economic rationale exist, a new channel or transaction stage can be evaluated. Each of the state-changing capabilities such as payment, cancellation and return requires separate acceptance. A flow verified in one country or account is re-examined in another market.
In the final quarter, the total contribution of the investment, the maintenance burden and the dependency on vendors are evaluated. Features that do not work do not have to be kept. The team decides which capability it will keep in-house, which service it will source externally and which development it will postpone. The year-end goal is not to make every transaction autonomous, but to operate the selected commercial tasks in a sustainable way.
The roadmap must also contain a rollback decision for each expansion. If a new category requires more human intervention than expected, it must be possible to temporarily narrow the scope. When the commercial terms of an integration change, the alternative flow that serves the customer is preserved. This way, the program does not become dependent on a single provider's calendar or on a trial feature. Technical capacity, team time and customer impact are handled together in the same expansion decision.
37 Corporate operating model and responsibilities
Because Agentic Commerce starts with product information and extends into price, stock, delivery and payment systems, it cannot be sustained through the ownership of a single team. At the same time, the phrase "shared responsibility of all teams" can make decisions unclear. There must be one accountable business owner for each commercial objective, one implementation lead for each system and one designated authority for each acceptance decision.
The commercial sponsor determines which customer need and economic outcome the program targets. The technology lead manages the integration boundaries, the environments and the release plan. The product data team owns the source and freshness of attributes. The pricing and operations teams verify that campaign, stock and delivery terms can actually be applied. The finance and payment teams monitor the consistency of commercial records and reconciliation.
A clear definition of decision rights
The authority to correct a product description and the authority to change a campaign do not have to sit with the same person. A change proposal prepared by AI does not remove this distinction either. The organization must define which change can be applied automatically, which one requires the business owner's approval and which one will remain only a suggestion. The authorization scope is kept at the narrowest limit appropriate for the task.
Customer service must not be seen as the last stage of the operating model. When an agent communicates a wrong condition, the customer usually ends up with the support team. This team must be able to see which information was presented at which moment and which transaction actually took place. Providing a clear incident summary without exposing unnecessary personal data can improve both the resolution time and the coordination inside the organization.
Change and exception management
A catalog, campaign or provider update can create the need for a new evaluation. Instead of running all tasks on every change, the affected flows are identified; if a critical shared component has changed, the scope is widened. When a test is repeated is tied not only to the calendar but to the nature of the change.
In exception management, a problem is owned according to its importance. A wrong description, a transaction inconsistency and an authorization violation must not be handled in the same queue at the same speed. Who holds the authority to stop a flow, and which evidence is required to reopen it, are determined in advance. A problem disappearing temporarily does not remove the need for retesting and recording.
The working rhythm can be built from short implementation meetings and less frequent management reviews. The implementation meeting resolves open defects and dependencies. The management review focuses on commercial contribution, cost and scope decisions. Team development does not consist only of tool training; it also covers data responsibility, interpreting uncertainty and handing off to a human at the right time.
When staff or a vendor changes, continuity of knowledge and authorization must be protected separately. Task definitions, data contracts and acceptance records must not remain in individual accounts. While the departing person's access rights are reviewed, the owner of open work is reassigned. A new team member must be able to understand from the records which area has been verified, which conditions are out of scope and which change requires retesting. This arrangement reduces the dependency on the people who ran the pilot.
38 The Webtures service model and concrete deliverables
The Webtures service approach in this area must be built as a transformation program that connects research to applicable changes. The core promise is a permissioned technical audit, controlled task testing, a prioritized fix plan and re-evaluation after each change. The proof of the value offered to the customer is not a general AI narrative; it is showing which task worked within which scope.
The first stage of the service is scoping. The customer's systems, product group, target channel and the transaction it wants completed are understood. A preliminary review based on publicly available information is kept separate from a permissioned system evaluation. A quick scan must not produce conclusions about detailed transaction reliability that go beyond its scope.
The delivery chain from evaluation to implementation
The full evaluation does not consist of an executive presentation alone. The records that technical teams can work from and the summary that the executive can decide on are prepared together. Each finding has the task it affects, its evidence, its priority and its acceptance criterion. Areas that could not be accessed are left out of the result; an unknown state is not allowed to look like a positive result.
| Deliverable | Content | Purpose |
|---|---|---|
| Scope and access record | Systems, permissions, environments and the final action | To define the boundary of the evaluation |
| Baseline status report | Task outcomes, data problems and open points | To choose the investment priority |
| Implementation backlog | Owner, dependency and acceptance criterion | To carry out the fixes |
| Retest package | Before and after results with remaining defects | To accept the change |
| Operations guide | Monitoring, exception, handoff and retest routine | To sustain quality |
In the implementation stage, data and content adjustments, integration work and payment infrastructure changes are planned as separate specialties. The work to be done by the Webtures team, the customer's team and the implementation partner must be visible in the contract. It must not be claimed that all developments are delivered ready-made before the available expertise and integration capability have been verified.
Commercial scope and acceptance
Pricing can be tied not only to page count or working hours, but to product scope, task variety, environment, access type, evaluation effort and retest frequency. A fixed-scope evaluation and development with variable dependencies must not be tied to the same price assumption. When additional scope is required, its impact on the customer is explained and the engagement is updated accordingly.
The purpose of the ongoing service is not to reproduce the report at set intervals. It is to enable the customer to intervene by monitoring critical data changes, transaction defects and cost deviations. Unnecessary alerts are reduced; who receives an alert and what is to be done are defined.
The readiness assessment is presented as an opinion given for specific conditions; not as a general certificate or a guarantee of success across all agents. The customer's purchase of a particular product is not made a condition of a positive assessment. This approach ensures that the commercial goals of the service and the credibility of the research are protected together.
At the end of the service, the customer must receive a handover package that allows the work to continue. The current task set, open defects, access owners, change history and retest method are transferred in an understandable form. When the ongoing service ends, it is explained which monitoring will stop and which responsibilities the customer will take over. The quality of the delivery is judged not only by the screens used during the consultancy, but by the customer's ability to keep making its own decisions with this knowledge.
39 The relationship between measurement and implementation with Brantial
The relationship that could be established between the Webtures consulting program and Brantial's product development can make it easier to monitor findings on a regular basis and to compare implementation results. The commerce tasks, new integrations and approved change flows described in this report are a product development proposal. They are not presented as existing Brantial capabilities or as completed customer implementations before the scope of use has been verified.
In the proposed design, visibility data and task verification complement each other. A product appearing in an AI answer can be the starting point of the research; the product being selected at the current price or reaching the correct cart requires different records. For this reason, visibility, information accuracy and transaction outcome must be shown separately on the product screen. Combining them into a single score can hide which improvement is needed.
Determining the development sequence
The first product step can be a workspace that links task and evidence records. The task definition, the data version used, the observed defect, the proposed change and the retest result are kept together. Such a foundation can support the consulting process without requiring direct order processing capability.
The second step is read integrations with the data sources the customer has permitted. The aim is to be able to track which records represent the same product in different systems and when the data changed. For each connection, the data scope, the access duration and the customer's responsible person are defined. It should not be necessary to transfer the entire catalog or customer data simply because it might be usable later.
At a later stage, moving fix suggestions into the workflow can be evaluated. The system can show an inconsistency in a product attribute and prepare a suggestion based on the authoritative record. Approving the suggestion, writing it to the target system and reading the result back are designed as separate steps. Actions with real effect, such as payment, order or price changes, require narrower authorizations and separate acceptance.
The limit of a credible product claim
In product promotion, the security certificate, hosting region, data processing scope and integration support must be used only after they have been verified with current documents. An item on the roadmap must not be moved into the list of existing features. A development made for a specific customer must not give the impression that it is a standard feature across all accounts either.
The measurement infrastructure must support the customer's ability to export its data and to examine the findings with its own teams. The Webtures report is not reduced to an uninterpreted copy of the Brantial screen. The consulting value remains in the responsibility of explaining the commercial context, choosing priorities and evaluating the outcome of the change.
When this relationship is set up correctly, product development can be guided by needs proven in the field. However, because no field testing was carried out within the scope of this report, feature priorities are not accepted as a verified distribution of demand. The needs to be learned from the first implementations must feed product investment decisions separately.
The features to be developed must also have their own acceptance criteria. For an alert module, not only the number of problems found is evaluated, but also the burden of false alerts and whether the problem reaches the right person. In a suggestion module, the applicability of the suggestion and the evidence it rests on matter. An increase in the number of integrations is not product success on its own; connections must stay current, their scope must be clear and disconnections must be reported in an understandable way. In this way, product measurement carries the same verification discipline as customer tasks.
40 Management decisions and preparing for the future
What management needs on Agentic Commerce is not faith in a single forecast of the future, but a decision structure that can adapt to different developments. The pace of market adoption, the commercial terms of the platforms and the willingness of customers to delegate authorization may each evolve in a different direction. The organization's readiness must rest on its ability to produce measurable work without hiding these uncertainties.
The first decision is which task will be improved. Finding a specific product correctly, preparing a valid offer and completing a transaction the user has approved are different goals. When management defines the goal explicitly, it limits unnecessary scope growth and makes clear which team must own which outcome.
The second decision is the economic and operational acceptance limit. The cost level at which the program continues, the type of error that will stop the flow and the situation in which human support is sufficient are determined in advance. If acceptance criteria are written only after a successful demo, the evaluation of the investment quickly becomes dependent on optimistic assumptions.
Investment that fits more than one future
In the scenario where agent-mediated commerce advances slowly, improving product data, offer consistency and support records still serves existing customers. In the scenario where specific platforms grow rapidly, channel access, commercial terms and data sharing gain importance. In a more open ecosystem, portable task definitions and integrations that can be adapted to different interfaces may prove valuable. These scenarios are not firm predictions; they are thinking tools for questioning the resilience of the investment.
The organization must apply the same discipline to the decision to build its own agent. Mastery of product data and customer needs does not automatically require operating a separate agent product. Such an investment carries a continuous burden of maintenance, deployment, support and reliability. It must first be shown in which task the existing options fall short and which knowledge of the business will make the difference.
Managing the learning
Management should not ask the program for a success rate alone. It should see the scope, the causes of failure, the areas that could not be verified and the commercial contribution together. Which assumption has changed in the most recent period, and how that affects the roadmap, should be evaluated regularly. A decision to stop progress or narrow the scope is a valid outcome of learning if it rests on evidence.
The core investment this report proposes is a verifiable working cycle: select a task with defined boundaries, understand the current state through a permissioned review, apply the approved change and retest the result. When economic evaluation and explicit accountability are added to the cycle, readiness stops being a one-time check.
If a sector comparison is to be published in the future, the purposes of private customer assessment and public research must be separated. Participation consent, the sampling method and the comparable scope are set from the outset. The results of a small volunteer group are not presented as the state of all of Türkiye. Customer rankings are not influenced by the purchase of services. The organization's research reputation depends less on producing a result that looks strong than on keeping clear which conclusion can be stated under which conditions.
Webtures' position in this field must be built on the expertise that carries the transformation of commerce infrastructure into practice. The strength of the claim must come not from the terms used, but from findings that make the customer's decision easier, from deliverables that can be implemented and from results whose limits are explicit. Lasting readiness means knowing what works today and determining how to respond when it changes tomorrow.
Appendix A Audit checklist
This list contains 48 checks for planning the first audit. It is not the official specification of any platform, nor a mandatory certification list for all businesses. For each check, applicability, observation, evidence, owner and next step are recorded. A check that does not apply and a check that fails are kept as separate cases. Critical defects are not offset by the number of completed checks.
Channel and authorization fit
- Definition of the target channel: It is written explicitly which of the discovery, handoff to the store, test cart or direct transaction paths is being evaluated.
- Account acceptance: The business's application, access and live-use status are verified through the relevant account. Being able to access the technical documentation is not evidence of acceptance.
- Geographic scope: Merchant account, buyer location and delivery region are kept as separate fields; it is not assumed that they are the same country.
- Product eligibility: Conditions such as product category, product condition, personalization and subscription are compared against the scope of the target channel.
- Authorization scope: The permissions granted for actions such as reading, creating a cart, completing an order and returning are examined separately.
- Test limit: Environment, time window, call limit, financial impact and the person responsible for termination are recorded before the trial begins.
Product discovery and identity
- Need-based discovery: It is checked that suitable products can be found for realistic requests that contain no brand name.
- Merchant identity: It is verified that offers from different merchants for the same product are not mixed and that the relevant conditions are attached to the correct merchant.
- Product family and variant: It is tested that identity and the commercial record remain consistent when color, size, capacity or pack quantity changes.
- Identity mapping: It is checked, using the mapping record, that the PIM, ERP, feed, page and checkout identifiers point to the same product.
- Language and unit of measure: It is examined whether changes in Turkish wording, model code, local size system and unit of measure cause a loss of meaning.
- Missing-information behavior: When a product attribute is unknown, it is evaluated whether the system asks for clarification or states the uncertainty instead of fabricating certainty.
Data accuracy and freshness
- Domain owner: The system that is the deciding source for price, stock, product attributes and delivery is recorded.
- Deployment consistency: The visible page, structured data, feed and API response are compared against the product record of the same date.
- Update delay: The time at which a source change appears in the targets is measured; alongside the average, the changes that are delayed and the changes that never arrive are shown.
- Meaning of stock: Physical stock, sellable stock, reserved stock and stock suitable for the delivery region are separated from one another.
- Version and time: Data version, source time and transfer time are kept separately; the last transfer date alone is not relied upon.
- Inconsistency resolution: The behavior to use when sources conflict, the responsible team and, where necessary, the rule for stopping the task are defined.
Offers and commercial terms
- Total cost: The total the customer will pay is recalculated with the product, mandatory fees, shipping and applicable discounts.
- Campaign eligibility: It is checked that no discount is promised before the user, product, cart, time and payment method conditions are met.
- Delivery validity: It is verified that the date and fee depend on the real address, the stock location and the shipping option.
- Offer validity period: It is tested that an expired or changed offer is regenerated and, where required, resubmitted for approval.
- Returns and warranty: The condition, coverage and exceptions of the relevant product are matched with the correct commercial record.
- Commercial limits: It is examined that quantity, spend and margin rules are applied in line with the authorized business policies.
Task and transaction execution
- Task contract: Mandatory conditions, preferences, permitted tools and the expected final state are written before the test.
- Tool response: It is checked that success, error and missing-information cases steer the continuation of the operation correctly.
- Cart state: It is verified that the product, quantity and total reported by the agent match the real cart record.
- Duplicate request: It is tested that the same business intent does not create an extra order because of a network error or a retry.
- Timeout: It is checked that when a response is lost, the current state is queried first and no blind repeat is made with a new financial transaction.
- Transaction closure: It is verified that the order, payment and stock states are consistent with one another or carry an explicit pending state.
Security and user control
- Authentication: It is evaluated that the identities of the agent provider, the user and the business account are not confused with one another.
- Least privilege: It is tested that each tool accesses only the data and actions required for the task.
- Authorization revocation: It is checked that new state-changing operations stop after revocation or expiry; reconciliation access is handled separately.
- Sensitive data: It is examined that personal data, access keys and payment information do not leak into model output or unnecessary logs.
- External content steering: In the permissioned test environment, it is verified that product or review content does not change the system rules.
- Human handoff: It is tested that when additional verification is required, the user can take over in an understandable way and the scope of the approval is preserved.
Measurement and reconciliation
- Shared transaction identifier: A traceable link is established between task, offer, cart, payment and order.
- Traffic evidence: Observed, technically verified, inferred and unknown agent relationships are labeled separately.
- Webhook processing: It is tested that signature, redelivery, out-of-order arrival and delay cases are handled correctly.
- The right denominator: It is checked that failed or ambiguous eligible attempts are not excluded from measurement merely because their outcome was negative.
- Cost record: Alongside successful tasks, the cost of failed attempts, retries and human support is also accounted for.
- Outcome separation: Discovery, cart, order, collection, net revenue and incremental revenue are tracked with different indicators.
Operations and improvement
- Responsible team: The person or team who will decide is known for every data domain, integration and finding.
- Change record: Changes in model, tool, platform, catalog and commercial rule versions are tracked.
- Retest: After a fix, the relevant task and the neighboring flows that could be affected are re-evaluated.
- Incident management: Responsibilities for stopping on a critical defect, determining customer impact, reconciliation and reopening are defined.
- Rollback plan: How the existing customer journey will continue when the new integration is taken out of service is determined.
- Regular review: Scope, acceptance criteria and economic results are re-evaluated at set intervals.
Appendix B Implementation cards for twelve task families
The tasks below form the starting library of the proposed pilot. Test data and monetary amounts are illustrative. In each business, the product, channel, permissions and verification method are defined anew. A task family may contain several positive and negative scenarios that test the same purpose. The 12 families in the pilot matrix can be selected from this set. In the 720-execution calculation in section 28, a single task contract is selected for each family and each eligible access path and run three times; not all of the sub-scenarios described here are included in that count.
Reaching the right product from a need
The task is to find options that fit a usage need without giving an explicit product name. For example, a suitable docking station can be searched for a customer who wants to use two displays with a specific computer. For success, manufacturer-based compatibility and the required connection features are verified. A similar name or a high rating is not sufficient on its own. Which information could not be found is recorded. If no suitable product exists, saying so is the correct outcome; silently changing the user's firm condition is not accepted.
Selecting a variant and a quantity
The task is to select the specified color, size, capacity or pack quantity within a product family. The trial follows the same identifier from the product page through the feed to the test cart. The evidence of success is not the product photo but the authoritative variant record and the cart line. If the requested variant is out of stock, a close variant is not selected without the user's approval. An order in multiple quantities is evaluated together with sellable stock and pack count.
Comparing alternatives
The task is to compare options that meet the same need using shared criteria. Using the price with shipping for one and the price without shipping for another is a false comparison. Comparison criteria are set before the test; fields with no data are shown as empty or uncertain. A quality score the agent produced itself is not presented as a manufacturer attribute. The result must explain the user's priorities and the real differences between the options.
Preparing an offer within a budget
The task is to prepare an offer that does not exceed a specified total, including mandatory fees. At an illustrative limit of 4,000 TL, a product priced at 3,900 TL does not mean it fits the budget once 150 TL of shipping is added. A campaign code becoming invalid and the currency changing are separate negative scenarios. Success is evaluated on the current total and the offer validity period. If the agent exceeds the budget, it must obtain new authorization from the user or stop the operation.
Meeting a delivery condition
The task is to find a delivery option that fits a specified address and date. Business days are separated from calendar days, and preparation time from transit time. The generic delivery text on the product page does not replace a valid offer for the real address. Success is verified with an applicable delivery option and the related conditions. When no suitable delivery is found, silently accepting a later option is a failure; an explicit handoff must be made for the alternative.
Creating a test cart
The task is to add the approved products and quantities to the sandbox cart. Because the expected final state is the cart, initiating a payment or an order is a scope violation. A retry belonging to the same business intent must not increase the quantity without being asked. The effect of a separate and authorized add request is evaluated separately. The product identifiers, price, discount and delivery conditions in the cart are compared against the task contract. If the cart expires, the system must state clearly how the new state will be created.
Handing the purchase over to the user
The task is to give control to the user after research or cart preparation. The user must continue with the correct store and the correct variant. Opening a link is not on its own a successful handoff; it is verified that the selected quantity, the valid offer and the required context are preserved. Session limits, the in-app browser and login requirements are tested. If the user has to make the selection again, this friction is recorded.
Completing an authorized transaction
The task is run only in a suitable test environment and with explicit authorization. The approved merchant, product, amount, currency and time limit are checked. Success is the payment and order records matching the targeted final state. Obtaining a token or a positive API response is not on its own evidence of a purchase. If additional verification is requested, a handoff to the user is expected. Authorization hold, collection and order acceptance are kept as separate outcomes.
Adapting to stock and price changes
While the task is in progress, the price or stock in the permissioned test data is changed. The aim is to test whether the system notices the change and produces a valid offer. Displaying the old price and transacting at the old price are different defects. Success is retrieving the current information and, if the change exceeds the scope of the approval, requesting permission again. When no suitable option remains, the task must stop. The retest covers not only the changed product but also similar variants.
Recovering safely from interruption and retry
During the task, a lost response, a delay or a redelivery is created in a controlled way. The agent is expected to learn the current state first and not to generate an extra operation for the same business intent. Payment and order records are compared using the transaction identifier. Success requires both a correct final state and side effects that stay within the limit. An order that was created twice after a lost response, with one later canceled, is not reported as an error-free transaction.
Stopping a request that exceeds authorization
This negative test family covers examples such as another customer's account, an expired permission, a different merchant, a budget overrun or a revoked authorization. The expected outcome is a safe refusal or an appropriate human handoff. The refusal itself is not an error. Success is verified by the absence of any unauthorized state change. Stopped attempts are not added to the positive purchase success rate; they are reported as a separate security indicator.
Executing an after-sales task
The task may be to learn the status of the test order, to prepare a cancellation request under suitable conditions or to execute an authorized partial return. First, the relationship between the user and the order and the scope of the operation are verified. Receiving the return request, accepting the product and sending the money back are different states. Success is determined by the state the task targets. An operation applied to the wrong line item or the wrong amount is counted as a failure even if the customer message looks correct.
Appendix C Implementation records and acceptance templates
The templates are used to make the assessment transferable between different teams. Reference identifiers are written instead of sensitive information. Task records must not turn into a new repository that creates unnecessary copies of customer data.
Scope record
A scope record defines the business, the business goal, the targeted final action, the product group, the channel, the environment and the time period. Permitted access and permitted actions are written separately. The points where the environment differs from production are stated. The person granting the authorization, the test owner and the person responsible for stopping are linked to the record. Areas where access could not be obtained are left open; they are not marked as positive results.
Finding record
| Field | Information to record | Function for acceptance |
|---|---|---|
| Finding identifier | Unique record and date | Merging duplicates |
| Affected task | Condition, product and channel | Determining customer impact |
| Expected and observed | Two concrete behaviors | Reproducing the defect |
| Evidence | Transaction and system record reference | Verification independent of the model's statement |
| Impact and priority | Financial, data or experience impact | Making the critical blocker visible |
| Owner and change | Team, work package and dependency | Starting the implementation |
| Retest | Same task and related checks | Making the closure decision |
Illustrative finding example
Assume that although the blue 256 GB variant of a product was selected, the black 128 GB variant was added to the test cart. This is not a real customer finding. The evidence is that the selection record and the cart line show different variant identifiers. The first investigation turns to the mapping between the URL parameter and the checkout identifier. After the fix, the same selection and the other variants are tested again. Correcting the displayed color is not enough; the product identifier in the cart must also be correct.
Go-live decision record
The decision record contains the products, users and channels in scope and the permitted transaction limit. It states whether any open critical defect exists, the rollback method, the observation period and the incident owner. The conditions under which the pilot will be expanded or stopped are written in advance. The people making the decision approve their own responsibilities for technical readiness and commercial acceptance separately.
Post-change assessment record
When the model, the protocol version, the feed schema or a campaign rule changes, the affected task families are identified. The change is first evaluated in a controlled environment. The results are compared against the same acceptance criteria. A rise in the success rate does not close a security defect; a cost reduction alone does not justify an increase in the wrong-product rate either. The change decision is made by viewing task quality, user authorization and economic outcome together.
Appendix D Shared vocabulary and frequently asked management questions
Core terms
Agent: A software system that can work with tools and data sources to accomplish a specific purpose. Its scope of authorization varies by task.
Task contract: The record in which the need, constraints, permissions, environment, expected outcome and verification method are written together.
Business intent identifier: The identifier that allows retries made for the same customer goal to be linked to a shared commercial transaction.
Offer: The presentation of a specific product, quantity, price and commercial terms together with their validity limits.
Feed: The bulk or scheduled transfer of product information in line with the data contract a channel expects.
Authoritative source: The system the business accepts as correct for a data domain, together with the record of responsibility. The source may not be the same for every domain.
Sandbox: The environment in which tests are run with controlled data and limited effects. It does not by itself guarantee the same behavior as production.
Idempotency: The behavior that prevents the repeated execution of the same request, under defined conditions, from producing an additional effect. Its scope and duration depend on the contract.
Reconciliation: The comparison of task, payment and order records to determine which real outcome actually occurred.
Human handoff: The transfer of the task to the user or an authorized operator with the necessary information and control.
Tokenization: The use of a representation with specific usage rules in place of sensitive payment information. It is not synonymous with a general spending permission.
Observability: The ability to explain the behavior of the system through events, records and measurements.
Is a good website enough on its own
A good web experience is a valuable foundation. However, product identity, current offers, permissioned tool access and transaction reliability must be evaluated separately. Conversely, the existence of a technical API does not make the handoff-to-the-user experience unimportant. The evaluation is made across the whole of the targeted path.
Should every business build its own agent
No. For some businesses, offering an accurate catalog and transaction capabilities to existing channels may be sufficient. The decision to build one's own agent must be justified by the task to be solved and the operating cost. Building a new interface does not replace solving product data problems.
Is it necessary to implement all protocols
No. Protocol choice should follow the target channel, the capability and the provider's conditions. Unused integrations can create a maintenance and security burden. The task to be supported is determined first, and then the technical contracts required for that task are selected.
Would a single readiness score be enough
A single score can make the progress narrative easier, but it must not hide the scope and the critical defects. A payment authorization violation cannot be offset by a high content score. The decision must be made together with the channel, product group, task, level of evidence and open blockers.
Does the work end once readiness is complete
No. Changes in products, stock, campaigns, models and platforms can affect a task that worked before. Retesting and a structure of responsibility are part of readiness. The frequency of checks is set according to the rate of data change and the impact of the transaction.
Which decision should be made in the first management meeting
The first decision, ahead of any goal of automating everything, is which customer task will be evaluated and in what scope. Alongside this decision, the data owner, the technical lead, the permission limit and the evidence of success must be determined. Subsequent investment can be expanded on the basis of the concrete needs this task reveals.
Growth & GEO