Skip to main content
T-Brain learns from your product information and shopper activity. You can upload data from your existing systems or, if you are a Topsort customer, use datasets already synced from Topsort. To get started, provide your catalog and at least one month of event history, including all purchases and paid product clicks and impressions.

Dataset types

A dataset is a named collection of data used to train a model. T-Brain supports three dataset types: You can name datasets to distinguish their purpose or source, such as products-main-store or events-september.

Upload data

  1. Go to Data.
  2. Click Add dataset.
  3. Select Categories, Products, or Events.
  4. Give the dataset a name.
  5. Upload your Parquet files and follow the instructions.
T-Brain displays the required structure for each dataset type and verifies that your files comply with it. The tables below explain the columns shown in the upload requirements. Fields marked Required must be populated. Other fields provide product information, relationships, or context needed for your use case.

Categories

Category uploads are snapshots. Each upload should contain the full set of categories. When a path is provided, its final segment must equal category_id. For example, a category with ID air-fryers could have the path home.kitchen.air-fryers. Path segments may contain letters, numbers, underscores, and hyphens. The path can be null for a root category or a catalog without a hierarchy.

Products

Product uploads are snapshots. Each upload should contain the full set of products. Although only product_id is structurally required, include descriptive information where available so the model has useful information about your products.

Events

Provide at least one month of event history. Include all purchases, not only purchases attributed to advertising, plus paid product clicks and impressions. Events must include a consistent shopper identifier so interactions and purchases can be connected. In the dataset, this identifier is stored in user_id and must remain stable across sessions and days. Event uploads are additive, allowing you to add activity from additional periods.

Supported event types

Keep identifiers consistent

Use exact, consistent identifiers across datasets:
  • An event’s product_id must match the product dataset’s product_id.
  • Values in a product’s category_ids must match category records’ category_id.
  • Use the same user_id for the same shopper across interactions and purchases.
  • Give each purchase line a unique event_id, and use order_id to group lines from the same order.

Refresh an uploaded dataset

In Data, click Add Files next to the dataset you want to update. For products and categories, provide a full snapshot. For events, add the new event records. Once the dataset is updated, retrain any model that should learn from the refreshed information. Adding files does not automatically refresh an already trained model.

Use synced Topsort datasets

If you are a Topsort customer, three datasets will already be available:
  • Topsort-categories
  • Topsort-products
  • Topsort-events
These datasets update in real time. You can select them when creating a task instead of uploading the same information yourself. Real-time dataset updates do not automatically retrain deployed models.

Train with your datasets

Go to Tasks, choose what you want the model to do, and select the relevant datasets. For example, a product-ranking task uses products and events. Start training, follow progress in Training, and review model quality in Evals before deployment.

Data privacy and access

Your uploaded data and fine-tuned models are isolated per retailer. Retention and deletion policies can be configured based on your requirements. Talk to a Topsort sales representative to arrange access. API documentation is available on request, including authentication, request and response formats, latency, and rate limits.