Ana içeriğe geç
Versiyon: 1.0.6

Data Generation

Generate realistic test data using AI-powered synthetic data generation, ensuring your tests have the data they need.

Data Generation

Coming Soon

This feature is currently in development.

Generation Methods​

AI Synthetic Data​

Let AI create contextually appropriate data:

FeatureDescription
Context-AwareUnderstands data relationships
RealisticMimics real-world patterns
DiverseVaried data distribution
ConsistentMaintains data integrity

Faker Library​

Built-in fake data generators:

CategoryExamples
PersonName, email, phone, address
CommerceProduct, price, company
InternetURL, IP, username
DatePast, future, range
FinanceCredit card, IBAN, currency
LocationCountry, city, coordinates

Custom Rules​

Define your own generation logic:

Rules:
order_total:
type: calculated
formula: 'sum(items.price * items.quantity)'

discount:
type: conditional
when: 'order_total > 100'
then: 'order_total * 0.1'
else: 0

status:
type: weighted
values:
completed: 70
pending: 20
cancelled: 10

Generation Profiles​

Quick Generation​

Simple data creation:

Profile: Basic User
Generate:
- name: faker.name()
- email: faker.email()
- phone: faker.phone()

Advanced Generation​

Complex data with relationships:

Profile: E-commerce Order
Generate:
User:
count: 100
fields:
name: faker.name()
email: faker.email(unique=true)

Product:
count: 500
fields:
name: faker.commerce.productName()
price: faker.price(10, 1000)

Order:
count: 1000
fields:
user_id: random(User.id)
status: weighted(completed:70, pending:30)
created_at: faker.date.past(30)

OrderItem:
count: 3000
fields:
order_id: sequential(Order.id, 1-5)
product_id: random(Product.id)
quantity: random(1, 10)

Data Patterns​

Distribution Patterns​

PatternDescriptionUse Case
UniformEqual probabilityRandom selection
NormalBell curveNatural variation
WeightedCustom probabilitiesStatus distribution
SequentialOrdered valuesIDs, dates

Realistic Patterns​

# Age distribution matching demographics
age:
type: normal
mean: 35
std: 15
min: 18
max: 80

# Purchase amount with realistic skew
amount:
type: lognormal
mean: 50
std: 30

# Time-based patterns
created_at:
type: time_series
pattern: business_hours
timezone: UTC

Constraints & Validation​

Uniqueness​

email:
type: string
generator: faker.email()
unique: true
retry: 10

Referential Integrity​

order:
user_id:
reference: users.id
on_missing: create # create, skip, error

Custom Validation​

age:
type: integer
validate:
- 'value >= 18'
- 'value <= 120'

email:
type: string
validate:
- regex: "^[a-z]+@[a-z]+\\.[a-z]+$"

Bulk Generation​

Large Datasets​

Generate millions of records efficiently:

SizeStrategy
< 10KIn-memory
10K - 1MBatched
> 1MStreaming

Performance Options​

Generation:
records: 1000000
batch_size: 10000
parallel: true
workers: 4
output: streaming

Progress Tracking​

  • Real-time progress
  • Estimated completion
  • Error reporting
  • Pause/resume capability

Templates​

Built-in Templates​

TemplateDescription
User AccountStandard user profile
E-commerce OrderOrder with items
Financial TransactionPayment records
Healthcare PatientMedical records
Inventory ItemProduct inventory

Custom Templates​

Create reusable generation templates:

Template: Banking Customer
Version: 1.0
Fields:
account_number:
type: string
pattern: '[A-Z]{2}[0-9]{18}'

balance:
type: decimal
min: 0
max: 1000000

account_type:
type: enum
values: [checking, savings, investment]

Export Options​

Formats​

FormatBest For
JSONAPI testing
CSVDatabase import
SQLDirect database insert
ExcelManual review
ParquetBig data

Streaming Export​

For large datasets:

  • Direct database insert
  • File streaming
  • API endpoint delivery

Best Practices​

Data Quality​

  • Validate generated data
  • Check distributions
  • Verify relationships

Performance​

  • Use appropriate batch sizes
  • Enable parallel generation
  • Stream large datasets

Reproducibility​

  • Use seed values
  • Save generation profiles
  • Document configurations