The SELJI Method
At SELJI.com, every recommendation is grounded in measurable data, not marketing claims. Where possible, we conduct hands on testing and direct product reviews to validate real world performance. Our proprietary review system then combines these results with AI, Natural Language Processing (NLP), and the Amazon Product Advertising API 5.0 to turn raw consumer feedback and product metrics into transparent, evidence based scores that shoppers can trust.
🧩 Data Sources
We aggregate and validate data from multiple independent, verifiable channels, then cross check them against one another.
Amazon Product Advertising API 5.0
Real time access to official Amazon data: pricing history, technical specifications, verified review metadata, and stock or variant identifiers. Version 5.0 improves categorization, price tracking, and attribute consistency across regions.
Verified User Reviews
Real feedback from major marketplaces and forums, including Amazon, Best Buy, Walmart, and Reddit.
Expert & Lab Testing
Independent evaluation and benchmark reports from third party testers.
Manufacturer Data
Supplied specifications and firmware update logs, tracked over time.
SELJI In-House Testing
Our own measured results, added wherever hands on testing is possible.
By combining structured data (API 5.0) with unstructured human feedback, we achieve both breadth and precision in analysis.
🧹 Data Cleaning & Normalization
Before any scoring occurs, every input passes through a rigorous cleaning process.
Deduplication & Canonical Mapping
Merge duplicates across SKUs, regions, and rebrands into a single product identity.
Fraud Detection
Identify synthetic or incentivized reviews through linguistic and temporal anomaly detection.
Unit Normalization
Convert disparate metrics (Pa, dB, mAh, Wh) into standardized, comparable units.
Version Control Alignment
Tag every dataset with firmware build numbers or release identifiers for fair, time consistent comparisons.
Only verified and time stamped data proceeds into the scoring pipeline.
🧠 Feature & Sentiment Extraction
Our NLP pipeline isolates meaningful performance signals from millions of words of user feedback.
Aspect Based Sentiment
Scores specific attributes such as suction strength, battery life, and noise level.
Negation & Contrast Detection
Correctly reads phrases like “not quiet” or “better than before the update.”
Weighted Trust Modeling
Verified buyers and consistently reliable reviewers carry more influence.
Topic Clustering
Surfaces recurring reliability and usability patterns across many reviews.
This creates a multidimensional feature map describing what users genuinely experience.
⚙️ Scoring Model Architecture
Each product is evaluated through a Category Specific Scoring Matrix that combines quantitative metrics, sentiment polarity, and confidence intervals. Weights below are the share each pillar contributes to the composite score.
Example: Smart Vacuum Category
Performance across mixed surfaces.
Docking, self cleaning, and refill systems.
AI object avoidance and mapping precision.
Durability and acoustic performance.
Price to performance and maintenance costs.
Connectivity and software reliability.
Confidence intervals are computed through bootstrap resampling to reflect data stability and reviewer variance.
🧮 Firmware as a Living Variable
Many modern products are software driven devices whose behavior evolves through firmware updates. In SELJI’s system, firmware is treated as a dynamic performance factor, not a static specification.
Performance Evolution
Each firmware build can alter suction, navigation, or power efficiency. We detect these shifts through time based sentiment changes and API 5.0 metadata.
Reliability Tracking
Update cadence, regression frequency, and fix latency feed our Product Stability Index (PSI). Brands that deliver consistent, stable updates earn higher reliability weight.
Version Normalization
Reviews and metrics are aligned to the latest stable firmware. Older data tied to outdated builds is down weighted so obsolete flaws do not skew results.
Firmware Confidence Multiplier (FCM)
Each product receives a multiplier between 0.8 and 1.1 based on update quality. Stable, improvement oriented firmware pushes scores upward, while erratic or regressive updates reduce them.
AdjustedScore = BaseScore × FCMBy integrating firmware behavior, SELJI scores always reflect current real world performance, not the launch day snapshot.
📊 Cost & Longevity Analysis
We quantify long term value using both static specs and dynamic consumption data.
Total Cost of Ownership
Consumables, energy, and accessory costs projected over 12 to 36 months.
Price Trajectory & Promo Frequency
Pulled directly via Amazon API 5.0 for precise MSRP trends.
Warranty & Return Rates
Derived from aggregated customer service data and sentiment.
Every recommendation reflects durability, affordability, and lifecycle value.
🔬 Reliability & Durability Modeling
Reliability is modeled statistically through our Product Stability Index (PSI).
Survival Curve Analysis
Tracks defect and return mentions over the product lifespan.
Post Update Trend Detection
Flags complaint spikes that follow specific firmware releases.
Weighted Historical Variance
Measures consistency across review periods, not just a single moment.
Firmware Event Markers
Isolate software related performance changes from hardware failures.
The PSI ensures each product’s reliability score mirrors its real trajectory over time.
🧪 Hands-On Validation
Whenever possible, SELJI performs its own testing to ground the data model in physical reality.
Noise Testing
Calibrated dBA measurements in controlled environments.
Cleaning Efficiency
Standardized debris compositions across multiple surface types.
Navigation Tests
Timed obstacle courses to evaluate mapping precision.
Durability Cycles
Simulated long term use to measure suction and battery decay.
Measured results feed back into the database to continuously refine scoring accuracy.
📈 Ranking Transparency & Governance
No Paid Placement
Affiliate relationships never alter scores or rankings.
Dynamic Re-Weighting
Continuous data ingestion keeps rankings aligned with reality.
Audit Trail
Every score change is logged with its cause and a timestamp.
Reproducibility
The same dataset through the same version yields identical results.
🔍 Why It Matters
Most review sites summarize opinions. SELJI quantifies them. By combining live API data, NLP based sentiment modeling, firmware tracking, and human verified testing, we turn the chaos of online reviews into clear, defensible evidence, empowering shoppers to make confident, data backed choices.




